EP4728508A1 - Methods, apparatus, and systems for processing audio scene information - Google Patents
Methods, apparatus, and systems for processing audio scene informationInfo
- Publication number
- EP4728508A1 EP4728508A1 EP24731910.6A EP24731910A EP4728508A1 EP 4728508 A1 EP4728508 A1 EP 4728508A1 EP 24731910 A EP24731910 A EP 24731910A EP 4728508 A1 EP4728508 A1 EP 4728508A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- item
- path information
- path
- voxel
- indication
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/167—Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/03—Application of parametric coding in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Mathematical Physics (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
The disclosure relates to methods of processing audio scene information. One such method comprises: obtaining a voxel-based audio scene representation of an audio scene; for first and second locations in a two-dimensional voxel grid, successively encoding items of path information, each item of path information specifying a first location, a second location, a path length of an acoustic path, and a corner voxel on the acoustic path; and for a current item of path information, generating an encoded item of path information based on the item of path information. The encoded item of path information includes an indication of the respective first and second locations. If the corner voxel specified by the current item of path information is different from a comer voxel specified by a preceding item of path information, the encoded item of path information includes an indication of the corner voxel; If the comer voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information includes an indication that the comer voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, instead of the indication of the comer voxel. The disclosure further relates to corresponding apparatus, computer programs, and computer- readable storage media.
Description
METHODS, APPARATUS, AND SYSTEMS FOR PROCESSING AUDIO SCENE INFORMATION
Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63/508,367, filed 15 June 2023, which is incorporated by reference herein in its entirety.
[0002] The present disclosure relates to techniques of processing audio scene information, for example for storage, transmission, and/or audio rendering. In particular, the present disclosure is directed to voxel-based scene representations and to coding of path information (e.g., acoustic path information, such as diffraction path information) for voxelbased scene representations.
Background
[0003] The Moving Picture Experts Group (MPEG) is an alliance of working groups established jointly by the International Organization for Standardisation (ISO) and International Electrotechnical Commission (IEC), that sets standards for media coding, including audio coding. MPEG is organized under ISO/IEC SC 29, and the audio group is presently identified as working group (WG) 6. WG 6 is currently working on a new audio standard (also known as MPEG-I Immersive Audio, ISO/IEC 23090-4).
[0004] The new MPEG-I standard enables an acoustic experience from different viewpoints and/or perspectives or listening positions by supporting scenes and various movements around such scenes, such as movements using various degrees of freedom such as three degrees of freedom (3DOF) or six degrees of freedom (6DoF) in Virtual reality (VR), augmented reality (AR), mixed reality (MR) and/or extended reality (XR) applications. A 6 DoF interaction extends a 3 DoF spherical video/audio experience that is limited to head rotations (pitch, yaw, and roll) to include translational movement (forward/back, up/down, and left/right), to allow for navigation within a virtual environment (e.g., physically walking inside a room), in addition to the head rotations.
[0005] For audio rendering in VR, AR, MR and XR applications, object-based approaches have been widely employed by representing a complex auditory scene as multiple separate audio objects, each of which is associated with parameters or metadata defining a location/position and trajectory of that object in the scene. Alternatively audio rendering in such environments also uses higher order ambisonics (HO A). However, a new usage of “voxels” for rendering audio scenes is now being explored, such as for use of new immersive audio experiences. Voxels for audio rendering are relevant for media environments implemented in both hardware and software, such as video game and/or VR, AR, MR and XR environments.
[0006] A Voxel is a space volume with acoustic properties or audio rendering instructions assigned to it. Voxel size may be an encoder configuration parameter, and it can be (manually or automatically) selected according to a scene geometry level of details (e.g., in the range of 10 cm - 1 m).
[0007] Voxels for audio rendering can be obtained by:
• voxelization (or conversion) of a mesh-based scene representation
• from scene representation used for scene generation (or even video rendering) (e.g., by down-sampling of voxels of smaller size)
[0008] However, conventional approaches for providing realistic sound for user experiences (including those involving movement) in VR, AR, MR and XR environments using voxels still remain challenging and computationally complex.
[0009] Typical techniques for diffraction modeling in three-dimensional audio scenes, such as for computer-mediated reality applications, require re-calculation of diffraction paths and other diffraction information whenever any of the audio scene, the user location, or the audio source location change. For example, the diffraction path may change when the user and/or the audio source move through the three-dimensional audio scene. Further, the diffraction path may change when the audio scene itself changes, for example by indicating a door or window that opens or closes, or the like. Frequent re-calculations of diffraction paths may be computationally expensive, which requires comparatively powerful computation devices for implementing computer-mediated reality applications and/or may negatively affect user experience in some
cases. On the other hand, storing pre-computed acoustic paths (e.g., acoustic diffraction paths) may require large amounts of storage and/or bandwidth.
[0010] For example, US 11,606,662 B2, US 6,313,841 Bl, US 10,275937 B2, and US 10,293,259 B2 each relate to (pre-)computation of paths in a given scene or environment. However, the amount of data generated by doing so, in particular for larger numbers of paths, may be comparatively large, in particular when lossless transmission or storage of the path information is of interest.
[0011] There is thus a need for improved techniques for coding path information (e.g., acoustic path information, such as diffraction path information) in voxel-based audio scenes. There is a particular need for such techniques that can reduce bandwidth or storage requirements for handling the encoded path information.
Summary
[0012] In view of this need, the present disclosure provides methods of processing audio scene information (in particular, voxel-based audio scene information), apparatus for processing audio scene information, computer programs, and computer-readable storage media, having the features of the respective independent claims.
[0013] One aspect of the present disclosure relates to a method of processing audio scene information, for example encoding audio scene information, in particular, acoustic path information. The method may include obtaining a voxel-based audio scene representation of an audio scene. The method may further include, for one or more (e.g., plural) first locations and one or more (e.g., plural) second locations in a two-dimensional voxel grid relating to the voxelbased audio scene representation, successively encoding items of path information for respective first locations and second locations. The voxel grid may relate to a two-dimensional projection map generated from the voxel-based audio scene representation. Each item of path information may specify a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel on the acoustic path for which the acoustic path changes direction. The corner voxel may be a voxel for which a direct line to the second voxel is not occluded. The method may further include, for a current item of
path information, generating an encoded item of path information based on the item of path information. The encoded item of path information may include an indication of the respective first location and an indication of the respective second location. If the corner voxel specified by the current item of path information is different from a corner voxel specified by a preceding item of path information, the encoded item of path information may include an indication of the comer voxel (e.g., location of the comer voxel). This indication of the corner voxel may be an absolute, or non-differential indication of the corner voxel, for example using voxel coordinates or a voxel index. On the other hand, if the comer voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, the encoded item of path information may include an indication that the corner voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, instead of the indication of the corner voxel.
[0014] By coding items of path information in sequence and re-using information of preceding items of path information, the proposed method can reduce bitstream and storage requirements for handling encoded path information. Still, the original path information can be fully recovered, i.e., the proposed method provides for lossless coding of path information in voxel-based scenes.
[0015] In some embodiments, the encoded item of path information may further include an indication of the path length. The indication of the path length may be an absolute or nondifferential indication of path length, for example in units of voxels or voxel edge length.
In some embodiments, if the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, the encoded item of path information may further include an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
[0016] Thereby, the amount of data needed for encoding the path information can be further reduced, while still allowing for lossless recovery of the original path information.
[0017] In some embodiments, if the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, the encoded item of path information may further include an indication of whether the encoded
item of path information includes an indication of the path length or an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information. In this situation, the encoded item of path information may further include the indication of the path length or the indication of the difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
[0018] Thereby, the data needed for encoding the path information can be further reduced, overall, while still allowing for lossless recovery of the original path information.
In some embodiments, the method may further include, for each first location, traversing the voxel grid according to a predetermined pattern to determine a sequence of second locations. The method may further include, for the respective first location, successively encoding items of path information for the determined sequence of second locations.
[0019] In some embodiments, the predetermined pattern may traverse the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
[0020] In some embodiments, the difference may be encoded by 2 bits. Alternatively, the difference may take one of four predetermined values (e.g., potential values of the difference may be limited to a set of four distinct values).
[0021] Using the above pattern for traversing the voxel grid, it can be ensured that once the comer voxel does not change from one item of path information to the next, the difference between respective path lengths can only take one of four values, which can be encoded with high efficiency, using only two bits.
[0022] In some embodiments, each encoded item of path information may include an indication of whether a first mode or a second mode is used. Each encoded item of path information may further include an indication of the first location. Each encoded item of path information may further include an indication of the second location. Each encoded item of path information may yet further include an indication of whether an acoustic path exists for the first location and the second location. Each encoded item of path information may further include, if the first mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner
voxel specified by a preceding item of path information. Further in the first mode, each encoded item of path information may further include, if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of the path length. Further in the first mode, each encoded item of path information may otherwise include the indication of the comer voxel and the indication of the path length. Each encoded item of path information may further include, if the second mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information. Further in the second mode, each encoded item of path information may further include, if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between a previous path length and a current path length. Therein, the previous path length may be a path length specified by the preceding item of path information and the current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Further in the second mode, each encoded item of path information may otherwise include the indication of the comer voxel and the indication of the path length.
[0023] In some embodiments, each encoded item of path information may include an indication of whether a first mode or a second mode is used. Each encoded item of path information may further include an indication of the first location. Each encoded item of path information may further include an indication of the second location. Each encoded item of path information may further include an indication of whether an acoustic path exists for the first location and the second location. Each encoded item of path information may further include, if the first mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by a preceding item of path information. Then, if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information may further include an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between a previous path
length and a current path length, together with the indication of the path length or the indication of the difference between the previous path length and the current path length. Here, the previous path length may be a path length specified by the preceding item of path information. The current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Further in the first mode, the encoded item of path information may otherwise include the indication of the corner voxel and the indication of the path length. Each encoded item of path information may further include, if the second mode is used, the indication of whether the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information. Further in the second mode, each encoded item of path information may further include, if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between a previous path length and a current path length. Therein, the previous path length may be a path length specified by the preceding item of path information and the current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Further in the second mode, each encoded item of path information may otherwise include the indication of the corner voxel and the indication of the path length.
[0024] In some embodiments, the method may further include outputting the encoded items of path information to a bitstream.
[0025] Another aspect of the present disclosure relates to a method of processing audio scene information, for example decoding audio scene information, in particular, acoustic path information. The method may include receiving a bitstream comprising a sequence of encoded items of path information for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to a voxel-based audio scene representation. Each encoded item of path information may correspond to a respective item of path information that specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a comer voxel on the acoustic path for which the acoustic path changes direction. The comer voxel may be a voxel for which a direct line to the second voxel is not occluded. The method may further include successively decoding encoded items of path information to generate corresponding items of path information. Therein, for a
current item of encoded path information, generating the corresponding item of path information may include determining whether the current encoded item of path information includes an indication that the corner voxel specified by the corresponding item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding item of encoded path information. Then, if the corner voxel specified by the corresponding item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, the method may include setting the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information as the comer voxel for the item of path information corresponding to the current encoded item of path information. On the other hand, if the corner voxel specified by the corresponding item of path information is different from the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, the method may include extracting an indication of the comer voxel from the current encoded item of path information.
[0026] In some embodiments, generating the corresponding item of path information may further include extracting an indication of the path length from the current encoded item of path information.
[0027] In some embodiments, if the corner voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information may further include extracting an indication of a difference between a previous path length and a current path length. Therein, the previous path length may be a path length specified by the item of path information corresponding to the preceding encoded item of path information. The current path length may be a path length specified by the item of path information corresponding to the current encoded item of path information.
[0028] In some embodiments, if the corner voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information may further include extracting an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference
between a previous path length and a current path length. Therein the previous path length may be a path length specified by the item of path information corresponding to the preceding encoded item of path information. The current path length may be a path length specified by the item of path information corresponding to the current encoded item of path information. Generating the corresponding item of path information may further include extracting the indication of the path length or the indication of the difference between the previous path length and the current path length.
[0029] In some embodiments, the one or more second locations may relate to locations obtained by, for each first location, traversing the voxel grid according to a predetermined pattern to define a sequence of second locations. Further, for the respective first location, the encoded items of path information may be successively decoded in accordance with the sequence of second locations.
[0030] In some embodiments, the predetermined pattern may traverse the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
[0031] In some embodiments, the difference may be encoded by 2 bits. Alternatively, the difference may take one of four predetermined values.
[0032] In some embodiments, each encoded item of path information may include an indication of whether a first mode or a second mode is used. Each encoded item of path information may further include an indication of the first location. Each encoded item of path information may further include an indication of the second location. Each encoded item of path information may yet further include an indication of whether an acoustic path exists for the first location and the second location. Each encoded item of path information may further include, if the first mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information. Each encoded item of path information may further include, if the first mode is used, and if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of the path length. Otherwise, each encoded item of path information may further include the
indication of the comer voxel and the indication of the path length. Each encoded item of path information may further include, if the second mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information. Each encoded item of path information may further include, if the second mode is used, and if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between a previous path length and a current path length. Therein, the previous path length may be a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Otherwise, each encoded item of path information may further include the indication of the comer voxel and the indication of the path length.
[0033] In some embodiments, each encoded item of path information may include an indication of whether a first mode or a second mode is used. Each encoded item of path information may further include an indication of the first location. Each encoded item of path information may further include an indication of the second location. Each encoded item of path information may yet further include an indication of whether an acoustic path exists for the first location and the second location. Each encoded item of path information may further include, if the first mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information. Each encoded item of path information may further include, if the first mode is used, and if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of whether the encoded item of path information includes the indication of the path length or an indication of a difference between a previous path length and a current path length, together with the indication of the path length or the indication of the difference between the previous path
length and the current path length. Therein, the previous path length may be a path length specified by the item of path information corresponding to the preceding encoded item of path information. The current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Otherwise, each encoded item of path information may further include the indication of the corner voxel and the indication of the path length. Each encoded item of path information may further include, if the second mode is used, the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information. Each encoded item of path information may further include, if the second mode is used, and if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between a previous path length and a current path length. Therein, the previous path length may be a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length may be a path length specified by the item of path information corresponding to the encoded item of path information. Otherwise, each encoded item of path information may further include the indication of the comer voxel and the indication of the path length.
[0034] According to another aspect, an apparatus for processing audio scene information is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform all steps of the methods according to preceding aspects and their embodiments.
[0035] According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., processor).
[0036] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device
(e.g., processor) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.
[0037] It should be noted that the methods and systems including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and systems disclosed in this document. Furthermore, all aspects of the methods and systems outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0038] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa.
Brief Description of the Drawings
[0039] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein
[0040] Fig. 1 schematically illustrates an example of a processing chain for processing audio scene information for audio rendering;
[0041] Fig. 2 schematically illustrates an example of a diffraction path for a source location and a listener location in a voxel-based three-dimensional audio scene;
[0042] Fig. 3 is a flowchart illustrating an example of a method of processing audio scene information for audio rendering according to embodiments of the disclosure;
[0043] Fig. 4 is a flowchart illustrating an example of an implementation detail of the method of Fig. 3 according to embodiments of the disclosure;
[0044] Fig. 5 to Fig. 7 schematically illustrate examples of processing chains for processing audio scene information for audio rendering according to embodiments of the disclosure;
[0045] Fig. 8 is a diagram illustrating complexity measures as functions of time for different operating modes/implementations of processing audio scene information for audio rendering according to embodiments of the disclosure;
[0046] Fig. 9 schematically illustrates an example of a possible use case for techniques according to embodiments of the disclosure;
[0047] Figs. 10A-10C schematically illustrate examples of part of a voxel-based audio scene according to embodiments of the disclosure;
[0048] Fig. 11 schematically illustrates an example of a voxel -based audio scene to which embodiments of the disclosure may be applied;
[0049] Fig. 12 is a flowchart illustrating an example of a method of processing audio scene information for encoding acoustic path information according to embodiments of the disclosure;
[0050] Fig. 13 is a flowchart illustrating an example of a method of processing audio scene information for decoding acoustic path information according to embodiments of the disclosure;
[0051] Fig. 14 is a diagram illustrating complexity measures as functions of time for different operating modes/implementations of processing audio scene information for audio rendering according to embodiments of the disclosure; and
[0052] Fig. 15 is a block diagram schematically illustrating an example of an apparatus implementing methods according to embodiments of the disclosure.
Detailed Description
[0053] In the following, example embodiments of the disclosure will be described with reference to the appended figures. Identical elements in the figures may be indicated by identical reference numbers, and repeated description thereof may be omitted.
Voxel-based Audio Scene Representations
[0054] First, an overview over voxel-related concepts for representation of audio scenes will be given.
What is a voxel for audio rendering?
[0055] A voxel is understood as a space volume with acoustic properties or audio rendering instructions assigned to it.
What is a voxel size for audio rendering?
[0056] The voxel size may be an encoder configuration parameter. It may be (manually or automatically) selected according to a scene geometry level of details (e.g., in the range of 10 cm - 1 m).
How large audio scenes can be handled?
[0057] Large audio scenes do not necessarily result in a large number of voxels and high rendering complexity. For example, a large audio scene can be represented as
• a set of independent sub-scenes (and method for “teleport” between these representations without a Tenderer “re-start”)
• a set of scenes updates (based on the user position)
How can discontinuity issues caused by voxel granularity be handled?
[0058] Any strong discontinuities in sound levels (and jumps of diffracted signal direction) can be avoided by application of interpolation (e.g., in time and space).
How to represent voxel-based audio scenes?
[0059] Any voxel-based representation of an audio scene may contain an indication of voxels that are not transmission voxels (e.g., that are occluder voxels), i.e., voxels in which sound cannot propagate or cannot freely propagate - a representation of occluding geometries. This indication may relate to an indication of coordinates (e.g., center coordinates, comer coordinates, etc.) of the respective voxels. The coordinates of these voxels may be represented by grid indices, for example. Additionally, the voxel-based representation may include indications of material properties of the voxels that are not transmission voxels, such as absorption coefficients, reflection coefficients, etc.. In addition to the occluder voxels, the voxelbased representation may also indicate transmission voxels (e.g., air voxels), i.e., voxels in which sound can propagate - a representation of sound propagation media. Accordingly, some implementations of voxel-based representations of audio scenes may include, for each voxel in a
predefined section of space (e.g., within boundaries enclosing the audio scene), and indication of a respective material property.
Techniques for Processing Audio Scene Information
[0060] Fig. 1 schematically illustrates a processing chain 100 that can be used for processing audio scene information for audio rendering. Specifically, the processing chain 100 can be used for converting voxel related data into parameters and signals needed for auralization (or audio rendering in general). The processing chain 100 may be implemented in software, hardware, or combinations thereof. For example, the processing chain 100 may be implemented by a renderer/decoder coupled to AR/VR/MR/XR equipment, such as AR/VR/MR/XR goggles. Specific implementations may include game consoles, set-top-boxes, personal computers, etc..
[0061] The processing chain receives an audio scene description 20 from a bitstream (or storage/memory) 10. The audio scene description 20 may comprise a representation of a three- dimensional audio scene and information on a source location of a sound source within the audio scene. The representation of the three-dimensional audio scene may be voxel-based, for example.
[0062] The processing chain 100 further receives an indication of a user position (listener location) 30 of a user (listener) within the audio scene. The audio scene description 20 and the user position 30 are provided to a diffraction direction calculation block (diffraction calculation block) 40 for determining (e.g., calculating) diffraction information. The diffraction information may relate to an acoustic diffraction path within the audio scene between the source location and the listener location. The diffraction information is then provided to a diffraction modeling tool 50 for applying diffraction modeling and optionally occlusion modeling, based on the diffraction information. The occlusion modeling calculates attenuation gains for the direct line between the listener and an audio source. The diffraction modeling tool 50 may output auralized audio data (3DoF auralizer data) that includes, for example, a location of an object to be rendered, an orientation, and frequency dependent gains. The diffraction modeling tool output may be further processed by other rendering stages such as Doppler, Directivity, Distance Attenuation, etc.. In general, the diffraction modeling tool 50 may be said to output diffraction information, as detailed below. The auralized audio data may then be used for audio replay, for example.
[0063] In summary, a processing chain as shown in Fig. 1 may be used to convert voxel related data into the parameters for parameters and signals for auralization. The diffraction
direction calculation block 40 and the diffraction modeling tool 50 may be seen as non-limiting examples of rendering tools. In general, the rendering tools may generate 3DoF auralizer data.
[0064] As noted above, the scene description may include a voxel matrix and associated coefficients (e.g., reflection coefficients, occlusion coefficients, absorption coefficients, transmission coefficients etc.). These coefficients may be indicative of a material or material property of the respective voxel. The rendering tools may include, for example, occlusion and diffraction modelling tools. The 3DoF auralizer data may include, for example, object position, orientation and frequency dependent gains.
[0065] As noted above, the voxel-based representation of the three-dimensional audio scene defines psycho-acoustically relevant geometric elements and sound propagation media. In some implementations, the scene description may use the following param eters/interfaces (e.g., the following agreed upon data format, or agreed upon point of data exchange) to provide the information to rendering tools:
Scene size: in absolute units (e.g., meters) in number of voxels and/or voxel size
Scene anchors: in terms of coordinate anchors (to map absolute coordinates to voxel indices) in terms of scene anchors (to map sub-scene to sub-set of voxels)
Scene content data: reference to material properties that approximates acoustic effects caused by occluders (sound obstacles) located in the corresponding volume (e.g., coefficients for transmission, reflection, etc.) reference to sound propagation media properties that approximates an acoustic effect caused by media located in the corresponding volume (e.g., speed of sound, energy absorption, distance attenuation curve, etc.) rendering control parameter describing intended occlusion modelling effects
E.g., “global” or “local” occluder type that determines the length (and shape) of occlusion effect shadow behind this voxel rendering control parameter describing intended sound diffraction modelling effects
E.g., voxel type that controls/causes the change of sound direction (i.e., path of the diffracted sound cannot penetrate this volume) content control parameters describing audio signal relevance and scene authoring E.g., audio signal IDs and/or signal gains the determines which signal is perceptually relevant (rendered) in the corresponding volume rendering control parameter describing intended reverberation modelling effects
E.g., voxel type that controls reverberation settings (e.g., RT60, DDR, RIR, etc.)
Scene content updates: referenced to update triggering events
All data can be audio object dependent (to support content creator intent in flexible audio scene authoring).
[0066] The 3DoF auralizer data may include the following information: parameters and associated signals for the set of audio objects (and HO A) o parameters include the metadata output of the rendering tools (i.e., position, orientation and gains simulating effects of occlusion, diffraction, early reflections, parameters for reverberation coefficients, IR, etc.) o associated signals represent the audio output of the rendering tools (i.e., downmixed or replicated audio signals) scene state identifier (i.e., metadata allowing to map the scene description and user input to the 3DoF auralizer data)
[0067] Fig. 2 illustrates an example of a possible scene state and a diffraction path for this scene state. It is understood that the scene state relates to or comprises the listener location 210 and the audio scene description (including the representation of the three-dimensional audio scene and the source location 220).
[0068] The example of Fig. 2 relates to a voxel based representation of the three- dimensional audio scene. This voxel based representation indicates “air” voxels or empty voxels (i.e., voxels in which sound can propagate, or transmission voxels) 230 and occluder voxels 240 (i.e., voxels in which sound cannot propagate or cannot freely propagate). Accordingly, occluder
voxels may be understood to relate to voxels filled with a material other than air, and that can reflect, block, or otherwise alter sound propagation. For the occluder voxels 240, the representation may further indicate respective transmission, reflections coefficients and potentially absorption coefficients relating to material properties of these voxels. These coefficients may be linked to ID’s or indices of their respective voxels in the voxel -based representation. In general, the voxel based representation may define psycho-acoustically relevant geometric elements and sound propagation media in the audio scene.
[0069] A listener location 210 is indicated by a parameter Lvox and a source location 220 is indicated by another parameter Svox.
[0070] A diffraction path (or in general, acoustic path) between the source location 220 and the listener location 210 may be determined using a pathfinding algorithm that takes the listener location 210, the source location 220, and the representation of the three-dimensional audio scene (or a two-dimensional representation, e.g., 2D projection or 2D matrix, derived therefrom, e.g., in the form of a voxel grid) as inputs. For example, an algorithm for determining the diffraction information may take the listener location 210, the source location 220, and the representation of the three-dimensional audio scene as inputs and may output a location of a comer voxel (e.g., diffraction comer) 250, indicated by Cvox and the variable nn representing the length of the diffraction path. For example, the diffraction information (or acoustic path information in general) may be determined based on:
[Cvox, rin] = DiffractionDirectionCalculation , o , Svox, VoxDataDiffractionMap) where DiffractionDirectionCalculation indicates the algorithm for determining the diffraction information (“pathfinding algorithm”) and VoxDataDiffractionMap indicates the voxel -based representation of the three-dimensional audio scene or a processed version thereof (e.g., 2D projection or 2D matrix derived therefrom). Cvox is understood to indicate the coordinates of the diffraction corner (e.g., coordinates, voxel/grid coordinates, or voxel/grid indices of the respective voxel including the diffraction corner). The diffraction comer may also be referred to as corner voxel.
[0071] Here, DiffractionDirectionCalculation may involve any viable pathfinding algorithm, such as the Fast traversal algorithm for ray tracing (cf. Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87.) and
the JPS algorithm (cf. Harabor, D.D. and A. Grastien, Online Graph Pruning for Pathfinding On Grid Maps. Proceedings of the Twenty -Fifth AAAI Conference on Artificial Intelligence, 2011.), for example. Further, one may directly apply a 3D path search algorithms to obtain the shortest path between the source location 220 and the listener location 210 using the voxel -based scene representation. Alternatively, one may apply a 2D path search algorithms for this task, using an appropriate 2D projection plane of the 3D voxel-based scene representation. For the indoor (e.g., multiroom room) sound simulation the corresponding 2D projection plane may be similar to a floor plan that describes a “sound propagation path topology”. For outdoor sound simulation scenarios, it may be of interest to consider a second (e.g., vertical) 2D projection plane to account for the diffraction paths going over sound obstacle(s) or occluding structure(s). The path finding approach remains the same for all projection planes, but its application delivers an additional path that can be used for the diffraction modelling.
[0072] The pathfinding algorithm is assumed to output an acoustic path (e.g., diffraction path) that connects the source location 220 to the listener location and that consist of a plurality of indications of voxels (e.g., voxel indices or voxel coordinates) at which the acoustic path may be said to change direction. For visualization, the acoustic path might in some cases be seen as relating to a plurality of straight path segments (line segments) that are sequentially linked end- to-end. Each transition from one path segment to another path segment relates to a change of direction of the diffraction path.
[0073] According to the algorithm for determining the diffraction information, the diffraction corner (or corner voxel) Cvox may be determined as a voxel that lies on or on the proximity of the diffraction path and is adjacent to an occluder voxel (in a set of voxels representing corner voxels on the diffraction map, Cset) of the diffraction map (indicated by the voxel-based representation). For example, the diffraction comer Cvox may be selected from a set of voxels (Pset) forming the diffraction path as a voxel that is close to a ‘visible’ (from the listener position Lc) occluder voxel (belonging to Cset) causing the path (Pset) to change direction. If there are more than one such corners, the one furthest away from the listener location along the diffraction path (Pset) is selected.
[0074] In general, the diffraction path algorithm may be said to determine diffraction information relating to the acoustic path (e.g., acoustic diffraction path) within the audio scene
between the source location and the listener location. In general, the diffraction path algorithm may be said to determine acoustic path information.
[0075] This diffraction information (or acoustic path information in general) may be sufficient information for the Tenderer to recover/determine a virtual source location of a virtual audio source that encapsulates effects of acoustic diffraction effects. This is the case for the coordinates of the diffraction comer Cvox and the diffraction path length nn. For example, the virtual source location may be recovered by calculating the direction (e.g., azimuth, or azimuth and elevation) of the diffraction corner when seen from the listener location. Using this direction and taking the path length nn of the diffraction path as the virtual source distance to the listener location, the virtual source location can be determined.
[0076] It is noted that the diffraction information can be represented in different ways. One option, as noted above, is diffraction information including/ storing the path length nn and the coordinates (e.g., grid coordinates, etc.) of the diffraction corner Cvox.
Based on the above, the following data elements may be defined:
An example of the scene state Ni may be represented by
Ni = {Lvox, Svox, V oxDataDiffracti onMap } , i.e., may relate to or comprise the listener location Lvox, the source location Svox and the voxelbased representation (e.g., VoxDataDiffractionMap) of the audio scene.
A scene state identifier for a scene state Ni may be defined as
SceneStateldentifier = HASH(Ni), where HASH is a hash function that generates a hash value for scene state Ni, e.g., that maps scene states to fixed-size values. In general, the scene state identifier may be said to be indicative of a certain scene state or to identify a certain scene state.
Further, an example of the diffraction information N2 may be represented by
N2 = {Cvox, rin{, where nn is the path length of the diffraction path and Cvox indicates the location (e.g., voxel location) of the diffraction corner, as described above.
A quantized version of the diffraction information N2 may be indicated by N3, where
N3 = voxSceneDiffractionPreComputedPathData(Ni),
N3 (~= N2) = DiffractionDirectionCalculation^i), where voxSceneDiffractionPreComputedPathData() is a bitstream syntax that parses the bitstream and retrieves the precomputed (stored and quantized) diffraction information (e.g., generated by the processing chain 500 of Fig. 6) and DiffractionDirectionCalculation() denotes a function that performs an online calculation of the diffraction information which may be implemented for example in diffraction direction calculation block 40 in Fig. 5 and Fig. 7.
The diffraction information, for example Cvox and nn, may also be seen as relating to 3DOF auralizer data, because user the position voxel coordinates Lvox are fixed.
[0077] An example of syntax element voxSceneDiffractionPreComputedPathData() according to the MPEG-I standard is given by Table 1. This voxel payload data structure may have the following elements: numberOfV oxDiffractionPathData
This element represents the number of pre-computed diffraction path data sets. voxDiffractionPathStartVoxelPacked
This element represents the packed form of the variable voxDiffractionPathStartVoxel indicating the voxel indices of the path start voxel of the pre-computed diffraction path (e.g., Svox or Lvox in Fig. 2). voxDiffractionPathStartVoxel may be a 2D position on the diffraction map indicating the start position of the diffraction path, for example. voxDiffractionPathEndVoxelPacked
This element represents the packed form of the variable voxDiffractionPathEndVoxel indicating the voxel indices of the path end voxel of the pre-computed diffraction path (e g-, L vox or Svox in Fig- 2). voxDiffractionPathEndVoxel
may be a 2D position on the diffraction map indicating the end position of the diffraction path, for example. voxDiffractionPathDataExistFlag
This element indicates whether the diffraction path exists or not. voxDiffractionSourceDirectionPacked
This element represents the packed form of the variable voxDiffractionSourceDirection indicating the voxel indices of the voxel for determining diffracted source azimuth value. This may correspond to the comer voxel C VOX, for example. voxDiffractionPathLength
This element represents the diffraction path length on the diffraction map 2D matrix. This may correspond to the path length nn, for example.
Note: voxSceneDimensions
The number of voxels per each scene dimension. escaped Value()
This element implements a general method to transmit an integer value using a varying number of bits. It features a two level escape mechanism which allows to extend the representable range of values by successive transmission of additional bits. Syntax of escaped Value() shall be as defined in ISO/IEC 23003-3.
Table 1 — Syntax of voxSceneDiffractionPreComputedPathDataQ
The variables retrieved from the voxSceneDiffractionPreComputedPathDataQ (e g., as in
Fig. 6) are further processed to output the diffraction information N3.
[0078] A technical benefit and effect according to techniques of the present disclosure is that the scene state identifier or other information derived from the scene state may be used to avoid application of the diffraction modeling tools or rendering tools if the corresponding processing was already done for this scene state and the diffraction information or 3DoF auralizer data are available. In this scenario the Tenderer can access the diffraction information/3DoF auralizer data (for a known scene state) without application of the rendering tools by: re-using the data calculated before (precomputed), or applying the data calculated by another Tenderer.
[0079] A technical benefit an effect is thus that techniques according to the present disclosure relate to lossless functionality aiming at the low complexity mode (complexity vs bitrate).
[0080] To fully implement such scheme, the present disclosure proposes to provide the processing chain for processing audio scene information for audio rendering (e.g., in a decoder/renderer) with an interface for providing/outputting the diffraction information for later use or use by a different decoder/renderer. This interface is understood to be a data interface for outputting data in a predefined format, to allow for consistent re-use especially by other
decoder s/renderers. The interface may be implemented and/or utilized in any combination of software and hardware. Specifically, this may relate to providing/outputting a data element that comprises the diffraction information and information on the scene state, such as the scene state identifier, for example. The data element may have a predefined format, for example with predefined data fields. Using this interface, the processing chain can provide the computed diffraction information or 3DoF auralizer data together with the scene state identifier to other decoder s/renderers and/or store it for later re-use.
[0081] In some implementations, the above may relate to providing/outputting (or on the receiving side, receiving) precomputed acoustic path information for one or more (e.g., all feasible) combinations of source locations and listener locations in a voxel-based scene representation (e.g., two-dimensional voxel grid).
[0082] Example 1: if the decoder/renderer has obtained acoustic path information / diffraction information (e.g., a diffraction path) for a given user position (listener location), the decoder/renderer can re-use it until the user leaves the corresponding voxel volume (or the scene description is updated).
[0083] Example 2: If the computed acoustic path information / diffraction information corresponds to a scene state unknown to the other decoders, they may re-use the acoustic path information / diffraction information and avoid running their own diffraction modeling tools or rendering tools.
[0084] Exchange and sharing of acoustic path information / diffraction information among different decoders can be done using a database, which can be included into the bitstream (to be accessed, for example, via application request).
[0085] Fig. 3 is a flowchart showing an example of a method 300 of processing audio scene information for audio rendering in accordance with embodiments of the present disclosure. Method 300 may be implemented in software, hardware, or combinations thereof. For example, the processing chain 100 may be implemented by a renderer/decoder coupled to AR/VR/MR/XR equipment, such as AR/VR/MR/XR goggles. Specific implementations may include game consoles, set-top-boxes, personal computers, etc..
[0086] Method 300 comprises steps S310 through S350 that may be performed, for example, by a decoder/renderer. These steps may be performed, for example, whenever the scene state changes. With the scene state understood as relating to or comprising the listener location 210 and the audio scene description (including the representation of the three-dimensional audio scene and the source location 220), for example implemented by scene state Ni above, a change of the scene state could relate to one or more of a change of the listener location 210, a change of the source location, and a change of the (representation of the) three-dimensional audio scene. Alternatively, steps S310 through S350 may be performed for each of a plurality of processing cycles of a decoder/renderer. If the audio scene description is unchanged, step S310 may however be omitted. It is also to be understood that steps S310 through S350 do not need to be performed in the order shown in Fig. 3.
[0087] At step S310, an audio scene description is received. The audio scene description comprises a representation of a three-dimensional audio scene and information on a source location of a sound source within the audio scene. For example, the audio scene description may comprise elements Svox and VoxDataDiffractionMap defined above, for example.
[0088] At step S320, information of a listener location of a listener within the audio scene is received. The listener location may correspond to element Lvox defined above, for example.
[0089] At step S330, diffraction information relating to an acoustic diffraction path within the audio scene between the source location and the listener location is obtained. The obtained diffraction information may be indicative of a virtual source location of a virtual sound source. For example, the virtual source location may have the same direction (e.g., azimuth, or azimuth and elevation), when seen from the listener location, as the diffraction comer Cvox. The virtual source distance may correspond to the length nn of the diffraction path. Accordingly, the diffraction information may comprise indications of Cvox and nn defined above.
[0090] At step S340, audio rendering is performed for the sound source based on the diffraction information. This may include, for example, diffraction modeling.
[0091] To this end, a virtual source location of a virtual source may be determined based on the diffraction information. The virtual source may be an audio source that encapsulates effects of acoustic diffraction between the source location and the listener location in the three-
dimensional audio scene. For example, the virtual source location may be determined based on Cvox and nn by
• determining the direction (e.g., azimuth, or azimuth and elevation) of the diffraction comer Cvox when seen from the listener location;
• using the determined direction as the virtual source direction of the virtual source when seen from the listener location; and
• using the diffraction path length nn as o the virtual source distance of the virtual source from the listener location, or o the virtual source gain compensation derived from the diffraction and direct path lengths.
[0092] Audio rendering may then include rendering the virtual sound source at the virtual source location, for example.
[0093] At step S350, a representation of the diffraction information is output. For example, outputting the representation of the diffraction information may comprise outputting a data element comprising the diffraction information and information on the scene state. The scene state may comprise the audio scene description (e.g., Svox and VoxDataDiffractionMap) and the listener location (e.g., Lvox).
[0094] The output may be provided to a look up table (LUT). The LUT includes, as its entries, different items of diffraction information indexed with information on respective scene states (e.g., indexed with respective scene state identifiers). This LUT thus may be said to include the diffraction information and information on the scene state. The LUT can be stored and/or provided to be later retrieved, for example by other decoders, from a bitstream or from a shared storage (e.g., cloud or server based), for example by application request. A hash value of the scene state or the scene state identifier can be used to retrieve the actually desired entry from the LUT.
[0095] Further, the representation of the diffraction information may be output to a bitstream (e.g., outgoing bitstream) and/or to a storage (e.g., a memory, cache, file, etc.). The storage may be local or it may be shared (e.g., cloud based). In general, the representation of the diffraction information may be output to a suitable medium for storing digital information or
computer related information. The output may at least partially be directed to an external or shared data source or data repository.
[0096] In some implementations, the representation of the diffraction information may be output as part of a voxSceneDiffractionPreComputedPathData() syntax element according to ISO/IEC 23090-4 (Coded representation of immersive media — Part 4: MPEG-I immersive audio, https://www.iso.org/standard/8471 l.html), or according to any future standard deriving therefrom.
For example, the voxSceneDiffractionMap() syntax element may be given by Table 2.
Table 2 — Syntax of voxSceneDiffractionMapQ
voxSceneDiffractionMap() provides a compact representation of a 2D diffraction map (VoxDataDiffractionMap). This 2D representation is similar to the 3D representation used for the voxel-based 3D audio scene.
[0097] A MapElement is defined by 2 points (x,y-indices) on the diffraction map and a corresponding value. The two points span a rectangle and all covered grid cells are assigned the value voxDiffractionMapValue.
[0098] The bitstream element numberOfVoxDiffractionMapElements signifies the number of MapElements.
[0099] The bitstream element voxDiffractionMapValue_signifies the binary value controlling the path finding algorithm. It is useful because the value indicates whether a path can go through the grid cell or not. This value is defined for all entries on the diffraction map.
[0100] The bitstream element voxDiffractionMapPosPackedS signifies a packed representation of 2 indices of the start grid cell of a MapElement. It may be an array that illustrates a collection of all start grid cells.
[0101] The bitstream element voxDiffractionMapPosPackedE signifies a packed representation of 2 indices of the end grid cell of a MapElement. It may be an array that illustrates a collection of all end grid cells.
[0102] Both the voxDiffractionMapPosPackedS and voxDiffractionMapPosPackedE are useful because they allow for a compact representation of the data where a single voxDiffractionMapValue is used for all grid cells between these two variables.
[0103] Fig. 4 is a flowchart illustrating an example of a method 400 including steps that may be performed for implementing steps of method 300. Method 400 comprises steps S410 through S460. Of these, steps S410 through S450 may implement step 330 of method 300. Further, step S460 may correspond to step S350.
[0104] At step S410, a current scene state is determined based on the audio scene description and the listener location.
[0105] At step S420, it is determined whether the current scene state corresponds to a known scene state for which precomputed diffraction information is available (e.g., can be retrieved). The precomputed diffraction information may be retrieved from a bitstream (incoming bitstream) or storage (including, in particular, an external or shared storage), for example. Determining whether the current scene state corresponds to a known scene state may comprise determining a hash value based on the current scene state. It may further comprise comparing the hash value of the current scene state against hash values of known (e.g., previously encountered) scene states.
[0106] If it is determined that the current scene state corresponds to a known scene state (YES at step S430), the method proceeds to step S440.
[0107] At step S440, the diffraction information is determined by extracting the precomputed diffraction information for the known scene state from the bitstream or storage. The storage may relate to local storage (e.g., memory, cache, file, etc.) or to a shared storage (e.g., cloud storage, server storage).
[0108] Extracting the precomputed diffraction information for the known scene state may include receiving a look up table or an entry of a look up table from the bitstream (incoming bitstream) or storage. The look up table may be seen as a representation of the diffraction information. It may comprise a plurality of items of precomputed diffraction information, each associated with a respective known scene state. The precomputed diffraction information and the associated known scene state may correspond to the aforementioned data elements. The known scene state may comprise or be indicative of a known audio scene description and a known listener location.
[0109] Selecting the relevant entry of a received look up table, or selecting the relevant entry to be received (if not all of the look up table, but only an entry thereof is received) may involve using hash values, as described above.
[0110] On the other hand, if it is determined that the current scene state does not correspond to a known scene state (NO at step S430), the method proceeds to step S450.
[OHl] At step S450, the diffraction information is determined using a pathfinding algorithm, based on the source location, the listener location, and the representation of the three- dimensional audio scene. This may be done in accordance with the procedure described above with reference to Fig. 2.
[0112] At step S460, the diffraction information obtained via step S440 or step S450 is output. This step may correspond to step S350 described above.
[0113] In summary, the proposed method may comprise (inter alia) the following:
• Check if the pre-computed diffraction path information (i.e., pre-computed diffraction information) can be retrieved from cache and re-applied for the current scene state and listener position.
• Check if the pre-computed diffraction path information (i.e., precomputed diffraction information) for the current scene state can be obtained from memory cache or bitstream and re-applied.
[0114] Therein, the scene state is defined via the input parameters Lvox, Svox, VoxDataDiffractionMap for the function DiffractionDirectionCalculation() comprising the pathfinding algorithm, voxel Cvox selection and diffraction path length estimation steps.
The diffraction path information (e.g., diffraction information) is defined via the output parameters Cvox, rin . This diffraction path information, if it is available, can be directly obtained from the bitstream syntax voxSceneDiffractionPreComputedPathData() for the corresponding scene state to avoid the function DiffractionDirectionCalculation() call.
[0115] When the diffraction path information Cvox, rin are obtained for the current scene state Lvox, Svox, VoxDataDiffractionMap, this information can be cached in memory (and provided outside the Tenderer) for later re-use by the Tenderer (or other Tenderer instances).
[0116] In other words, “Diffracted path finding" according to the disclosure (e.g., embodied by method 300 and/or method 400) may involve the following processing:
Check if the pre-computed diffraction path information for the current scene state can be obtained from memory cache or bitstream and re-applied.
The scene state is defined via the input parameters Lvox, Svox, VoxDataDiffractionMap for the function DiffractionDirectionCalculation() comprising the path-finding algorithm, voxel Cvox selection and diffraction path length estimation steps.
[Cvox, rin] = DiffractionDirectionCalculation(fr o , Svox, VoxDataDiffractionMap)
The diffraction path information is defined via the output parameters Cvox, rin. This diffraction path information, if it is available, can be directly obtained from the bitstream syntax voxSceneDiffractionPreComputedPathData() for the corresponding scene state to avoid the function DiffractionDirectionCalculation() call.
- When the diffraction path information Cvox, rin are obtained for the current scene state Lvox, Svox, VoxDataDiffractionMap, this information can be cached in memory (and provided outside the Tenderer) for later re-use by the Tenderer (or other Tenderer instances).
[0117] In the above, a bitstream syntax definition may be written in a function() style in MPEG standard document. It defines how to read/parse data (bitstream elements) from the bitstream. In this case, it is used to obtain necessary variables/information to recover the diffraction path information.
[0118] Fig. 5, Fig. 6, and Fig. 7 show examples of a processing chain 500 in accordance with the above that can be used for processing audio scene information for audio rendering.
Specifically, the processing chain 500 can be used for converting voxel related data into parameters and signals needed for auralization.
[0119] Fig. 5 relates to the case that the current scene state is an unknown scene state. Different from the processing chain 100 of Fig. 1, the bitstream/memory 510 additionally includes data elements comprising diffraction information and associated scene states, for example in the form of look up tables, as described above.
[0120] Same as the processing chain 100, the processing chain 500 receives an audio scene description 20 from the bitstream (or storage/memory) 510. The processing chain 500 further receives an indication of a user position (listener location) 30 of a user (listener) within the audio scene.
[0121] The diffraction direction calculation block (diffraction calculation block) 40 for determining (e.g., calculating) diffraction information and the diffraction modeling tool 50 for applying diffraction modeling and optionally occlusion modeling, based on the diffraction information, may be the same as for the processing chain 100.
[0122] However, different from the processing chain 100 of Fig. 1, the audio scene description 20 and the listener location 30 are used to determine a scene state 515 or scene state identifier. This scene state 515 (e.g., scene state Ni defined above) or scene state identifier (e.g., HASH(Ni)) is provided/input to a scene state analyzing block 520 that determines whether the current scene state 515 corresponds to a known scene state 530 (yes at block 535) or not (no at block 535). In the present example, the current scene state 515 does not correspond to a known scene state (i.e., the current scene state 515 is an unknown scene state). Thus, the audio scene description 20 and the listener location 30 are input to the diffraction direction calculation block 40, in the same manner as for the processing chain 100, for generating the diffraction information 550. Subsequently, the diffraction information 550 is used for rendering/diffraction
modeling, as in the case of processing chain 100. Additionally however, the diffraction information 550 (e.g., diffraction information N2 or quantized version N3 thereof as defined above) is output via an interface, for later re-use by the Tenderer or other (external) rendering instances. Specifically, the diffraction information 550 may be output together with the corresponding scene state via an output interface 555 to the bitstream (or memory/storage) 510.
[0123] Fig. 6 relates to the case of that the current scene state 515 is a known scene state.
[0124] Again, the current scene state 515 is provided/input to the scene state analyzing block 520 to determine whether the current scene state 515 corresponds to a known scene state 530 or not. In the present example, the current scene state 515 corresponds to a known scene state. Thus, instead of inputting the audio scene description 20 and the listener location 30 to the diffraction direction calculation block 40 for calculating/generating the diffraction information, the diffraction information is extracted/received from the bitstream (or storage/memory) 510, as described above (e.g., via step S450 of method 400). Still, even though the diffraction information is not locally calculated, it may be output to the bitstream (or memory/storage) 510, as in the case of Fig. 5. The reason is that if the diffraction information is obtained from one source (e.g., among the bitstream or storage), it can be made available to the respective other source(s) in this way.
[0125] Fig. 7 shows the full processing chain 500, including data paths for both a known and an unknown scene state 515.
[0126] Fig. 8 is a diagram illustrating complexity measures for different implementations of processing audio scene information or audio rendering as functions of time, assuming a simple maze as the audio scene. It is further assumed that the user randomly moves through the maze, thus revisiting previously visited locations. Graph 810 relates to the case that no pre-computed diffraction information whatsoever is available (e.g., no diffraction information provided with the bitstream, memory/cache disabled). In this case, the computational load on the Tenderer is substantially constant and comparatively high. Graph 820 relates to the case that pre-computed diffraction information is locally available (e.g., no diffraction information provided with the bitstream, local memory/cache enabled). In this case, the processing load on the Tenderer decays over time, since more and more items of diffraction information are locally accumulated. In other words, more and more scene states that are encountered will relate to (locally) known scene
states. Graph 830 finally relates to the case that pre-computed diffraction information is externally provided (e.g., full diffraction information provided with the bitstream). In this case, the computation load on the Tenderer is constantly low, as a significant portion of scene states relates to known scene states and the diffraction information can be externally retrieved (e.g., from the bitstream or by request from an external/ shared storage), without local calculation.
[0127] Fig. 9 schematically illustrates an example of a possible use case for techniques according to embodiments of the disclosure. Shown are two listeners (users) A and B at different locations within an audio scene (e.g., a house with different areas and levels). Users A and B may be users that individually or jointly explore a VR environment including the audio scene, for example as part of a game, virtual tour, etc.. The users exploring a common VR environment may be running a social VR, for example. Having different listener locations within the audio scene, users A and B will produce different rendering results and different diffraction information. The present disclosure foresees that each user (or their respective device/ decoder/ render er) makes their calculated diffraction information available to other users. Once user B enters an area of the audio scene in which user A had been present previously, they may benefit from user A’s precomputed diffraction information, and vice versa. For example, user A’s diffraction information may be made available to user B via a LUT that indexes different items of diffraction information with corresponding scene states or scene state identifiers. By exchanging diffraction information between different devices/decoders/renderers, computational load for both users’ devices/decoders/renderers can be reduced, depending on their movement patterns within the audio scene.
[0128] In addition, since users (listeners) tend to behave similarly, diffraction information (diffraction data) is accumulated in particular for relevant (e.g., frequently occurring) scene states. This would be very difficult to achieve for encoder-side precomputation of diffraction information since the encoder does not have access to the actual listener positions and therefore can only assume them. Further, use of data storage (e.g., physical/shared storage or bitstream bandwidth) would be much more inefficient for encoder-side precomputation, due to part of the precomputed diffraction information relating to irrelevant or less relevant scene states in this case.
[0129] For example, the proposed functionality and techniques can create LUTs that correspond to the real 6D0F behavior of users (and not an assumed one at the encoder side), and thus may be said to relate to smart user-oriented LUT creation.
Representation and Coding of Acoustic Path Information
[0130] In the above, methods for providing, outputting, storing, exchanging, and/or reusing precomputed scene state information have been described. In conjunction with or addition to this, it may be of interest to provide for an interface for providing, outputting, storing, exchanging, and/or re-using acoustic path information between encoders and decoders / renderers (e.g., encoder - decoder, decoder 1 - decoder 2, or decoder 1 - decoder 1). For doing so, independently of specifics of computation and exchange of the acoustic path information, the present disclosure provides a scheme for efficient representation (e.g., coding) of acoustic path information, for example for storage or transmission. As such, this scheme could be used as a standalone, or in conjunction with techniques described elsewhere in the present disclosure.
[0131] Here, the acoustic path information may relate to initially precomputed acoustic path information, for example for all feasible pairs of start and end locations (start and end voxels) in the voxel grid (noting that for some pairs a valid path may not exist). In this case, the acoustic path information may be provided by an encoder, for example. Alternatively, the acoustic path information may relate to acoustic path information computed by a decoder / renderer at runtime, for example for later re-use by the same decoder or for use by a different decoder. The present disclosure provides different modes for representing (e.g., encoding) the acoustic path information, depending on the respective use case, for example depending on the amount and/or nature of the precomputed acoustic path information.
[0132] In a reference model (RM) for representing or encoding acoustic path information, pairs of “from -to” voxel coordinates (e.g., user and object voxel coordinates) are encoded per each acoustic diffraction path.
• if an acoustic diffraction path exists for the given voxel coordinate pair: a diffraction comer coordinate (determining position of the diffracted source) and diffraction path length (determining level/gain of the diffracted source) are explicitly coded and transmitted
otherwise: no additional information is coded and transmitted
[0133] However, the RM approach for the pre-computed acoustic path data (acoustic diffraction path data) representation has high redundancy due to:
• repetition of voxel coordinates for static scene states (in terms of either user or object position) and
• redundant path length representation accuracy (encoded as IEEE float32)
[0134] The present disclosure addresses both issues by:
• avoiding repetition of voxel coordinates (e.g., by a nested order of voxel pair coding) and
• application of differential diffraction path length coding (e.g., by sequential traversing of the voxel grid and representation of the path length in terms of number of voxels or its difference to the previous value)
[0135] In general, acoustic path information may comprise one or more items of path information (items of acoustic path information), each relating to a respective acoustic path. Each item of path information specifies a first location in a voxel grid, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel. The first location (e.g., first voxel) may relate to a start location (e.g., start voxel) of the acoustic path, such as the source location Svox defined above. The second location (e.g., second voxel) may relate to an end location (e.g., end voxel) of the acoustic path, such as the listener location Lvox defined above. The start location may be given by syntax element pcpdStartVoxelPacked defined below and the end location may be given by syntax element pcpdEndVoxelPacked defined below, for example. In any case, the assignment of first and second locations to start and end locations may also be the reverse in some implementations, depending on use cases and requirements. The corner voxel may be a voxel on the acoustic path for which the acoustic path changes direction. In some implementations, the comer voxel may additionally be required to be visible from the second location (e.g., end location), in the sense
that there must be a line of sight between the second location and the corner voxel in the voxel grid. If there should me more than one voxel meeting this definition, the voxel among the voxels meeting the definition that is farthest from the second location (or closest to the first location) may be designated as the comer voxel.
[0136] The present disclosure seeks to efficiently encode sequences of items of path information. In general, when encoding a given item of path information, techniques according to the present disclosure seek to re-use information related to a preceding item of path information. For example, even when one or both of the first location and the second location differ from one item of path information to the next, the corner voxel and/or the path length may be the same or similar.
[0137] To increase probability that information can be re-used among items of path information, techniques according to the present disclosure perform encoding and decoding in a nested manner. That is, items of path information in the sequence of path information are grouped in the aforementioned sequence by their respective first locations. Further, for each first location, items of path information are preferably grouped such that respective second locations of items of path information adjacent in said sequence are close to each other, e.g., adjacent in the voxel grid.
[0138] One example implementation of techniques according to the present disclosure seeks to re-use at least information on the corner voxel. An example of a corresponding method of encoding acoustic path information will now be described with reference to Fig. 12 and Fig. 13
[0139] Fig. 12 is a flowchart showing a method 1200 of processing audio scene information, in particular, encoding acoustic path information, including plural items of path information, for a given voxel-based audio scene representation.
[0140] At step S1210, a voxel-based audio scene representation of an audio scene is obtained (e.g., received, extracted from a bitstream, read from storage, etc.).
[0141] At step SI 220, for one or more (e.g., multiple) first locations and one or more (e.g., multiple) second locations in a two-dimensional voxel grid relating to the voxel-based audio scene representation, items of path information are successively encoded for respective
first locations and second locations. That is, the items of path information may be arranged in a given sequence (e.g., predefined sequence), and they may be encoded one after another in accordance with this sequence. This sequence may be such that items of path information specifying the same first location are subsequent to each other, i.e., for a sub-sequence within the sequence. In this sense, the aforementioned sequence may be said to relate to going through the relevant first and second locations in a nested manner, where items of path information specifying the same first location are grouped. Preferably, for each such group, the second locations specified by the items of path information in the group are gone through in accordance with a predefined pattern, as described in more detail below.
[0142] As noted above, each item of path information may specify a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel on the acoustic path for which the acoustic path changes direction. It is understood that the voxel grid may relate to a two-dimensional projection map generated from the voxel-based audio scene representation.
[0143] At step SI 230, for a current item of path information, an encoded item of path information is generated based on the (current) item of path information. The generated encoded item of path information includes at least an indication of the respective first location and an indication of the respective second location. It may include additional encoded information, as detailed below.
[0144] The further encoding procedure for the current item of path information differs (and thus, the further content of the encoded item of path information) depends on whether the comer voxel specified by the current item of path information is different from a comer voxel specified by a preceding item of path information (i.e., preceding in the sequence). It is understood that method 1200 may include a step of determining whether or not this is the case (not shown in the figure).
[0145] Step SI 240 relates to the case that the corner voxel specified by the current item of path information is different from the comer voxel specified by the preceding item of path information. Then, the encoded item of path information includes an indication of the (location of the) comer voxel. In other words, the indication of the corner voxel is included in (or added to) the encoded item of path information. This indication of the corner voxel may be an absolute,
explicit, and/or non-differential indication of the comer voxel, for example using voxel coordinates or a voxel index of the corner voxel. The indication may relate to syntax element pcpdSourceDirectionPacked defined below, for example. Additionally, the encoded item of path information may include an indication that the corner voxel is different from the one of the preceding item of path information, for example in the form of a single bit flag. This bit flag may relate to flag pcpdUsePrevSourceDirection == “false” defined below (in first mode, e.g., selective-mode defined below), or to flag pcpdUsePrevData == “false” defined below (in second mode, e.g., full-mode defined below).
[0146] Step S1250 relates to the case that the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information (i.e., the corner voxel specified by the current item of path information is identical to the comer voxel specified by the preceding item of path information). Then, the encoded item of path information includes an indication that the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, instead of the indication of the comer voxel. In other words, the indication that the corner voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information is included in (or added to) the encoded item of path information. This indication may relate to a single bit flag, for example. This bit flag may relate to flag pcpdUsePrevSourceDirection == “true” defined below (in first mode, e.g., selectivemode), or to flag pcpdUsePrevData == “true” defined below (in second mode, e.g., full-mode).
[0147] Method 1200 may further comprise a step of outputting the encoded items of path information to a bitstream (not shown in the figure), for example for transmission or storage.
[0148] Starting from the technique set out above, the present disclosure provides two different modes (coding modes) that may be selected for encoding the acoustic path information: a first mode (e.g., selective-mode) that can be used for example when only a relatively small number of items of path information are to be encoded, and a second mode (e.g., full-mode) that can be used for example when large numbers of items of path information, for example for all feasible pairs of first and second locations in an acoustic scene, are to be encoded. Whether the first or second mode is used for an item of path information (or a sequence of items of path information) may be signaled by a flag included in the bitstream, valid for the whole sequence,
or in the encoded item of path information. This flag may be, for example, flag pcpdFullMode as defined below.
[0149] In the first mode (e.g., pcpdFullMode == “false”), the generated encoded item of path information further includes an indication of the path length (e.g., pcpdPathLength defined below). This indication of the path length may be an absolute, explicit, and/or non-differential indication of path length, for example in units of voxels or voxel edge length. Thus, the first mode may re-use a previous indication of the comer voxel, but may include an indication of the path length regardless of whether the corner voxel changes from one item of path information to the next.
[0150] As an alternative implementation of the first mode, the generated encoded item of path information may include, if the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, an indication (e.g., 1 -bit flag) of whether the encoded item of path information includes the indication of the path length or an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information. Depending on this indication, the encoded item of path information may then include the indication of the path length or the indication of the difference between the path length specified by the current item of path information and the path length specified by the preceding item of path information. The latter may relate to a 2 -bit value, for example.
[0151] In the second mode (e.g., pcpdFullMode == “true”), the generated encoded item of path information not necessarily includes an indication of the path length. Namely, in the second mode, if the comer voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information further includes an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information. This indication may be in the form of syntax element pcpdPathLengthDelta defined below, for example.
[0152] For certain orders of the items of path information, the above difference between path lengths can be very efficiently encoded. Thus, for encoding in the second mode, for each first location, the voxel grid may be traversed according to a predetermined pattern to determine
a sequence of second locations and, for the respective first location, items of path information are successively encoded for the determined sequence of second locations. That is, the traversing of the voxel grid determines a sequence of items of path information, and encoding is performed according to this sequence.
[0153] One example of such predetermined pattern traverses the voxel grid in a rasterscanning manner, along rows and columns of the voxel grid.
[0154] When using such pattern, the aforementioned difference of path lengths can be encoded very efficiently, by using only 2 bits. In other words, the aforementioned difference may (only) take one of four predetermined values, in the sense that these four predetermined values are sufficient for encoding the aforementioned difference (assuming that the comer voxel between the current and preceding items of path information is the same). An example of these four predetermined values is given below in Table 3, where it is assumed that the voxel edge length is 1.
[0155] In line with the above, the bitstream may include, for a sequence of encoded items of path information, an indication of whether a first mode or a second mode is used (e.g., bit flag pcpdFullMode defined below). Further, each encoded item of path information may have the following content:
• an indication of the first location (e.g., pcpdStartVoxelPacked defined below)
• an indication of the second location (e.g., pcpdEndVoxelPacked defined below), and
• an indication of whether an acoustic path exists for the first location and the second location (e.g., bit flag pcpdPathExists defined below)
[0156] If the first mode (e.g., selective-mode) is used (e.g., pcpdFullMode == “false”), the encoded item of path information further comprises:
• the aforementioned indication of whether the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by a preceding item of path information (e.g., a 1 -bit flag, such as bit flag pcpdUsePrevSourceDirection defined below)
• if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information (e.g., pcpdUsePrevSourceDirection == “true”), an indication of the path length (e.g., pcpdPathLength defined below)
• otherwise (e.g., pcpdUsePrevSourceDirection == “false”), the indication of the comer voxel (e.g., pcpdSourceDirectionPacked defined below) and the indication of the path length (e.g., pcpdPathLength defined below)
[0157] In an alternative embodiment for the first mode (e.g., selective-mode), the encoded item of path information may further comprise, instead of the above:
• the aforementioned indication of whether the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by a preceding item of path information (e.g., a 1 -bit flag, such as bit flag pcpdUsePrevSourceDirection defined below)
• if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information (e.g., pcpdUsePrevSourceDirection == “true”), an indication (e.g., a 1 -bit flag) of whether the path length is coded non-differentially (e.g., in absolute terms, or explicitly) or differentially (e.g., in relative terms, as a delta value) o if the path length is coded non-differentially, a (non-differential, absolute, or explicit) indication of the path length (e.g., pcpdPathLength defined below) o if the path length is coded differentially, an indication of a difference between a previous path length and a current path length (e.g., a 2 -bit value), wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information
• otherwise, if the corner voxel specified by the item of path information corresponding to the encoded item of path information is not the same as the comer voxel specified by the preceding item of path information (e.g., pcpdUsePrevSourceDirection == “false”), the
indication of the corner voxel (e.g., pcpdSourceDirectionPacked defined below) and the indication of the path length (e.g., pcpdPathLength defined below)
[0158] If the second mode (e.g., full-mode) is used (e.g., pcpdFullMode == “true”), the encoded item of path information further comprises:
• the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the comer voxel specified by the preceding item of path information (e.g., bit flag pcpdUsePrevData defined below)
• if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information (e.g., pcpdUsePrevData == “true”), an indication of a difference between a previous path length and a current path length (e.g., a 2 -bit value, such as pcpdPathLengthDelta defined below), wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information
• otherwise (e.g., pcpdUsePrevData == “false”), the indication of the comer voxel (e.g., pcpdSourceDirectionPacked defined below) and the indication of the path length (e.g., pcpdPathLength defined below).
[0159] Fig. 13 is a flowchart showing a method 1300 of processing audio scene information, in particular, decoding acoustic path information, including plural (encoded) items of path information, for a given voxel -based audio scene representation. Decoding method 1300 may include steps mirroring those of the corresponding encoding method 1200. It is therefore understood that the data elements (e.g., indications, flags, etc.) mentioned below may correspond to those defined above in the context of method 1200. In other words, the bitstream received by method 1300 may be the bitstream output by method 1200.
[0160] At step S1310, a bitstream is received. The bitstream comprises a sequence of encoded items of path information for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to a voxel-based audio scene representation.
[0161] Each encoded item of path information corresponds to a respective item of path information that specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a comer voxel on the acoustic path for which the acoustic path changes direction. It is understood that the bitstream received at this step may be a bitstream as generated or output by method 1200 described above.
[0162] At step SI 320, encoded items of path information are successively decoded to generate (e.g., recover) corresponding items of path information.
[0163] As described above, the encoded items of path information may be arranged in a given sequence (e.g., predefined sequence) in the bitstream, and they may be decoded one after another in accordance with this sequence. This sequence may be such that encoded items of path information specifying the same first location are subsequent to each other, i.e., for a subsequence within the sequence. In this sense, the aforementioned sequence may be said to relate to going through the relevant first and second locations in a nested manner, where encoded items of path information specifying the same first location are grouped. Preferably, for each such group, the second locations specified by the encoded items of path information in the group are gone through in accordance with a predefined pattern, as described above.
[0164] The further decoding procedure for a current item of path information differs depending on whether the current encoded item of path information includes an indication that the comer voxel specified by the corresponding item of path information is the same as the comer voxel specified by an item of path information corresponding to a preceding item of encoded path information.
[0165] Thus, at step S1330, for the current item of encoded path information, it is determined whether the current encoded item of path information includes the indication that the comer voxel specified by the corresponding item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding item of encoded path information. This indication may be a 1 -bit flag, for example. This bit flag may relate to flag pcpdUsePrevSourceDirection (in first mode, e.g., selective-mode), or to flag pcpdUsePrevData (in second mode, e.g., full-mode).
[0166] Step SI 340 relates to the case that the corner voxel specified by the corresponding item of path information is the same as the corner voxel specified by the item of path information
corresponding to the preceding encoded item of path information (e.g., pcpdUsePrevSourceDirection == “true” in first mode or pcpdUsePrevData == “true” in second mode). In this case, the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information is set as the comer voxel for the item of path information corresponding to the current encoded item of path information. That is, the corner voxel of the preceding decoded item of path information is re-used as the corner voxel for the current decoded item of path information.
[0167] Step S1350 relates to the case that the corner voxel specified by the corresponding item of path information is different from the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information (e.g., pcpdUsePrevSourceDirection == “false” in first mode or pcpdUsePrevData == “false” in second mode). In this case, an (absolute, explicit, and/or non-differential) indication of the corner voxel (e.g., pcpdSourceDirectionPacked) is extracted from the current encoded item of path information.
[0168] As noted above, the present disclosure provides two different modes that may be selected for encoding and decoding the acoustic path information: the first mode (e.g., selectivemode) that can be used for example when only a relatively small number of items of path information are to be encoded, and the second mode (e.g., full-mode) that can be used for example when large numbers of items of path information, for example for all feasible pairs of first and second locations in an acoustic scene, are to be coded.
[0169] In the first mode, the encoded item of path information obtained from the bitstream further includes an indication of the path length (e.g., pcpdPathLength), as described above. This indication of the path length may be an absolute, explicit, and/or non-differential indication of path length, for example in units of voxels or voxel edge length. Thus, generating the item of path information corresponding to the current encoded item of path information further comprises extracting the indication of the path length from the current encoded item of path information.
[0170] In the second mode, the encoded item of path information not necessarily includes the indication of the path length, but it may alternatively include an indication of a difference (e.g., pcpdPathLengthDelta) between the path length specified by the current item of path
information and a path length specified by the preceding item of path information, if the comer voxel remains unchanged.
[0171] Accordingly, if the comer voxel specified by the corresponding item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, decoding in the second mode (e.g., at step SI 340) may further comprise extracting an indication of a difference between a previous path length and a current path length. Here, the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the current encoded item of path information.
[0172] As above, the one or more second locations may relate to locations obtained by, for each first location, traversing the voxel grid according to a predetermined pattern to define a sequence of second locations and, for the respective first location. Then, it is understood that the encoded items of path information are successively decoded in accordance with the sequence of second locations.
[0173] For example, the predetermined pattern may traverse the voxel grid in a rasterscanning manner, along rows and columns of the voxel grid. In this case, the aforementioned difference may be encoded by 2 bits, or in other words, may (only) take one of four predetermined values.
[0174] Next, an example implementation of the above scheme will be described.
[0175] This implementation, in line with the above, supports coding of pre-computed acoustic path information (e.g., acoustic diffraction path data) in:
• a “full” scene data mode, where the diffraction data is pre-computed for every voxel of the scene (i.e., for the low complexity rendering scenario), corresponding to the aforementioned second mode (e.g., full-mode) and
• a “selective” scene data mode, where the diffraction data is computed for a sub-set of voxels of the scene, corresponding to the aforementioned first mode (e.g., selective-mode)
[0176] These two modes provide capability to support different scenarios for applying coding according to the present disclosure, for example:
• “full” mode: “encoder-to-renderer” (e.g., all data is pre-computed and then used by Tenderers) and
• “selective” mode: “encoder/renderer-to-renderer” (e.g., some data is pre-computed, but new data can be added or exchanged during the rendering process)
[0177] Accordingly, the “selective” mode provides more real-time related capabilities (e.g., one Tenderer can act as an encoder for another Tenderer), but the “full” mode typically can provide a smaller coded data size and/or lower rendering complexity.
[0178] In one specific example, the bitstream syntax (data interface) may be defined as follows: if(pcpdDataPresent) { 1 bit pcpdNumStartPostions = escapedValue(8, 16,32) if(!pcpdFullMode) { // selective-mode 1 bit for(int i = 0; i < pcpdNumStartPostions; ++i) { pcpdNumEndPositions = escapedValue(8, 16,32) pcpdStartVoxelPacked; NBitsMap for(int j = 0; j < pcpdNumEndPositions; ++j) { pcpdEndVoxelPacked; NBitsMap if(pcpdPathExists) { 1 bit
// differential coding if(!pcpdUsePrevSourceDirection) { 1 bit pcpdSourceDirectionPacked; NBitsMap
} pcpdPathLength; 32 bit float
} } }
} else { // full-mode for(int i = 0; i < pcpdNumStartPostions; ++i) { if(pcpdDataPresentForPos) { 1 bit pcpdStartVoxelPacked; NBitsMap for(int x = 1; x <= voxSceneDimensions[0]; ++x) { for(int y = 1; y <= voxSceneDimensions[l]; ++y) { // pcpdEndVoxelPacked is defined by x and y if(pcpdPathExists) { 1 bit
// differential coding if(!pcpdUsePrevData) { 1 bit
pcpdSourceDirectionPacked;
NBitsMap pcpdPathLength; 32 bit float
} else { pcpdPathLengthDelta;
2 bit
}
}
}
}
}
}
}
}
Note: NbitsMap = cez7(/og2(voxSceneDimensions[0]*voxSceneDimensions[l]-l)
The above bitstream syntax may be seen as an alternative to the bitstream syntax given in
Table 1 above.
Table 3 — Value of pcpdPathLengthDelta
For the above example bitstream syntax, the encoding and decoding process may be as follows.
[0179] The encoder (or Tenderer) can select the pre-computed acoustic diffraction path data coding mode (signaled by pcpdFullMode or another suitable flag) according to the resulting or desired compression performance for the current application scenario or scene.
If an acoustic diffraction path exists (e.g., pcpdPathExists == “true”) in the “selective” coding mode (e.g., pcpdFullMode == “false”) and previous data is not available (e.g., pcpdUsePrevSourceDirection == “false”), the coordinate (e.g., packed coordinate) of the comer voxel and path length data is explicitly read from the bitstream (as in the RM); on the other hand, if previous data is available (e.g., pcpdUsePrevSourceDirection == “true”), the last transmitted voxel comer coordinate data is used in calculations. Notably, in the “selective” coding mode, the path length must be transmitted for each path (e.g., if pcpdPathExists == “true”).
pcpdSourceDirectionPacked = pcpdSourceDirectionPackedprev
[0180] Alternatively, in the “selective” coding mode, if previous data is available, the bitstream may include an indication (e.g., 1 -bit flag) of whether the path length is explicitly transmitted or can be coded differentially with reference to the previous data (e.g., using a 2 -bit value).
[0181] If an acoustic diffraction path exists (e.g., pcpdPathExists == “true”) in the “full” coding mode (e.g., pcpdFullMode == “true”) and previous data is not available (e.g., pcpdUsePrevData == “false”), the coordinate (e.g., packed coordinate) of the comer voxel and path length data is explicitly read from the bitstream (as in the RM); if previous data is available (e.g., pcpdUsePrevData == “true”), the last transmitted voxel comer and path length values are used in calculations. pcpdSourceDirectionPacked = pcpdSourceDirectionPackedprev pcpdPathLength = pcpdPathLength rev + Delta(pcpdPathLengthDelta)
[0182] The path length difference value (Delta) is encoded with 2 bits (e.g., by pcpdPathLengthDelta) because the path-length difference to the previous path length can be either +/-1 or +/-d, where d=sqrt(2)-l; since the path-length is calculated as the sum of lateral + diagonal steps on the uniform voxel grid.
[0183] Next, data compression performance estimation results for techniques according to the present disclosure will be described.
[0184] The bitrate comparison (VoxData payload only) for all MPEG-I CfP Testi scenes are shown in Table 4 and Fig. 14. Decoder output remains bit-exact with respect to the RM. Overall Testi bitrate savings -38% (full) and -16% (selective).
Conditions:
RM: Current reference model (v25 bitstream)
PCPD (full): Pre-computed path data coding (full-mode)
PDPD (selective): Pre-computed path data coding (selective-mode)
Table 4 — Date Compression Performance Results
Representation of Voxel Coordinates / Indices
[0185] The following efficient representation of voxel indices may be applicable for transmission or storage of both voxel grid and Diffraction Map (VoxDataDiffractionMap) entries, for example. It may substitute any fixed-length representations of voxel indices (voxel coordinates).
[0186] The following steps may be performed in the context of the proposed representation:
[0187] Step 1 : Determine the amount (i.e., number, count) of bits needed for the current grid resolution / diffraction map dimension. For a three dimensional voxel grid and a two dimensional diffraction map, these numbers NbitsVox and NbitsMap, respectively, may be determined for example as follows:
NbitsVox = ceil(log2(L*W*H-l))
NbitsMap = ceil(log2(L*W-l)), where L, W, H (Length, Width, Height) is the dimension of voxel grid and diffraction map. The values may differ for the voxel grid and diffraction map.
Step 1 may apply to both the encoder side and the decoder side.
[0188] Step 2: the voxel indices (x, y, z) and diffraction map indices (x, y) are mapped onto a packed representation index (Idx) and encoded using Nbits vox and Nbits map bits, respectively. In one embodiment (x, y, z) is zero-based and the packed representation indices may range from 0 to L*W*H-1 for voxels and from 0 to L*W-1 for the diffraction map. The mapping from the indices (x, y, z) onto the packed representation indices may be for example as follows:
Idx(x, y, z) = (x-1) + ((y-1) * L) + ((z-1) * L * W).
Step 2 may be performed at the encoder side only.
[0189] In the above, the packed representation index is an index that can uniquely identify a voxel in the voxel grid or diffraction map. Put differently, the voxels in the voxel grid may have assigned thereto a unique consecutive index, so that each voxel in the voxel grid can be uniquely identified by a single integer number. Accordingly, a packed representation index may be used for any indication of a voxel location in the voxel grid or in a two-dimensional map. In particular, the packed representation index may be used for indicating any voxel locations mentioned throughout the disclosure.
[0190] The assignment of unique indices to the voxels may be according to a predefined pattern. For example, the voxel grid may be scanned/traversed in x, y, and z directions, in this order, for consecutively assigning the unique index to respective voxels.
[0191] The mapping from the packed representation indices back onto the voxel and diffraction map indices may be for example as follows: x = floor(Idx) % L + 1 y = floor(Idx / L) % W + 1 z = floor(Idx / L / W) % H + 1, where % denotes the modulo operator.
Apparatus
[0192] While methods and processing chains have been described above, it is understood that the present disclosure likewise relates to apparatus (e.g., computer apparatus or apparatus
having processing capability in general) for implementing these methods and processing chains (or techniques in general).
[0193] An example of such apparatus 1500 is schematically illustrated in Fig. 15. The apparatus 1500 comprises a processor 1501 and a memory 1502 coupled to the processor 1501. The memory 1502 may store instructions for execution by the processor 1501. The processor 1501 may be adapted to implement the processing chains described throughout the disclosure and/or to perform methods (e.g., methods of processing audio scene information for audio rendering) described throughout the disclosure. The apparatus 1500 may receive inputs (e.g., audio scene description, listener location, etc.) and generate outputs (e.g., representations of diffraction information, acoustic path information, etc.).
Interpretation
[0194] Aspects of the systems described herein may be implemented in an appropriate computer-based sound processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of these systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0195] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and/or as data and/or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and/or other characteristics. Computer-readable media in which such formatted data and/or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0196] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in
hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and/or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, computer-implemented neural networks described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input/output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0197] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art.
Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0198] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
Enumerated Example Embodiments
[0199] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0200] EEE1. A method of processing audio scene information, the method comprising: obtaining a voxel-based audio scene representation of an audio scene;
for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to the voxel-based audio scene representation, successively encoding items of path information for respective first locations and second locations, wherein each item of path information specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel on the acoustic path for which the acoustic path changes direction; and for a current item of path information, generating an encoded item of path information based on the item of path information, wherein the encoded item of path information includes an indication of the respective first location and an indication of the respective second location, wherein, if the comer voxel specified by the current item of path information is different from a comer voxel specified by a preceding item of path information, the encoded item of path information includes an indication of the comer voxel; wherein, if the comer voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information includes an indication that the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, instead of the indication of the comer voxel.
[0201] EEE2. The method according to EEE1, wherein the encoded item of path information further includes an indication of the path length.
[0202] EEE3. The method according to EEE1, wherein, if the corner voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information further includes an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
[0203] EEE4. The method according to EEE1, wherein, if the corner voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information further includes:
an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information; and the indication of the path length or the indication of the difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
[0204] EEE5. The method according to any one of the preceding EEEs, further comprising: for each first location, traversing the voxel grid according to a predetermined pattern to determine a sequence of second locations and, for the respective first location, successively encoding items of path information for the determined sequence of second locations.
[0205] EEE6. The method according to EEE5, wherein the predetermined pattern traverses the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
[0206] EEE7. The method according to EEE5 or EEE6 when depending on EEE3, wherein the difference is encoded by 2 bits; or wherein the difference takes one of four predetermined values.
[0207] EEE8. The method according to any one of the preceding EEEs, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location, wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by a preceding item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of the path length; and
otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information; and otherwise, the indication of the corner voxel and the indication of the path length.
[0208] EEE9. The method according to any one of EEE1 to EEE7, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location, wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by a preceding item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information, together with the indication of the path length or the indication of the difference between the previous path length and the current path length; and
otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length.
[0209] EEE10. The method according to any one of EEE 1 to EEE9, further comprising: outputting the encoded items of path information to a bitstream.
[0210] EEE11. A method of processing audio scene information, the method comprising: receiving a bitstream comprising a sequence of encoded items of path information for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to a voxel-based audio scene representation, each encoded item of path information corresponding to a respective item of path information that specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel on the acoustic path for which the acoustic path changes direction; and successively decoding encoded items of path information to generate corresponding items of path information; wherein for a current item of encoded path information, generating the corresponding item of path information comprises: determining whether the current encoded item of path information includes an indication that the comer voxel specified by the corresponding item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding item of encoded path information; if the corner voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded
item of path information, setting the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information as the corner voxel for the item of path information corresponding to the current encoded item of path information; and if the corner voxel specified by the corresponding item of path information is different from the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, extracting an indication of the comer voxel from the current encoded item of path information.
[0211] EEE12. The method according to EEE11, wherein generating the corresponding item of path information further comprises: extracting an indication of the path length from the current encoded item of path information.
[0212] EEE13. The method according to EEE11, wherein, if the corner voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information further comprises: extracting an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the current encoded item of path information.
[0213] EEE14. The method according to EEE11, wherein, if the corner voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information further comprises: extracting an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the current encoded item of path information; and
extracting the indication of the path length or the indication of the difference between the previous path length and the current path length.
[0214] EEE15. The method according to any one of EEE 11 to EEE 14, wherein the one or more second locations relate to locations obtained by, for each first location, traversing the voxel grid according to a predetermined pattern to define a sequence of second locations and, for the respective first location, and the encoded items of path information are successively decoded in accordance with the sequence of second locations.
[0215] EEE16. The method according to EEE15, wherein the predetermined pattern traverses the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
[0216] EEE17. The method according to EEE 15 or EEE 16 when depending on EEE13, wherein the difference is encoded by 2 bits; or wherein the difference takes one of four predetermined values.
[0217] EEE18. The method according to any one of EEE 11 to EEE 17, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location; wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of the path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used:
the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information; and otherwise, the indication of the corner voxel and the indication of the path length.
[0218] EEE19. The method according to any one of EEE 11 to EEE 17, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location; wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of whether the encoded item of path information includes the indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information, together with
the indication of the path length or the indication of the difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information; if the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length.
[0219] EEE20. An apparatus, comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE 1 to EEE 19.
[0220] EEE21. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE19.
[0221] EEE22. A computer-readable storage medium storing the program of EEE21.
Claims
1. A method of processing audio scene information, the method comprising: obtaining a voxel-based audio scene representation of an audio scene; for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to the voxel-based audio scene representation, successively encoding items of path information for respective first locations and second locations, wherein each item of path information specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a corner voxel on the acoustic path for which the acoustic path changes direction; and for a current item of path information, generating an encoded item of path information based on the item of path information, wherein the encoded item of path information includes an indication of the respective first location and an indication of the respective second location, wherein, if the comer voxel specified by the current item of path information is different from a comer voxel specified by a preceding item of path information, the encoded item of path information includes an indication of the comer voxel; wherein, if the comer voxel specified by the current item of path information is the same as the comer voxel specified by the preceding item of path information, the encoded item of path information includes an indication that the corner voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, instead of the indication of the comer voxel.
2. The method according to claim 1, wherein the encoded item of path information further includes an indication of the path length.
3. The method according to claim 1, wherein, if the comer voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, the encoded item of path information further includes an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
4. The method according to claim 1, wherein, if the comer voxel specified by the current item of path information is the same as the corner voxel specified by the preceding item of path information, the encoded item of path information further includes: an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information; and the indication of the path length or the indication of the difference between the path length specified by the current item of path information and a path length specified by the preceding item of path information.
5. The method according to any one of the preceding claims, further comprising: for each first location, traversing the voxel grid according to a predetermined pattern to determine a sequence of second locations and, for the respective first location, successively encoding items of path information for the determined sequence of second locations.
6. The method according to claim 5, wherein the predetermined pattern traverses the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
7. The method according to claim 5 or 6 when depending on claim 3, wherein the difference is encoded by 2 bits; or wherein the difference takes one of four predetermined values.
8. The method according to any one of the preceding claims, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location,
wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by a preceding item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of the path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information; and otherwise, the indication of the corner voxel and the indication of the path length.
9. The method according to any one of claims 1 to 7, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location, wherein each encoded item of path information further comprises, if the first mode is used:
the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by a preceding item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the preceding item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information, together with the indication of the path length or the indication of the difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the preceding item of path information, an indication of a difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length.
10. The method according to any one of claims 1 to 9, further comprising: outputting the encoded items of path information to a bitstream.
11. A method of processing audio scene information, the method comprising: receiving a bitstream comprising a sequence of encoded items of path information for one or more first locations and one or more second locations in a two-dimensional voxel grid relating to a voxel-based audio scene representation, each encoded item of path information
corresponding to a respective item of path information that specifies a first location, a second location, a path length of an acoustic path between the first location and the second location in the voxel grid, and a comer voxel on the acoustic path for which the acoustic path changes direction; and successively decoding encoded items of path information to generate corresponding items of path information; wherein for a current item of encoded path information, generating the corresponding item of path information comprises: determining whether the current encoded item of path information includes an indication that the comer voxel specified by the corresponding item of path information is the same as the comer voxel specified by an item of path information corresponding to a preceding item of encoded path information; if the comer voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, setting the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information as the comer voxel for the item of path information corresponding to the current encoded item of path information; and if the comer voxel specified by the corresponding item of path information is different from the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, extracting an indication of the corner voxel from the current encoded item of path information.
12. The method according to claim 11, wherein generating the corresponding item of path information further comprises: extracting an indication of the path length from the current encoded item of path information.
13. The method according to claim 11, wherein, if the comer voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of
path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information further comprises: extracting an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the current encoded item of path information.
14. The method according to claim 11, wherein, if the comer voxel specified by the corresponding item of path information is the same as the comer voxel specified by the item of path information corresponding to the preceding encoded item of path information, generating the corresponding item of path information further comprises: extracting an indication of whether the encoded item of path information includes an indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the current encoded item of path information; and extracting the indication of the path length or the indication of the difference between the previous path length and the current path length.
15. The method according to any one of claims 11 to 14, wherein the one or more second locations relate to locations obtained by, for each first location, traversing the voxel grid according to a predetermined pattern to define a sequence of second locations and, for the respective first location, and the encoded items of path information are successively decoded in accordance with the sequence of second locations.
16. The method according to claim 15, wherein the predetermined pattern traverses the voxel grid in a raster-scanning manner, along rows and columns of the voxel grid.
17. The method according to claim 15 or 16 when depending on claim 13,
wherein the difference is encoded by 2 bits; or wherein the difference takes one of four predetermined values.
18. The method according to any one of claims 11 to 17, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location; wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of the path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded
item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information; and otherwise, the indication of the corner voxel and the indication of the path length.
19. The method according to any one of claims 11 to 17, wherein each encoded item of path information comprises: an indication of whether a first mode or a second mode is used; an indication of the first location; an indication of the second location; and an indication of whether an acoustic path exists for the first location and the second location; wherein each encoded item of path information further comprises, if the first mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by an item of path information corresponding to a preceding encoded item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of whether the encoded item of path information includes the indication of the path length or an indication of a difference between a previous path length and a current path length, wherein the previous path length is a path length specified by the item of path information corresponding to the preceding encoded item of path information and the current path length is a path length specified by the item of path information corresponding to the encoded item of path information, together with the indication of the path length or the indication of the difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length; and wherein each encoded item of path information further comprises, if the second mode is used: the indication of whether the corner voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified
by the item of path information corresponding to the preceding encoded item of path information; if the comer voxel specified by the item of path information corresponding to the encoded item of path information is the same as the corner voxel specified by the item of path information corresponding to the preceding encoded item of path information, an indication of a difference between the previous path length and the current path length; and otherwise, the indication of the corner voxel and the indication of the path length.
20. An apparatus, comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 19.
21. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 19.
22. A computer-readable storage medium storing the program of claim 21.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363508367P | 2023-06-15 | 2023-06-15 | |
| PCT/EP2024/065486 WO2024256238A1 (en) | 2023-06-15 | 2024-06-05 | Methods, apparatus, and systems for processing audio scene information |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4728508A1 true EP4728508A1 (en) | 2026-04-22 |
Family
ID=91465378
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24731910.6A Pending EP4728508A1 (en) | 2023-06-15 | 2024-06-05 | Methods, apparatus, and systems for processing audio scene information |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4728508A1 (en) |
| KR (1) | KR20260023046A (en) |
| CN (1) | CN121532825A (en) |
| WO (1) | WO2024256238A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6313841B1 (en) | 1998-04-13 | 2001-11-06 | Terarecon, Inc. | Parallel volume rendering system with a resampling module for parallel and perspective projections |
| US10293259B2 (en) | 2015-12-09 | 2019-05-21 | Microsoft Technology Licensing, Llc | Control of audio effects using volumetric data |
| TWI636422B (en) | 2016-05-06 | 2018-09-21 | 國立臺灣大學 | Indirect illumination method and 3d graphics processing device |
| US11606662B2 (en) | 2021-05-04 | 2023-03-14 | Microsoft Technology Licensing, Llc | Modeling acoustic effects of scenes with dynamic portals |
-
2024
- 2024-06-05 CN CN202480047238.1A patent/CN121532825A/en active Pending
- 2024-06-05 EP EP24731910.6A patent/EP4728508A1/en active Pending
- 2024-06-05 KR KR1020267001178A patent/KR20260023046A/en active Pending
- 2024-06-05 WO PCT/EP2024/065486 patent/WO2024256238A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| KR20260023046A (en) | 2026-02-20 |
| WO2024256238A1 (en) | 2024-12-19 |
| CN121532825A (en) | 2026-02-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250203316A1 (en) | Methods, apparatus, and systems for processing audio scenes for audio rendering | |
| JP7330306B2 (en) | Transform method, inverse transform method, encoder, decoder and storage medium | |
| CN110166757B (en) | Method, system and storage medium for compressing data by computer | |
| Xu et al. | Dynamic point cloud geometry compression via patch-wise polynomial fitting | |
| EP4728508A1 (en) | Methods, apparatus, and systems for processing audio scene information | |
| HK40118970A (en) | Methods, apparatus, and systems for processing audio scenes for audio rendering | |
| WO2025056788A1 (en) | Methods and apparatus for processing voxel-based scene representations | |
| KR20250022845A (en) | Method, system and device for acoustic 3D range modeling on voxel-based geometric representations | |
| JP2026505661A (en) | Method, apparatus, and system for processing an audio scene for audio rendering | |
| WO2024179939A1 (en) | Multi-directional audio diffraction modeling for voxel-based audio scene representations | |
| TWI878974B (en) | Apparatus and method for encoding or decoding ar/vr metadata with generic codebooks | |
| JP2025528679A (en) | Apparatus and method for encoding or decoding pre-computed data for rendering early reflections in an AR/VR system | |
| EP4674141A2 (en) | Apparatus and method for rendering multi-path sound diffraction with multi-layer raster maps | |
| IL317485A (en) | Real nodes extension in scene description | |
| WO2024234159A1 (en) | Point cloud coding method, point cloud decoding method, decoder, coder, and computer-readable storage medium | |
| WO2025015465A1 (en) | Encoding and decoding method, decoder, encoder, and computer readable storage medium | |
| HK40064136B (en) | Method and device for point cloud encoding and decoding | |
| CN117793353A (en) | Binocular image compression method and electronic equipment |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251205 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |