EP4666595A2 - Methods, apparatus, and systems for processing audio scenes for audio rendering - Google Patents
Methods, apparatus, and systems for processing audio scenes for audio renderingInfo
- Publication number
- EP4666595A2 EP4666595A2 EP24706949.5A EP24706949A EP4666595A2 EP 4666595 A2 EP4666595 A2 EP 4666595A2 EP 24706949 A EP24706949 A EP 24706949A EP 4666595 A2 EP4666595 A2 EP 4666595A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- voxel
- diffraction
- scene
- voxels
- audio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/50—Controlling the output signals based on the game progress
- A63F13/54—Controlling the output signals based on the game progress involving acoustic signals, e.g. for simulating revolutions per minute [RPM] dependent engine sounds in a driving game or reverberation against a virtual wall
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/55—Controlling game characters or game objects based on the game progress
- A63F13/57—Simulating properties, behaviour or motion of objects in the game world, e.g. computing tyre load in a car race game
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/165—Management of the audio stream, e.g. setting of volume, audio stream path
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
- H04S7/304—For headphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/305—Electronic adaptation of stereophonic audio signals to reverberation of the listening space
- H04S7/306—For headphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/307—Frequency adjustment, e.g. tone control
Definitions
- the present disclosure relates to techniques of processing audio scene information for audio rendering.
- the present disclosure is directed to voxel-based scene representation for a three-dimensional audio scene and audio rendering, taking into account diffraction effects caused by elements of the three-dimensional audio scene, for example, by applying diffraction modelling for audio occlusion.
- MPEG-I Moving Picture Experts Group Immersive Audio
- VR Virtual reality
- AR augmented reality
- MR mixed reality
- XR extended reality
- a 6DoF interaction extends a 3DoF spherical video/audio experience that is limited to head rotations (pitch, yaw, and roll) to include translational movement (forward/back, up/down, and left/right), to allow for navigation within a virtual environment (e.g., physically walking inside a room), in addition to the head rotations.
- head rotations pitch, yaw, and roll
- translational movement forward/back, up/down, and left/right
- Voxels For audio rendering in VR, AR, MR and XR applications, object-based approaches have been widely employed by representing a complex auditory scene as multiple separate audio objects, each of which is associated with parameters or metadata defining a location/position and trajectory of that object in the scene. Alternatively, audio rendering in such environments also uses higher order ambisonics (HO A). However, a new usage of “voxels” for rendering audio scenes is now being explored, such as for use of new immersive audio experiences. Voxels for audio rendering are relevant for media environments implemented in both hardware and software, such as video game and/or VR, AR, MR and XR environments.
- Voxel is a space volume with acoustic properties or audio rendering instructions assigned to it.
- Voxel size may be an encoder configuration parameter, and it can be (manually or automatically) selected according to a scene geometry level of details (e.g., in the range of 10 cm - 1 m).
- Voxels for audio rendering can be obtained by:
- diffraction modelling especially modelling of acoustic diffraction effects in virtual environments (e.g. virtual reality or game worlds) is often completely discarded or substituted by a direct signal propagation approach.
- the diffraction path may change when the user and/or the audio source move through the three-dimensional audio scene. Further, the diffraction path may change when the audio scene itself changes, for example by indicating a door or window that opens or closes, or the like. Frequent re- calculations of diffraction paths may be computationally expensive, which requires comparatively powerful computation devices for implementing computer-mediated reality applications and/or may negatively affect user experience in some cases. There is thus a need for improved techniques for diffraction modeling in three-dimensional audio scenes, particularly three-dimensional audio scenes utilizing voxels.
- the present disclosure provides methods of processing audio scene information (in particular, voxel-based audio scene information) for audio rendering, apparatus for processing audio scene information for audio rendering, computer programs, and computer- readable storage media, having the features of the respective independent claims.
- the method may include obtaining a voxelized representation of the audio scene.
- the voxelized representation may include a set of voxels that forms a connected geometric region on a voxel grid of the voxelized representation.
- the geometric region may have a cuboid shape and the set of voxels may include at least a first boundary voxel and a second boundary voxel defining the cuboid shape of the geometric region.
- the voxels in the geometric region may share a common voxel property.
- the audio scene may be updated and, in response to an update of the audio scene, the method may further include determining an updated pair of boundary voxels defining an updated cuboid region for the geometric region. Alternatively or additionally, the method may further include determining an updated voxel property for the voxels in the cuboid region. The method may yet further include generating a representation of the audio scene based on the set of voxels with the updated pair of boundary voxels and the updated voxel property.
- the proposed method provides a simple approach for compression of voxelbased scene representation (i.e. a voxelized representation for an audio scene), which allows simple and efficient scene update mechanism during a change of a three-dimensional audio scene, e.g. when a user and/or an audio source move through the three-dimensional audio scene.
- an updating method can be implemented at an encoder for providing an (updated) audio scene, i.e. an updated representation of the audio scene, to a decoder for rendering the updated representation of the audio scene.
- such proposed voxel-based scene representation approach may allow to create complex dynamic scenes and encode them efficiently.
- the method may further include determining, from a plurality of voxels of the voxelized representation, at least the first boundary voxel and the second boundary voxel for the set of voxels.
- the boundary voxels may describe a scene element which can be a part to be updated within the audio scene.
- existing data relating to the voxel property for the scene element may be overwritten depending on the updated pair of boundary voxels and/or the updated voxel property.
- the update of the audio scene may be based on an update condition including one or more of: time, user input, input from a presentation engine, or input from an application logics.
- an update attribute indicating the update condition may be included as metadata in a bitstream along with a compressed representation of the audio scene based on the determined set of voxels.
- the update of the audio scene may be further based on a trigger received from a user in real time.
- a method of decompressing a compressed voxel-based audio scene from a bitstream may include receiving the bitstream comprising a voxelized representation of the audio scene.
- the voxelized representation may include a plurality of voxels arranged in a voxel grid, and each voxel may have an associated voxel property.
- the method may include decoding a set of voxels that form a connected geometric region on the voxel grid and an indication that the voxels in the geometric region share a first voxel property.
- the method may also include decoding a subset of the set of voxels associated with a scene sub-element within the geometric region and an indication that the subset of the set of voxels is assigned with a second voxel property. Moreover, the method may further include generating an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels. The generation of the updated representation of the audio scene may involve overwriting the voxels of the subset with the second voxel property. Configured as above, the proposed method provides a simple approach for decompression of voxel-based scene representation (i.e.
- a voxelized representation for an audio scene which allows simple and efficient scene update mechanism during a change of a three-dimensional audio scene, e.g. when a user and/or an audio source move through the three-dimensional audio scene. It may be appreciated that such a decompression method can be implemented at a decoder for rendering an (updated) audio scene and an updated representation of the audio scene may be output to a Tenderer for rendering the updated representation of the audio scene.
- a method of processing an audio scene for three-dimensional audio rendering may include receiving a voxelized representation of the audio scene.
- the voxelized representation may comprise a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation.
- the method may further include decoding the received voxelized representation of the audio scene.
- the method may also include applying at least one three-dimensional (3D) smoothing filter to the decoded voxelized representation of the audio scene.
- the voxels in the geometric region may share a common voxel property.
- the common voxel property may comprise a reference indicative of a material describing an occlusion property associated with the set of voxels for the geometric region.
- the method may further include applying the at least one 3D smoothing filter to one or more occlusion coefficients indicative of the occlusion property.
- the at least one 3D smoothing filter may be defined based on data transmitted in a bitstream or may be hard-coded in a Tenderer.
- the at least one 3D smoothing filter may be applied to the decoded set of voxels of the voxelized representation of the audio scene.
- the at least one 3D smoothing filter may be associated with a scene element identifier for a geometric region within the audio scene. In this case, the at least one 3D smoothing filter may be applied to the geometric region having that scene element identifier.
- the at least one 3D smoothing filter may be applied depending on a user position.
- the proposed method enables an extra representation capability that can transform a regular shape (e.g. a rectangular shape, a cuboid shape, etc.) of the geometric region formed by the set of voxels into a smoother shape for an object, thereby achieving a better correspondence to the intended geometry (e.g. from the content creator) without requiring a large amount of metadata overhead for transmission.
- applying the 3D filter(s) is a simple and convenient way for smoothing (averaging) the shape for the object, since it can be easily implemented at the decoder side. Accordingly, an improved audio scene representation approximating to a more realistic acoustic environment can be provided to the Tenderer in a bandwidth- and cost-efficient manner.
- a method of processing an audio scene for three-dimensional audio rendering may include obtaining a scene configuration indicative of the audio scene.
- the scene configuration may include a source position of an audio source, a given user position, and a scene description. More specifically, the scene description may include a voxel matrix for a voxelized representation of the audio scene, and associated occlusion and diffraction coefficients.
- the method may also include obtaining a two-dimensional (2D) projection map related to the voxelized representation of the audio scene.
- the method may also include determining a diffraction path between the audio source and the given user position based on the 2D projection map.
- the method may further include obtaining auralization data for rendering the audio scene based on a result of the determination.
- the determined (diffraction) path (which may be output by the path finding algorithm) contains sufficient information for generation of a virtual sound source at a virtual source position that realistically simulates the impact of sound diffraction in the original three-dimensional audio scene.
- this allows to provide a realistic listening experience in a three-dimensional audio scene at reasonable computational effort.
- this enables realistic sound rendering in three- dimensional audio scenes even for real time applications, such as virtual reality applications or computer/console games.
- the two-dimensional projection map may be obtained by applying a projection operation to the voxelized representation of the audio scene.
- the diffraction path may be determined based on one or more two-dimensional (2D) path search algorithms using the 2D projection map for the voxelized representation of the audio scene.
- the 2D projection map may be either received from a bitstream transmitted by an encoder or calculated at a Tenderer. In some embodiments, the 2D projection map may be calculated at a Tenderer by applying filtering to the voxel matrix of the scene description based on the source position and/or the given user position.
- the 2D projection map may be obtained by selecting among at least one of: a horizontal projection relating to a diffraction path and a vertical projection relating to the diffraction path.
- the horizontal projection may be related to the voxelized representation by a horizontal projection operation that projects onto a horizontal plane.
- the vertical projection may be related to the voxelized representation by a vertical projection operation that projects onto a vertical plane.
- the selection may be based on a route and/or a direction of the diffraction path.
- a set of diffraction coefficients may be assigned based on occluding object geometry and/or respective material properties.
- the method may further include selecting a subset of voxels in the voxel matrix based on the determined diffraction path. Specifically, the subset of voxels may cause one or more changes in a diffracted sound direction. Besides, the method may further include calculating coefficients of an EQ filter for a diffracted audio signal based on the selected subset of voxels. In particular, the calculation of the coefficients of the EQ filter may depend on at least one of a direct distance between an audio object and a listener, a length of the diffraction path, the diffraction coefficients of the corresponding voxels, and an angle of a change in a direction of the diffraction path.
- the method may further include applying diffraction modelling to the voxelized representation of the audio scene.
- the proposed method applying the 2D projection plane (map) of the 3D voxelbased scene representation for diffraction modelling allows to significantly reduce computational workload for the rendering tool (i.e. the Tenderer).
- the proposed method further allows to flexibly choose one or more suitable projection planes for the diffraction calculation based on the occluding structure.
- a further projection plane may be taken into account for introducing an additional path possible for the diffraction modelling, thereby increasing the accuracy of modelling the perceptual effect of acoustic occlusion/diffraction.
- a method of processing an audio scene at a rendering device for three-dimensional audio rendering may include receiving a three- dimensional audio scene and information on a sound source at a source position. The method may also include determining a rendering mode of the rendering device. The method may further include determining, for a given listener position, a virtual sound source at a virtual source position based on the source position to simulate an impact of acoustic diffraction by the three- dimensional audio scene on a source signal of the sound source at the source position.
- the determination of a virtual sound source may be based on one or more of: precomputed data along a predefined user path, and real-time data computed using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling, either of which may be selected depending on the given listener position and/or the determined rendering mode of the rendering device.
- the proposed method of processing an audio scene at a rendering device for three-dimensional audio rendering exploits diffraction modeling on a voxel-based audio scene representation to reduce computational complexity.
- the proposed method provides a possibility for selecting a suitable implementation for diffraction modelling to achieve complexity reduction and/or increased bandwidth efficiency, which allows to further optimize the hardware/software usage for rendering in real-time an audio scene of varying, complex three-dimensional virtual environments. It is further noted that the selection of an implementation for diffraction modelling may also depend on the device that performs the rendering in which different modes may be set to meet respective hardware/software requirements.
- the method may further include determining that a scene state of the three-dimensional audio scene is a known state and in response, determining the virtual sound source based on the pre-computed data along the predefined user path.
- the rendering mode may include a complexity-reduced mode and/or a bandwidth-efficient mode.
- the method may further include determining the virtual sound source based on the precomputed data along the predefined user path or applying diffraction modeling to the three- dimensional audio scene using the pre-computed data along the predefined user path.
- the method may further include determining the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling. Whether to use the pre-computed data along the predefined user path or the real-time data computed using the voxel-based diffraction modeling may depend on the given listener position.
- the method may further include determining the virtual sound source based on the realtime data computed using the mesh-based diffraction modeling and/or the voxel-based diffraction modeling. In this case, whether to apply the mesh-based diffraction modeling or the voxel-based diffraction modeling, or both, may depend on the given listener position.
- the method may also include determining a diffraction modeling order indicative of a complexity of a diffraction path for the given listener position. Moreover, the method may further include determining whether to use the mesh-based diffraction modeling or the voxel-based diffraction modeling based on the determined diffraction modeling order for the given listener position.
- the method may further include determining the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling.
- the method may further include determining, at the given listener position, the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling.
- the present disclosure proposes a computationally efficient method for modelling a three-dimensional (audio) scene, in particular, for acoustic diffraction modelling of a three-dimensional (audio) scene.
- the present disclosure utilizes a simplified (but sufficiently accurate) geometry representation using a voxelization-based method.
- a two-dimensional space for diffraction modelling of the relevant geometry representation may be used together with means to control sound effects approximating acoustical occlusion/diffraction phenomena for content creators and encoder operators, thereby allowing for perceptually realistic acoustic occlusion/diffraction effect simulation for dynamic and interactive three-dimensional virtual environments. Consequently, overall user experience can be enhanced to promote a broader deployment of virtual reality (VR) applications.
- VR virtual reality
- an apparatus for processing audio scene information for audio rendering may include a processor and a memory coupled to the processor and storing instructions for the processor.
- the processor may be configured to perform all steps of the methods according to preceding aspects and their embodiments.
- the computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., processor).
- a computing device e.g., processor
- the storage medium may store a computer program adapted for execution on a computing device (e.g., processor) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.
- a computing device e.g., processor
- Fig. 1(a) schematically illustrates an example of a processing chain for processing audio scene information for audio rendering
- Fig. 1(b) schematically illustrates an example of a diffraction path for a source location and a listener location in a voxel-based three-dimensional audio scene
- Fig. 1(c) is a flowchart illustrating an example of a method of processing audio scene information for audio rendering according to embodiments of the disclosure
- FIGs. 2(a)-2(c) schematically illustrate examples of part of a voxel-based audio scene according to embodiments of the disclosure
- Fig. 3 schematically illustrates an example of a voxel-based audio scene to which embodiments of the disclosure may be applied;
- Figs. 4(a)-4(c) schematically illustrate examples of an audio scene update describing the effect of opening/closing a sliding door according to embodiments of the disclosure
- Figs. 5(a) and 5(b) are flowcharts illustrating examples of a method of updating an audio scene for three-dimensional audio rendering according to embodiments of the disclosure
- Fig. 6 schematically illustrates an example of the filtering effect on a voxel block element according to embodiments of the disclosure
- Fig. 7 is a flowchart illustrating an example of a method of processing an audio scene using 3D filtering for three-dimensional audio rendering according to embodiments of the disclosure
- Figs. 8(a)-8(c) schematically illustrate examples of 2D projection planes of a 3D voxel-based scene representation for diffraction modelling according to embodiments of the disclosure
- Figs. 9 shows a flowchart illustrating an example of a method of processing an audio scene for three-dimensional audio rendering according to embodiments of the disclosure
- FIG. 10(a) schematically illustrates an example of a possible use case for techniques according to embodiments of the disclosure
- Fig. 10(b) is a diagram illustrating complexity measures as functions of time for different operating modes/implementations of processing audio scene information for audio rendering according to embodiments of the disclosure
- Fig. 11 illustrates a non-limiting example of rendering an indoor audio scene using different implementation approaches according to embodiments of the disclosure
- Fig. 12 is a flowchart illustrating an example of a method of processing an audio scene at a rendering device for three-dimensional audio rendering according to embodiments of the disclosure.
- Fig. 13 is a block diagram schematically illustrating an example of an apparatus implementing methods according to embodiments of the disclosure.
- a voxel is understood as a space volume with acoustic properties or audio rendering instructions assigned to it.
- the voxel size may be an encoder configuration parameter. It may be (manually or automatically) selected according to a scene geometry level of details (e.g., in the range of 10 cm - 1 m).
- a large audio scene can be represented as
- Any voxel-based representation of an audio scene may contain an indication of voxels that are not transmission voxels (e.g., that are occluder voxels), i.e., voxels in which sound cannot propagate or cannot freely propagate - a representation of occluding geometries.
- This indication may relate to an indication of coordinates (e.g., center coordinates, corner coordinates, etc.) of the respective voxels.
- the coordinates of these voxels may be represented by grid indices, for example.
- the voxel-based representation may include indications of material properties of the voxels that are not transmission voxels, such as absorption coefficients, reflection coefficients, etc..
- the voxel-based representation may also indicate transmission voxels (e.g., air voxels), i.e., voxels in which sound can propagate - a representation of sound propagation media.
- transmission voxels e.g., air voxels
- voxels in which sound can propagate - a representation of sound propagation media e.g., air voxels
- some implementations of voxel-based representations of audio scenes may include, for each voxel in a predefined section of space (e.g., within boundaries enclosing the audio scene), and indication of a respective material property.
- FIG. 1(a) schematically illustrates a processing chain 100 that can be used for processing audio scene information for audio rendering.
- the processing chain 100 can be used for converting voxel related data into parameters and signals needed for auralization (or audio rendering in general).
- the processing chain 100 may be implemented in software, hardware, or combinations thereof.
- the processing chain 100 may be implemented by a render er/decoder coupled to AR/VR/MR/XR equipment, such as AR/VR/MR/XR goggles.
- Specific implementations may include game consoles, set-top-boxes, personal computers, etc..
- the processing chain receives an audio scene description 20 from a bitstream (or storage/memory) 10.
- the audio scene description 20 may comprise a representation of a three- dimensional audio scene and information on a source location of a sound source within the audio scene.
- the representation of the three-dimensional audio scene may be voxel-based, for example.
- the processing chain 100 further receives an indication of a user position (listener location) 30 of a user (listener) within the audio scene.
- the audio scene description 20 and the user position 30 are provided to a diffraction direction calculation block (diffraction calculation block) 40 for determining (e.g., calculating) diffraction information.
- the diffraction information may relate to an acoustic diffraction path within the audio scene between the source location and the listener location.
- the diffraction information is then provided to a diffraction modeling tool 50 for applying diffraction modeling and optionally occlusion modeling, based on the diffraction information.
- the occlusion modeling calculates attenuation gains for the direct line between the listener and an audio source.
- the diffraction modeling tool 50 may output auralized audio data (3DoF auralizer data) that includes, for example, a location of an object to be rendered, an orientation, and frequency dependent gains.
- the diffraction modeling tool output may be further processed by other rendering stages such as Doppler, Directivity, Distance Attenuation, etc..
- the diffraction modeling tool 50 may be said to output diffraction information, as detailed below.
- the auralized audio data may then be used for audio replay, for example.
- a processing chain as shown in Fig. 1(a) may be used to convert voxel related data into the parameters for parameters and signals for auralization.
- the diffraction direction calculation block 40 and the diffraction modeling tool 50 may be seen as non-limiting examples of rendering tools.
- the rendering tools may generate 3DoF auralizer data.
- the scene description may include a voxel matrix and associated coefficients (e.g., reflection coefficients, occlusion coefficients, absorption coefficients, transmission coefficients etc.). These coefficients may be indicative of a material or material property of the respective voxel.
- the rendering tools may include, for example, occlusion and diffraction modelling tools.
- the 3DoF auralizer data may include, for example, object position, orientation and frequency dependent gains.
- the voxel-based representation of the three-dimensional audio scene defines psycho-acoustically relevant geometric elements and sound propagation media.
- the scene description may use the following parameters/interfaces (e.g., the following agreed upon data format, or agreed upon point of data exchange) to provide the information to rendering tools:
- Scene size in absolute units (e.g., meters) in number of voxels and/or voxel size
- Scene anchors in terms of coordinate anchors (to map absolute coordinates to voxel indices) in terms of scene anchors (to map sub-scene to sub-set of voxels)
- Scene content data reference to material properties that approximates acoustic effects caused by occluders (sound obstacles) located in the corresponding volume (e.g., coefficients for transmission, reflection, etc.) reference to sound propagation media properties that approximates an acoustic effect caused by media located in the corresponding volume (e.g., speed of sound, energy absorption, distance attenuation curve, etc.) rendering control parameter describing intended occlusion modelling effects
- audio signal IDs and/or signal gains the determines which signal is perceptually relevant (rendered) in the corresponding volume rendering control parameter describing intended reverberation modelling effects
- voxel type that controls reverberation settings e.g., RT60, DDR, RIR, etc.
- Scene content updates referenced to update triggering events
- All data can be audio object dependent (to support content creator intent in flexible audio scene authoring).
- the 3DoF auralizer data may include the following information: parameters and associated signals for the set of audio objects (and HO A) o parameters include the metadata output of the rendering tools (i.e., position, orientation and gains simulating effects of occlusion, diffraction, early reflections, parameters for reverberation coefficients, IR, etc.) o associated signals represent the audio output of the rendering tools (i.e., downmixed or replicated audio signals) scene state identifier (i.e., metadata allowing to map the scene description and user input to the 3DoF auralizer data)
- Fig. 1(b) illustrates an example of a possible scene state and a diffraction path for this scene state. It is understood that the scene state relates to or comprises the listener location 210 and the audio scene description (including the representation of the three-dimensional audio scene and the source location 220).
- Fig. 1(b) relates to a voxel based representation of the three-dimensional audio scene.
- This voxel based representation indicates “air” voxels or empty voxels (i.e., voxels in which sound can propagate, or transmission voxels) 230 and occluder voxels 240 (i.e., voxels in which sound cannot propagate or cannot freely propagate).
- occluder voxels may be understood to relate to voxels filled with a material other than air, and that can reflect, block, or otherwise alter sound propagation.
- the representation may further indicate respective transmission, reflections coefficients and potentially absorption coefficients relating to material properties of these voxels. These coefficients may be linked to ID’s or indices of their respective voxels in the voxel-based representation.
- the voxel based representation may define psycho-acoustically relevant geometric elements and sound propagation media in the audio scene.
- a listener location 210 is indicated by a parameter Lvox and a source location 220 is indicated by another parameter Svox.
- a diffraction path between the source location 220 and the listener location 210 may be determined using a pathfinding algorithm that takes the listener location 210, the source location 220, and the representation of the three-dimensional audio scene (or a two-dimensional representation, e.g., 2D projection or 2D matrix, derived therefrom) as inputs.
- a pathfinding algorithm that takes the listener location 210, the source location 220, and the representation of the three-dimensional audio scene (or a two-dimensional representation, e.g., 2D projection or 2D matrix, derived therefrom) as inputs.
- an algorithm for determining the diffraction information may take the listener location 210, the source location 220, and the representation of the three-dimensional audio scene as inputs and may output a location of a diffraction corner 250, indicated by Cvox and the variable n n representing the length of the diffraction path.
- the diffraction information may be determined based on:
- [Cvox, r in ] DiffractionDirectionCalculation Li NO x, Svox, VoxDataDiffractionMap)
- DiffractionDirectionCalculation indicates the algorithm for determining the diffraction information (“pathfinding algorithm”)
- VoxDataDiffractionMap indicates the voxel-based representation of the three-dimensional audio scene or a processed version thereof (e.g., 2D projection or 2D matrix derived therefrom).
- Cvox is understood to indicate the coordinates of the diffraction corner (e.g., coordinates, voxel/grid coordinates, or voxel/grid indices of the respective voxel including the diffraction corner).
- DiffractionDirectionCalculation may involve any viable pathfinding algorithm, such as the Fast traversal algorithm for ray tracing (cf. Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87.) and the JPS algorithm (cf. Harabor, D.D. and A. Grastien, Online Graph Pruning for Pathfinding On Grid Maps.
- the Fast traversal algorithm for ray tracing cf. Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87.
- JPS algorithm cf. Harabor, D.D. and A. Grastien, Online Graph Pruning for Pathfinding On Grid Maps.
- the corresponding 2D projection plane may be similar to a floor plan that describes a “sound propagation path topology”.
- the pathfinding algorithm is assumed to output a diffraction path that connects the source location 220 to the listener location and that consist of a plurality of straight path segments (line segments) that are sequentially linked end-to-end. Each transition from one path segment to another path segment relates to a change of direction of the diffraction path.
- the diffraction corner Cvox may be determined as a voxel that lies on or on the proximity of the diffraction path and is adjacent to a corner voxel (in a set of voxels representing corner voxels on the diffraction map, Cset) of the diffraction map (indicated by the voxel-based representation).
- the diffraction corner Cvox may be selected from a set of voxels (Pset) forming the diffraction path as a voxel that is close to a ‘visible’ (from the listener position Lc) corner voxel (belonging to Cset) causing the path (Pset) to change direction. If there are more than one such corners, the one furthest away from the listener location along the diffraction path (P se t) is selected.
- Pset set of voxels
- the diffraction path algorithm may be said to determine diffraction information relating to the acoustic diffraction path within the audio scene between the source location and the listener location.
- This diffraction information may be sufficient information for the Tenderer to recover/determine a virtual source location of a virtual audio source that encapsulates effects of acoustic diffraction effects. This is the case for the coordinates of the diffraction corner Cvox and the diffraction path length Tin.
- the virtual source location may be recovered by calculating the direction (e.g., azimuth, or azimuth and elevation) of the diffraction corner when seen from the listener location. Using this direction and taking the path length Tin of the diffraction path as the virtual source distance to the listener location, the virtual source location can be determined.
- the diffraction information can be represented in different ways. One option, as noted above, is diffraction information including/storing the path length n n and the coordinates (e.g., grid coordinates, etc.) of the diffraction corner Cvox.
- An example of the scene state Ni may be represented by
- Ni ⁇ Lvox, Svox, VoxDataDiffractionMap ⁇ , i.e., may relate to or comprise the listener location Lvox, the source location Svox and the voxelbased representation (e.g., VoxDataDiffractionMap) of the audio scene.
- a scene state identifier for a scene state Ni may be defined as
- SceneStateldentifier HASH(Ni), where HASH is a hash function that generates a hash value for scene state Ni, e.g., that maps scene states to fixed-size values.
- the scene state identifier may be said to be indicative of a certain scene state or to identify a certain scene state.
- diffraction information N2 may be represented by
- N2 ⁇ Cvox, Ln ⁇ N2 ⁇ Cvox, Ln ⁇ , where n n is the path length of the diffraction path and C vox indicates the location (e.g., voxel location) of the diffraction corner, as described above.
- a quantized version of the diffraction information N2 may be indicated by N3, where
- N3 voxSceneDiffractionPreComputedPathData(Ni)
- the diffraction information for example Cvox and n n , may also be seen as relating to 3DOF auralizer data, because user the position voxel coordinates L vox are fixed.
- a technical benefit and effect according to techniques of the present disclosure is that the scene state identifier or other information derived from the scene state may be used to avoid application of the diffraction modeling tools or rendering tools if the corresponding processing was already done for this scene state and the diffraction information or 3DoF auralizer data are available.
- the Tenderer can access the diffraction information/3DoF auralizer data (for a known scene state) without application of the rendering tools by: re-using the data calculated before (precomputed), or applying the data calculated by another Tenderer.
- a technical benefit is thus that techniques according to the present disclosure relate to lossless functionality aiming at the low complexity mode (complexity vs bitrate). Depending on the virtual environments to be rendered, it is allowed to flexibly choose an implementation mode suitable for efficient rendering of an audio scene in real-time.
- Fig. 1(c) is a flowchart showing an example of a method 300 of processing audio scene information for audio rendering in accordance with embodiments of the present disclosure.
- Method 300 may be implemented in software, hardware, or combinations thereof.
- the processing chain 100 may be implemented by a render er/decoder coupled to AR/VR/MR/XR equipment, such as AR/VR/MR/XR. goggles.
- Specific implementations may include game consoles, set-top-boxes, personal computers, etc..
- Method 300 comprises steps S310 through S350 that may be performed, for example, by a decoder/renderer. These steps may be performed, for example, whenever the scene state changes.
- a change of the scene state could relate to one or more of a change of the listener location 210, a change of the source location, and a change of the (representation of the) three-dimensional audio scene.
- steps S310 through S350 may be performed for each of a plurality of processing cycles of a decoder/renderer.
- step S310 an audio scene description is received.
- the audio scene description comprises a representation of a three-dimensional audio scene and information on a source location of a sound source within the audio scene.
- the audio scene description may comprise elements S vox and VoxDataDiffractionMap defined above, for example.
- step S320 information of a listener location of a listener within the audio scene is received.
- the listener location may correspond to element Lvox defined above, for example.
- diffraction information relating to an acoustic diffraction path within the audio scene between the source location and the listener location is obtained.
- the obtained diffraction information may be indicative of a virtual source location of a virtual sound source.
- the virtual source location may have the same direction (e.g., azimuth, or azimuth and elevation), when seen from the listener location, as the diffraction corner Cvox.
- the virtual source distance may correspond to the length n n of the diffraction path.
- the diffraction information may comprise indications of C vox and n n defined above.
- audio rendering is performed for the sound source based on the diffraction information. This may include, for example, diffraction modeling.
- a virtual source location of a virtual source may be determined based on the diffraction information.
- the virtual source may be an audio source that encapsulates effects of acoustic diffraction between the source location and the listener location in the three-dimensional audio scene.
- the virtual source location may be determined based on Cvox and n n by
- Audio rendering may then include rendering the virtual sound source at the virtual source location, for example.
- a representation of the diffraction information is output.
- outputting the representation of the diffraction information may comprise outputting a data element comprising the diffraction information and information on the scene state.
- the scene state may comprise the audio scene description (e.g., S vox and VoxDataDiffractionMap) and the listener location (e.g., Lvox).
- the output may be provided to a look up table (LUT).
- the LUT includes, as its entries, different items of diffraction information indexed with information on respective scene states (e.g., indexed with respective scene state identifiers). This LUT thus may be said to include the diffraction information and information on the scene state.
- the LUT can be stored and/or provided to be later retrieved, for example by other decoders, from a bitstream or from a shared storage (e.g., cloud or server based), for example by application request.
- a hash value of the scene state or the scene state identifier can be used to retrieve the actually desired entry from the LUT.
- the representation of the diffraction information may be output to a bitstream (e.g., outgoing bitstream) and/or to a storage (e.g., a memory, cache, file, etc.).
- the storage may be local or it may be shared (e.g., cloud based).
- the representation of the diffraction information may be output to a suitable medium for storing digital information or computer related information. The output may at least partially be directed to an external or shared data source or data repository.
- the representation of the diffraction information may be output as part of a voxSceneDiffractionPreComputedPathDataQ syntax element according to ISO/IEC 23090-4 (Coded representation of immersive media — Part 4: MPEG-I immersive audio, https://www.iso.org/standard/84711.html), or according to any future standard deriving therefrom.
- the scene geometry is determined by a set of such pairs of points (as examples of representations of geometric regions), for example:
- ⁇ VoxBox id, material, Point s, Point_E/> may correspond to the scene element defined above, that is, a geometric region (e.g., defined by the pair of points) together with its voxel property (e.g., P_ID) and optionally its identifier (e.g., V ID).
- voxel property e.g., P_ID
- V ID e.g., V ID
- the voxel size may be determined by the number of voxels as
- a voxel-based representation of an audio scene may include representations or indications of one or more cuboid geometric regions (cuboid space regions, cuboid volumes) that have identical (i.e., same, common) acoustic properties (e.g., material, absorption coefficients, reflection coefficients, etc.) or identical rendering instructions.
- the acoustic property or rendering instruction for a given voxel may be non-limiting examples of a voxel property of the given voxel.
- the representations of indications of the cuboid geometric regions may relate to scene elements, for example. It is understood that the cuboid regions are each non-trivial, in the sense that they each comprise more than a single voxel, and consist of a connected (i.e., contiguous) set of voxels.
- first and second boundary voxels e.g., the above pair of points.
- first and second boundary voxels may relate to diametral corners (extreme-corner voxels) of the cuboid, such as the extreme- corner voxel with the smallest x, y, z coordinate values or coordinate indices, and the extreme-corner voxel with the largest x, y, z coordinate values or coordinate indices, for example.
- diametral extremecorner voxels are feasible as well.
- each geometric region may be represented by at least an indication of the first and second boundary voxels and an indication of the common voxel property of the voxels within the geometric region.
- the representation of the geometric region may include an identifier (ID) of the geometric region.
- Fig. 2(a) shows an example of a geometric region 1002 in a voxel grid 1007.
- the geometric region 1002 comprises a plurality of voxels 1003 that form a cuboid.
- the shape or size of the geometric region can be represented by first and second boundary voxels 1004-1, 1004-2, such as extremecorners or extreme-corner voxels, which in this example correspond to the lower left front corner and the upper right back corner, respectively.
- Fig. 2(b) shows a side view of the geometric region 1002, looking from the front right in Fig. 2(a).
- each next pair of points may re-define the voxel properties in the corresponding cuboid region. That is, voxel properties of subsequent geometric regions may overwrite any previously assigned voxel properties for the voxels of the subsequent geometric region.
- smaller geometric regions that are fully contained within larger geometric regions may redefine or overwrite voxel properties of voxels in the smaller geometric region with the voxel property of the smaller geometric region.
- corresponding voxel properties are overwritten, while any other voxel properties are maintained.
- the subsequent geometric region defines acoustic properties of its voxels
- these acoustic properties will be used to overwrite the acoustic properties defined for the voxels of the previous geometric region, but any rendering instructions of the voxels of the previous geometric region will be maintained.
- the order among geometric regions may be derived from whether geometric regions are fully contained within each other, may be derived from an order of representations or indications of the geometric regions in a bitstream, or may be derived from a predefined order referencing identifiers (IDs) of the geometric regions, for example.
- IDs identifiers
- Fig. 2(c) schematically illustrates how voxel properties assigned to a first geometric region defined by extreme-corner voxels 1004-1, 1004-2 may be overwritten by voxel properties of a second geometric region 1002.
- Fig. 3 schematically illustrates an example of a voxel-based description of a three-dimensional audio scene that includes a plurality of cuboid volumes of common voxel properties, potentially with smaller cuboid sub-volumes that redefine voxel properties.
- the audio scene may include voxels 1101 that are local sound occluders that occlude inside a given acoustic environment of the audio scene.
- the audio scene may additionally include voxels 1102 that are local sound occluders that separate acoustic environments from each other.
- voxel-based representation can be compressed by defining the shape of the scene geometry (geometric regions) by a set of pairs of points (i.e. boundary voxels), which may also be referred to scene definitions each of which may relate to a specific scene state.
- a scene update mechanism that allows to switch between different scene states resulting in a scene update, for example, when a user and/or an audio source move through the three-dimensional audio scene.
- Such an updating mechanism can be implemented at an encoder for providing an updated audio scene, i.e. an updated representation of the audio scene, to a decoder, or can be implemented at a decoder, for rendering the updated representation of the audio scene.
- the above-described compression approach for voxel-based scene representation may further allow a simple and efficient scene update mechanism, which can be based on the similar concept of redefining or overwriting voxel properties of voxels (i.e. the old data) in the smaller geometric region with the voxel property (i.e. the new data) for the smaller geometric region.
- the new set of the voxel geometry specification can be associated with the following attributes defining the conditions for the update application:
- time e.g. absolute, time offset, etc.
- an updated region for the scene geometry may also be defined by an updated pair of boundary voxels.
- An update of the audio scene or the scene geometry having an updated region may relate to another scene definition different to the previous (non-updated) ones.
- scene definitions based on the voxel-based representation as described above can likewise be used while taking also into account additional conditions regarding e.g. when the respective scene definitions may take place.
- Fig. 4 a simple example of an audio scene update describing the effect of opening/closing a sliding door is illustrated in Fig. 4.
- a geometric region 1002 in a voxel grid 1007 within an audio environment 1000 is shown in Fig. 4(a).
- the geometric region 1002 comprises a plurality of voxels 1003, and the shape or size of the geometric region 1002 can be represented by first and second boundary voxels 1004-1, 1004-2, such as extreme-corners or extreme-corner voxels, which in this example correspond to the lower left front corner and the upper right back corner, respectively.
- the geometric region 1002 can be viewed as, for example, a wall within the audio environment 1000.
- the geometric region 1002 of Fig. 4(a) contains a door part 1005 which can be closed or opened by, for example, sliding it, based on a trigger of the user.
- the user may come to the sliding door 1005 and press a button for triggering the opening or closing of the sliding door 1005.
- Fig. 4(b) and Fig. 4(c) refer to two different scene states in response to an update representing the sliding door effect: in case of opening the sliding door 1005, the audio scene will be switched/updated from Fig. 4(b) to Fig. 4(c), while, in case of closing the sliding door 1005, the audio scene will be switched/updated from Fig. 4(c) to Fig. 4(b).
- more than two scene states representing the process of moving (opening or closing) the sliding door 1005 may also be taken into account for rendering the corresponding movement of the sliding door 1005.
- the door part 1005 may be formed by a set of voxels among the plurality of voxels 1003 as a sub-region of the geometric region 1002.
- the set of voxels within the sub-region may be redefined or overwritten to create the door part 1005. That is, the set of voxels within the sub-region may be associated with the door part 1005.
- the shape or size of the door part 1005 can be represented by e.g. extreme-corners or extreme-corner voxels as indicated by the boundary voxel at the upper left front corner 1005-1 and the lower right back corner (not shown) in Fig. 4(a). It is noted that the voxel positions as shown in Fig.
- 4(a) are non-limiting examples illustrating positions of possible boundary voxel which can be selected for representing a specific geometric region, and voxels at other positions such as at the lower left front corner and the upper right back corner may also be selected as boundary voxels for representing the same geometric region.
- the state of the door 1005 will be switched from 1005b to 1005c, corresponding to a reduced door size caused by the opening of the door as shown in Fig. 4(b) and Fig. 4(c), respectively.
- a gap portion 1006 between the door portion 1005b, 1005c and the wall 1002 will be enlarged from the state 1006b to state 1006c as shown in Fig. 4(b) and Fig. 4(c), respectively.
- the voxel property of the voxels within the door portion and/or within the gap portion may be overwritten (or redefined) with an updated voxel property to obtain an updated scene definition corresponding to the scene state of Fig. 4(c).
- the voxel property associated with the sliding door 1005 to be updated/overwritten may include, for example, the material of the door portion (such as wood, metal, etc.) and/or the material of the air.
- the material of wood/metal for the door portion may be overwritten with the material of the air from one voxel to another within a region to be updated (i.e. from door to gap) which may be defined by an updated pair of boundary voxels.
- voxels 1006b-l, 1006b-2 and 1006c-l, 1006c-2 may be determined as boundary voxels for the region being updated to be the gap portion.
- boundary voxels 1006b-l and 1006b-2 define the corresponding gap portion 1006b for the scene state of Fig. 4(b)
- boundary voxels 1006c-l and 1006c-2 define the corresponding gap portion 1006c for the scene state of Fig. 4(c).
- the updated pair of boundary voxels 1006c-l and 1006c-2 defining the updated gap portion (region) 1006c will be applied, within which the voxel property indicative of the material of wood/metal for the door portion may be overwritten with the material of the air for turning the associated voxels into the gap portion.
- the material of the air within the associated voxel property may be overwritten with the material of wood/metal from one voxel to another within a region to be updated (i.e. from gap to door) which may be defined by an updated pair of boundary voxels. That is, when the scene state is to be switched from Fig.
- the pair of boundary voxels defining the gap portion (region) will be updated, i.e., the gap portion will be updated from 1006c (defined by 1006c-l and 1006c- 2) to 1006b (defined by 1006b-l and 1006b-2).
- the updated pair of boundary voxels 1006b-l and 1006b-2 defining the updated gap portion (region) 1006b will be applied to specify the voxels (e.g., the voxels within the gap portion 1006c but outside the gap portion 1006b) of which the associated voxel property relating to the material of air shall be overwritten with the material of wood/metal for the door portion.
- the above described audio scene update representing the effect of opening/ closing a sliding door shall be regarded as a non-limiting example for the scene update mechanism as proposed by the present disclosure.
- the proposed scene update mechanism may also be applied to any other audio scenes obtainable by the voxel-based representation taking into account a variety of mater ial/medium properties for a given voxel and varying shape/geometry/arrangement of audio scene elements as mentioned above.
- the proposed scene update mechanism also takes into account one or more update conditions for the respective scene definitions (states).
- the different scene states may be defined together with the update conditions to specify when or how a specific scene state may take place.
- the following definition may start which overwrites the material of door portion with the material of the air to obtain a gap between the door portion and the wall.
- This change of scene may take place immediately, or based on a time condition specifying that the change shall take place rather after a period of time, e.g. in 0.5 second. For example, the gap will become larger in (another) 0.5 second, and the door will be opened completely in another 0.5 second.
- Similar definitions may be assigned to any possible (update) conditions, e.g. time, user interaction, etc., as mentioned above.
- the previous concept for representing scene definitions without a condition can be regarded as just initializing/predefining the initial state of the scene which may then be realized using conditional updates, i.e. different states of the scene are to be carried out based on the updated condition(s).
- the definitions may also include time delay, space, the condition(s) that leads to the present state, etc. which may be included in the bitstreams.
- a new definition/state may be obtained after an update, which may also depend on the external triggering by the user (i.e. not included in the bitstreams).
- updates may take place in real-time based on the input from the representation engine or computer engine or application logic.
- the presentation engine receives the bitstreams that contains information regarding the update conditions, in response to a user triggering, animation of the door will then be performed to provide the “gap portion” to the Tenderer based on the update conditions.
- the presentation engine may also be used even without taking into account the bitstreams. Accordingly, one can just provide initialization of the scene data by means of e.g. a simple interface and then can update or change the audio scene according to the operation logic of e.g. the representation engine (for example, based on the property, rules, coordinates, conditions, etc. as defined in the logic), thereby simplifying update processing for a three-dimensional audio scene. Similar interface may also be used when taking into account a bitstream coming from a remote engine and stored locally.
- the update conditions contained in the bitstream may include a number of different scene states/definitions at different time instants during the opening/closing process.
- the update conditions contained in the bitstream may further include a time variable that indicates when the opening/closing process shall be carried out, e.g. the duration after which the process shall start or whether the process shall start immediately in response to a user triggering.
- a time variable that indicates when the opening/closing process shall be carried out, e.g. the duration after which the process shall start or whether the process shall start immediately in response to a user triggering.
- the bitstream may provide information that the material property of the voxels between the two points (voxels) defining the door will be changed to the material of the air for opening the door, as well as the time information indicating the duration one has to wait for opening the door.
- the same voxel points may be applied but the material property between those points shall be replaced with the door material, and the time variable shall indicate when to have the effect of closing door. For example, if the time variable indicates a duration of 1 second, the update of the material between those two points will take place after 1 second, i.e. animation of door closing will take place in 1 second after the user triggering the door closing process.
- the user triggering is externally generated (e.g. from the user) and is not included in the bitstream. However, the user triggering may be sent directly to the rendering engine.
- the bitstream to be transferred to the Tenderer may contain different scene states (e.g. information on different states the sliding door may have), the update conditions and time information indicating the rendering engine of when the update shall be carried out.
- the engine may be provided with some interface for receiving the update conditions from the bitstream as well as the external triggering.
- the Tenderer may be provided with bitstreams coming from the content creator side and interaction (triggers) coming from the user, and the rendering may be caused not directly by the bitstreams but by the interface defining the update conditions (e.g. by a short delay) in response to the external triggering.
- Fig. 5(a) shows an example flowchart of a method 500a of updating an audio scene for three- dimensional audio rendering according to embodiments of the disclosure.
- Method 500a may be implemented in software, hardware, or combinations thereof at an encoder for providing an updated representation of the audio scene (i.e. an updated audio scene) to a decoder for rendering the updated representation of the audio scene.
- Method 500a comprises steps S510 through S530 that may be performed, for example, by an encoder.
- these steps may be performed, for example, when an update of the audio scene takes place, i.e. when the scene state (or scene definition) of the audio scene changes.
- the scene state may relate to or comprise the listener location and the audio scene description (including the representation of the three-dimensional audio scene and the source location), and a change of the scene state could relate to one or more of a change of the listener location, a change of the source location, and a change of the (representation of the) three-dimensional audio scene.
- steps S510 through S530 may be performed for each of a plurality of processing cycles of an encoder and do not need to be performed in the order shown in Fig. 5(a).
- a voxelized representation of the audio scene is obtained.
- the voxelized representation comprises a set of voxels that forms a connected geometric region on a voxel grid of the voxelized representation.
- the voxelized representation of the audio scene may be similar to the examples shown in Fig. 2 where the geometric region has a cuboid shape and the set of voxels comprises at least a first boundary voxel and a second boundary voxel defining the cuboid shape of the geometric region.
- the voxels in the geometric region may share a common voxel property.
- an updated pair of boundary voxels defining an updated region and/or an updated voxel property for the voxels in the updated region is determined.
- the updated pair of boundary voxels defines an updated cuboid region for the geometric region, and the updated voxel property is associated with the voxels in the updated cuboid region. It is noted that at this step, it may be sufficient to determine one of the updated pair boundary voxels and the updated voxel property for updating the audio scene. However, both the updated pair boundary voxels and the updated voxel property may be determined to provide the user better perceptual effect in response to the scene update.
- a representation of the (updated) audio scene is generated based on the set of voxels with the updated pair of boundary voxels and/or the updated voxel property. Subsequently, the generated representation of the updated audio scene may be provided to a decoder for rendering the updated representation of the audio scene. It is further appreciated that the proposed updating method may generate a plurality of representations of (updated) audio scenes for a plurality of scene states (definitions) which will be delivered (via e.g. bitstreams) to the decoder for rendering complex dynamic scenes efficiently.
- a similar method 500b may be implemented at the decoder side for rendering an (updated) audio scene to be output to a Tenderer.
- the method 500b may provide a function of decompressing a compressed voxel-based (updated) audio scene from a received bitstream.
- the method 500b comprises steps S540 through S570 that may be performed, for example, in response to a user triggering or interaction indicating of an update of the audio scene. Similar to method 500a for the encoder, steps S540 through S570 may be performed for each of a plurality of processing cycles of a decoder and do not need to be performed in the order shown in Fig. 5(b).
- a bitstream comprising a voxelized representation of the audio scene is received.
- the voxelized representation comprises a plurality of voxels arranged in a voxel grid, and each voxel has an associated voxel property.
- the method 550b further comprises step S550 of decoding a set of voxels that form a connected geometric region on the voxel grid and an indication that the voxels in the geometric region share a first voxel property.
- the method 550b also comprises decoding (S560) a subset of the set of voxels associated with a scene sub-element within the geometric region and an indication that the subset of the set of voxels is assigned with a second voxel property.
- the method 550b comprises generating (S570) an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels, involving overwriting the voxels of the subset with the second voxel property.
- the updated representation of the audio scene may be output to a Tenderer for rendering the updated representation of the audio scene, allowing a simple and efficient scene update mechanism during a change of a three-dimensional audio scene.
- the decompressed voxel matrices may be additionally post-processed to enhance the representation capability.
- 3D filters e.g. gaussian filters
- the filter(s) By means of averaging provided by the filter(s), bulky objects/elements in the scene can become smoother and closer to a more realistic acoustic environment (e.g. to look or sound more nature).
- the block element may represent a tree crown (601a, 601b) as the leaf portion of the tree.
- the proposed method may also be applied to block elements of other kinds of shape within an audio scene.
- Fig. 6(a) shows the representation of the voxel block element after decompression but before applying the 3D filter.
- the voxel block element may have a rectangular shape (601a) for the tree crown which is defined by a cube/block assigned to the transmission properties of tree leaves. For example, two points (voxels) may be used to code the cube of voxels representing the tree crown with a reference to a material describing occlusion properties of its leaves.
- Fig. 6(b) shows the representation of the voxel block element after applying the 3D filter(s).
- Application of a 3D smoothing filter can transform the rectangular shape (601a) of the tree crown into an object with smoother shape (601b) to achieve a better correspondence to the intended geometry without big metadata overhead.
- the 3D filter(s) may be predefined or transmitted via the bitstream.
- a 3D smoothing filter may comprise a matrix (of a small size) depending on the resolution and the number of times to be applied. In an example, a 3x3 matrix applied three times may be sufficient for a tree around 10-meter away.
- such a smoothing (e.g. averaging) filter may be applied to the corresponding occlusion coefficients (during the direct sound occlusion modelling) to, for example, make the periphery of the tree crown more acoustically transparent than its core.
- such 3D filter(s) may be referenced to the data transmitted in the bitstream or hard-coded in the Tenderer.
- the 3D filter(s) may be associated with a filter identifier for specifying the filter to be applied and also how many times said filter shall be applied.
- the filtering can also be changed depending on the user position. For example, if the user approaches the object (e.g. the tree) from a far-away position, the filtering (e.g. the smoother or average in this case) may be activated or changed for better rendering the scene accordingly. Also, different filtering shapes may be defined and selected for improving the user perception in the dynamic audio environment.
- the filtered scene description may be delivered to the Tenderer and may not be changed.
- the filtering may be applied to the decompressed voxel matrices before applying the rendering matrix. Accordingly, the filtering operation may be separated from the scene description received in the bitstream to enhance application flexibility. For example, different filtering may be applied to the same (e.g. original) scene description which can be repeatedly retrieved from the bitstream, based on the user movement in the audio scene.
- Fig- 7 shows an example flowchart of a method 700 of processing an audio scene for three- dimensional audio rendering according to embodiments of the disclosure.
- Method 700 comprising steps S710 through S730 may be implemented at a decoder for providing an improved representation of the audio scene to the Tenderer.
- a voxelized representation of the audio scene is received.
- the voxelized representation comprises a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation.
- the received voxelized representation of the audio scene is decoded.
- at least one 3D smoothing filter is applied to the decoded voxelized representation of the audio scene.
- Such an appropriate 2D projection plane may be adapted to the location for the sound simulation, such as the use of a floor plan (e.g. horizontal) for the indoor scenario, and the use of a second (e.g., vertical) 2D projection plane for the outdoor scenario to account for the diffraction paths going over sound obstacle(s) or occluding structure(s). That is, the same path finding approach (algorithm) may be adapted to different projection planes, taking into account an additional path for more accurate diffraction modelling.
- a horizontal projection of a three-dimensional audio scene has been applied for calculating a virtual source position of a virtual sound source. Since humans are in general more sensitive to the horizontal dimension (i.e. azimuth) in audio spatial localization than to the vertical dimension, the resolution of the vertical axis is lower than the resolution of the horizontal axis when applying the corresponding 2D projection planes for diffraction modelling. For indoor scenes, it may be sufficient to apply only the horizontal plane or the vertical plane with lower resolution for calculating the diffraction path. For outdoor scenes, however, the present disclosure proposes to apply additionally the vertical projection plane (with better resolution) to improve the representation for the outdoor scenes.
- the same path finding algorithms may be used for the vertical analysis where the scene geometry may be divided over the vertical axis for calculating the diffraction paths on the vertical axis.
- the same path finding algorithms may be used for the vertical analysis where the scene geometry may be divided over the vertical axis for calculating the diffraction paths on the vertical axis.
- the application of the horizontal and vertical projection as the 2D projection planes for diffraction modelling, two or more virtual audio objects may be obtained which can be found on the horizontal and vertical planes, respectively.
- Fig. 8 schematically illustrates two examples of 2D projection planes of a 3D voxel-based scene representation for diffraction modelling according to embodiments of the disclosure.
- the audio scene may comprise a house, for example, as shown in Fig. 8(a), which depicts an occluding structure for the outdoor scenario.
- Fig. 8(b) and Fig. 8(c) show the corresponding vertical and horizontal planes, respectively, as the 2D projection planes for diffraction modelling, where two respective virtual audio sources (VS_v, VS_h) and their positions can be calculated based on the information on the original audio source (S) and the listener position (L).
- the 2D projection plane(s) can be calculated at the encoder side, which will then be transmitted in the bitstream, or at the Tenderer side (e.g. by a “default” projection plane calculation method).
- filtering or “matrix slicing cut” can be applied to the 3D voxel matrix taking in account the user and/or audio source positions.
- the obtained diffraction path(s) may be used by taking into account the diffraction coefficients.
- the same (or similar) tools and interfaces as for reflection modelling may be applied, assuming to use a virtual source rendering method. It is noted that while the reflection coefficients are used for reflection modelling, modelling the diffraction effect shall be based on the diffraction coefficients.
- the diffraction coefficients may be obtained at the encoder where a set of diffraction coefficients (similar to those of reflection coefficients) is assigned to each voxel based on the occluding object geometry and its material properties.
- the diffraction coefficients may also be obtained at the Tenderer where a sub-set of voxels which causes diffracted sound direction change(s) is selected based on the calculated diffraction path.
- the EQ coefficients of the diffraction audio signal may be further calculated based on the obtained subset of voxels (and the corresponding assigned diffraction coefficients).
- the EQ filters for the diffracted audio signal can be calculated based on a direct distance between the audio object and the listener.
- the calculation of the EQ filters for the diffracted audio signal can also depend on a diffracted path length (or its approximation), the diffraction coefficients of the corresponding voxels (or projection plane matrix elements), and/or the direction change angle(s) of the diffraction path.
- Fig- 9 shows an example flowchart of a method 900 of processing an audio scene for three- dimensional audio rendering according to embodiments of the disclosure.
- Method 900 may be implemented in e.g. the diffraction calculation block 40 and/or the diffraction modeling tool 50 within the processing chain 100.
- the method 900 comprises steps S910 through S940 for providing auralized audio data (e.g. 3DoF auralizer data) to be further processed by other rendering stages.
- auralized audio data e.g. 3DoF auralizer data
- a scene configuration indicative of the audio scene is obtained.
- the scene configuration includes a source position of an audio source, a given user position, and a scene description including a voxel matrix for a voxelized representation of the audio scene, and associated occlusion and diffraction coefficients.
- the method 900 further comprises step S920 which obtains a two-dimensional projection map related to the voxelized representation of the audio scene.
- the method 900 further comprises step S930 that determines a diffraction path between the audio source and the given user position based on the 2D projection map.
- the method 900 further comprises step S940 that obtaining auralization data for rendering the audio scene based on a result of the determination.
- the diffraction information e.g. diffraction path information relating to an acoustic diffraction path within the audio scene between the source location and the listener location
- the results of the diffraction modelling e.g. auralized audio data
- a pre-computed result i.e. the diffraction information and the results of the diffraction modelling which have been obtained at a previous time may be re-used at a later time.
- Such a pre-computed result may be obtained from the bitstream, and the use of the pre-computed diffraction data can avoid complex calculations at the decoder side.
- Fig. 10(a) schematically illustrates an example of a possible use case for reuse of pre-computed diffraction data according to embodiments of the disclosure.
- Users A and B may be users that individually or jointly explore a VR environment including the audio scene, for example as part of a game, virtual tour, etc..
- the users exploring a common VR environment may be running a social VR, for example. Having different listener locations within the audio scene, users A and B will produce different rendering results and different diffraction information.
- each user makes their calculated diffraction information available to other users.
- user B Once user B enters an area of the audio scene in which user A had been present previously, they may benefit from user A’s precomputed diffraction information, and vice versa.
- user A’s diffraction information may be made available to user B via a LUT that indexes different items of diffraction information with corresponding scene states or scene state identifiers.
- computational load for both users’ devices/decoders/renderers can be reduced, depending on their movement patterns within the audio scene.
- diffraction information (diffraction data) is accumulated in particular for relevant (e.g., frequently occurring) scene states. This would be very difficult to achieve for encoder-side precomputation of diffraction information since the encoder does not have access to the actual listener positions and therefore can only assume them. Further, use of data storage (e.g., physical/shared storage or bitstream bandwidth) would be much more inefficient for encoder-side precomputation, due to part of the precomputed diffraction information relating to irrelevant or less relevant scene states in this case.
- data storage e.g., physical/shared storage or bitstream bandwidth
- the proposed functionality and techniques can create LUTs that correspond to the real 6D0F behavior of users (and not an assumed one at the encoder side), and thus may be said to relate to smart user-oriented LUT creation.
- Fig. 10(b) is a diagram illustrating complexity measures for different implementations of processing audio scene information or audio rendering as functions of time, assuming a simple maze as the audio scene. It is further assumed that the user randomly moves through the maze, thus revisiting previously visited locations.
- Graph 810 relates to the case that no pre-computed diffraction information whatsoever is available (e.g., no diffraction information provided with the bitstream, memory/cache disabled). In this case, the computational load on the Tenderer is substantially constant and comparatively high.
- Graph 820 relates to the case that pre-computed diffraction information is locally available (e.g., no diffraction information provided with the bitstream, local memory/cache enabled).
- Graph 830 finally relates to the case that pre-computed diffraction information is externally provided (e.g., full diffraction information provided with the bitstream).
- the computation load on the Tenderer is constantly low, as a significant portion of scene states relates to known scene states and the diffraction information can be externally retrieved (e.g., from the bitstream or by request from an external/shared storage), without local calculation.
- the above-described diffraction modelling method may be implemented both on the mesh-based and voxel-based audio scene representation.
- the diffraction modeling method working on the voxel-based audio scene representation is computationally advantageous and can be used as an alternative method (e.g. running instead of the mesh- based one) for the low complexity/ bitrate rendering mode.
- this can also be used as an axillary method (e.g. running in parallel to the mesh-based one) for the cases when the mesh-based diffraction modeling cannot reliably compute (in real-time) diffraction data due to complexity-related reasons (e.g.
- Fig. 11 illustrates a non-limiting example of rendering an indoor audio scene using the above-indicated approaches, i.e. reuse of the pre-computed diffraction data, mesh-based diffraction modelling and voxel-based diffraction modelling, according to embodiments of the disclosure.
- Three respective virtual audio source objects (of diffraction modelling related data) obtained by the corresponding approaches are indicated by the point A for pre-computed data, point B for mesh-based data and point C for voxel-based data.
- the diffraction data may be pre-computed along the envisioned user path P between SI and S2 (and stored in the bitstream).
- the diffraction data may be computed (in real-time) using the mesh-based diffraction modeling method (around the sound sources, region S) and/or the voxel-based diffraction modeling method (in the rest of the scene, region R).
- two audio objects SI and S2 e.g. the TV set located in the living room 1 and the computer located in the sleeping room 2 are shown as original audio sources which emit sound (i.e. audio signals).
- the user may move, for example, from the living room 1 to the sleeping room 2, or vice versa.
- the scene description of this audio environment for audio rendering may be created by dividing the whole environment into a plurality of regions depending on the moving path of the user. For example, it can be divided into the region of user path P itself, the region S where the user can easily reach during his movement (e.g. the probable area where the user can be), and the region R being the rest of the scene.
- the diffraction data required for the rendering needs not be calculated in real time but can simply be obtained from the bitstream (i.e. the previous results such as diffraction paths).
- object coordinates and gains may also be obtained, so that the rendering tool can be skipped partially or even completely.
- the predefined user path P may relate to a known scene state for which the rendering can be carried out by externally providing the precomputed diffraction information (via the bitstream) without local calculation by the rendering tool (i.e. reduction of the computation load on the Tenderer), as explained above.
- the user is located within the region of the predefined path P (i.e. the scene state is known)
- it is just requested to retrieve the diffraction data externally (e.g. from the bitstream or by request from an external/shared storage) while skipping the computation of the rendering tools, which provides high quality rendering with minimum computational efforts.
- the reuse of the pre-computed data may require higher bandwidth for the transmission.
- the mesh-based diffraction modeling method may be applied, namely, the rendering of a mesh representation of the geometry in an audio scene.
- this region may be limited by the order of the diffraction modeling e.g. up to 3 rd diffraction.
- the calculated path may go over the obstacle of the corner, and above the limitation it may become cost-inefficient for calculation in real time.
- Such a region may be determined by the scene creator and may be based on a trade-off between the computational efforts and the size of a scene element, i.e. less computation is required for a smaller size.
- the mesh-based diffraction modelling approach can also provide high quality rendering, while requiring complex local computation by the rendering tools. On the other hand, such an approach does not need high bandwidth for the transmission (i.e. increased bandwidth efficiency).
- the voxel-based diffraction modeling approach may be applied, since voxels are not limited by the diffraction orders, which makes it easy to calculate the sound far away with sufficient accuracy (e.g. not necessarily for providing very detailed diffraction information such as including the vertical analysis as indicated above, but at least the modelling results based on the azimuth/horizontal plane are very representative).
- the voxel-based diffraction modelling approach allows to simplify the computation by the rendering tools compared to the mesh-based diffraction modelling, and can also increase the bandwidth efficiency for transmission compared to the reuse of the pre-computed data, as in this case the diffraction information is not included in the bitstream but will be calculated in real-time at the decoder/renderer side.
- the application of the voxel-based approach can still ensure the rendering of the audio scene with sufficient accuracy.
- the above indicated approaches may be combined with each other for rendering a complete audio scene.
- any combination of these approaches may be selected/determined to be applied in parallel, or alternatives, for rendering the same scene.
- the same scene can be rendered at different devices with various hardware/software requirements.
- the voxel-based modelling may be used for a device of low power usage
- the mesh-based modelling may be used for a device of higher power usage.
- a device may be provided with modes of reuse of precomputed data, of low complexity (power/computation), and/or of low bit rates (reduced transmission bandwidth). These different modes may be further assigned to different parts of the scene, as explained above. It is also noted that, for the transition areas, interpolation may be applied, or specific regions may be defined to coincide the nature (e.g. the obstacles and the boundaries) to ensure to smooth the rendering of the scene. Different implementations may be dependent on changes of the environment.
- Fig. 12 shows an example flowchart of a method 1200 of processing an audio scene at a rendering device for three-dimensional audio rendering according to embodiments of the disclosure.
- Method 1200 may be implemented in software, hardware, or combinations thereof e.g. at a render er/decoder coupled to AR/VR/MR/XR equipment as illustrated by the processing chain 100, or more specifically, e.g. by the diffraction calculation block 40 and/or the diffraction modeling tool 50 within the processing chain 100.
- the method 1200 comprises steps S1210 through SI 230 for providing auralized audio data (e.g. 3DoF auralizer data) to be further processed by other rendering stages.
- auralized audio data e.g. 3DoF auralizer data
- the method 1200 comprises step S 1210 of receiving a three-dimensional audio scene and information on a sound source at a source position.
- the method further comprises step SI 220 of determining a rendering mode of the rendering device.
- the method comprises step SI 230 of determining, for a given listener position, a virtual sound source at a virtual source position based on the source position to simulate an impact of acoustic diffraction by the three- dimensional audio scene on a source signal of the sound source at the source position.
- the determination of a virtual sound source is based on one or more of: pre-computed data along a predefined user path; and • real-time data computed using at least one of mesh-based diffraction modeling and voxelbased diffraction modeling, depending on the given listener position and/or the determined rendering mode of the rendering device.
- the rendering mode comprises a complexity-reduced mode and/or a bandwidth-efficient mode.
- the rendering mode may also include a high-quality rendering mode.
- the virtual sound source may be determined based on the pre-computed data along the predefined user path, or diffraction modeling may be applied to the three-dimensional audio scene using the precomputed data along the predefined user path.
- the virtual sound source may be determined based on the real-time data computed using the voxel-based diffraction modeling, depending on the given listener position.
- the virtual sound source may be determined based on the real-time data computed using the mesh- based diffraction modeling and/or the voxel-based diffraction modeling, depending on the given listener position.
- the virtual sound source in response to determining that a scene state of the three-dimensional audio scene is a known state (e.g. within the predefined path P as indicated in Fig. 11), the virtual sound source may be determined based on the pre-computed data along the predefined user path.
- the determination of the virtual sound source using the pre-computed data or the real-time computation may be based on the area/region to be rendered within the scene (i.e. the area/region where the listener is located). For certain areas where the diffraction modelling order is high, e.g. higher than 3, the voxel-based diffraction modelling may be used for determining the virtual sound source. Otherwise, the mesh-based diffraction modelling may be applied for the real-time computation of the virtual sound source.
- the following efficient representation of voxel indices may be applicable for transmission or storage of both voxel grid and Diffraction Map (VoxDataDiffractionMap) entries, for example. It may substitute any fixed-length representations of voxel indices (voxel coordinates).
- Step 1 Determine the amount (i.e., number, count) of bits needed for the current grid resolution / diffraction map dimension.
- these numbers NbitsVox and NbitsMap, respectively, may be determined for example as follows:
- NbitsVox ceil(log2(L*W*H-l))
- NbitsMap ceil(log2(L*W-l)), where L, W, H (Length, Width, Height) is the dimension of voxel grid and diffraction map. The values may differ for the voxel grid and diffraction map.
- Step 1 may apply to both the encoder side and the decoder side.
- Step 2 the voxel indices (x, y, z) and diffraction map indices (x, y) are mapped onto a packed representation index (Idx) and encoded using Nbits vox and Nbits map bits, respectively.
- Idx packed representation index
- (x, y, z) is zero-based and the packed representation indices may range from 0 to L*W*H-1 for voxels and from 0 to L*W-1 for the diffraction map.
- the mapping from the indices (x, y, z) onto the packed representation indices may be for example as follows:
- Idx(x, y, z) (x-1) + ((y-1) * L) + ((z-1) * L * W).
- Step 2 may be performed at the encoder side only.
- the packed representation index is an index that can uniquely identify a voxel in the voxel grid or diffraction map.
- the voxels in the voxel grid may have assigned thereto a unique consecutive index, so that each voxel in the voxel grid can be uniquely identified by a single integer number.
- a packed representation index may be used for any indication of a voxel location in the voxel grid or in a two-dimensional map.
- the packed representation index may be used for indicating any voxel locations mentioned throughout the disclosure.
- the assignment of unique indices to the voxels may be according to a predefined pattern. For example, the voxel grid may be scanned/traversed in x, y, and z directions, in this order, for consecutively assigning the unique index to respective voxels.
- VoxDataDiffractionMap[x][y] voxDiffractionMapValuefi]
- VoxDataDiffractionMap[x][y] VoxDataMatrix[x][y][H] where voxDiffractionMapPosPackedS and voxDiffractionMapPosPackedE indicated packed representation indices.
- An entropy coding method can be applied to the sequence of integer numbers representing acoustic properties, voxel grid coordinates, voxel grid indices, etc.
- an entropy encoding method can be applied to the sequence of integer numbers (representing acoustic property or audio rendering instruction set reference P_ID and grid indices XI , Y1 , Z 1 and X2, Y2, Z2) described above, or to packed representations thereof. Further, entropy coding may be applied to a sequence if integer numbers derived from the aforementioned representation of diffraction path information.
- the apparatus 1300 comprises a processor 1301 and a memory 1302 coupled to the processor 1301.
- the memory 1302 may store instructions for execution by the processor 1301.
- the processor 1301 may be adapted to implement the processing chains described throughout the disclosure and/or to perform methods (e.g., methods of processing audio scene information for audio rendering) described throughout the disclosure.
- the apparatus 1300 may receive inputs (e.g., audio scene description, listener location, etc.) and generate outputs (e.g., representations of diffraction information, etc.).
- aspects of the systems described herein may be implemented in an appropriate computer-based sound processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files.
- Portions of these systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers.
- Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
- WAN Wide Area Network
- LAN Local Area Network
- One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and/or as data and/or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and/or other characteristics.
- Computer-readable media in which such formatted data and/or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
- embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware.
- the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and/or application specific integrated circuits (“ASICs”).
- ASICs application specific integrated circuits
- computer-implemented neural networks described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input/output interfaces, and various connections (e.g., a system bus) connecting the various components.
- connections e.g., a system bus
- a method of updating an audio scene for three-dimensional audio rendering comprising: obtaining a voxelized representation of the audio scene, the voxelized representation comprising a set of voxels that forms a connected geometric region on a voxel grid of the voxelized representation, wherein the geometric region has a cuboid shape and the set of voxels comprises at least a first boundary voxel and a second boundary voxel defining the cuboid shape of the geometric region, wherein the voxels in the geometric region share a common voxel property; in response to an update of the audio scene, determining an updated pair of boundary voxels defining an updated cuboid region for the geometric region, and/or determining an updated voxel property for the voxels in the cuboid region; and generating a representation of the audio scene based on the set of voxels with the updated pair of boundary vox
- EEE2 The method according to EEE1, the method further comprising determining, from a plurality of voxels of the voxelized representation, at least the first boundary voxel and the second boundary voxel for the set of voxels.
- EEE3 The method according to EEE1 or EEE2, wherein the boundary voxels describe a scene element being a part to be updated within the audio scene, and wherein existing data relating to the voxel property for the scene element is to be overwritten depending on the updated pair of boundary voxels and/or the updated voxel property.
- EEE4 The method according to any one of EEE1 to EEE3, wherein the update of the audio scene is based on an update condition comprising one or more of: time, user input, input from a presentation engine, or input from an application logics.
- EEE5. The method according to EEE4, wherein an update attribute indicating the update condition is included as metadata in a bitstream along with a compressed representation of the audio scene based on the determined set of voxels.
- EEE6 The method according to any one of EEE1 to EEE5, wherein the update of the audio scene is further based on a trigger received from a user in real time.
- An encoder comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the encoder to carry out the method according to any one of EEE1 to EEE6.
- a method of decompressing a compressed voxel-based audio scene from a bitstream comprising: receiving the bitstream comprising a voxelized representation of the audio scene, the voxelized representation comprising a plurality of voxels arranged in a voxel grid, each voxel having an associated voxel property; decoding a set of voxels that form a connected geometric region on the voxel grid and an indication that the voxels in the geometric region share a first voxel property; decoding a subset of the set of voxels associated with a scene sub-element within the geometric region and an indication that the subset of the set of voxels is assigned with a second voxel property; and generating an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels, involving overwriting the voxels of the subset with the second
- a method of processing an audio scene for three-dimensional audio rendering comprising: receiving a voxelized representation of the audio scene, the voxelized representation comprising a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation; decoding the received voxelized representation of the audio scene; and applying at least one 3D smoothing filter to the decoded voxelized representation of the audio scene.
- EEE10 The method according to EEE9, wherein the voxels in the geometric region share a common voxel property, the common voxel property comprising a reference indicative of a material describing an occlusion property associated with the set of voxels for the geometric region.
- EEE11 The method according to EEE10, further comprising applying the at least one 3D smoothing filter to one or more occlusion coefficients indicative of the occlusion property.
- EEE 12 The method according to any one of EEE9 to EEE11, wherein the at least one 3D smoothing filter is defined based on data transmitted in a bitstream or is hard-coded in a renderer.
- EEE13 The method according to any one of EEE9 to EEE 12, wherein the at least one 3D smoothing filter is applied to the decoded set of voxels of the voxelized representation of the audio scene, and/or wherein the at least one 3D smoothing filter is associated with a scene element identifier for a geometric region within the audio scene, and wherein the at least one 3D smoothing filter is applied to the geometric region having that scene element identifier.
- EEE 14 The method according to any one of EEE9 to EEE13, wherein the at least one 3D smoothing filter is applied depending on a user position.
- EEE15 A non- transitory computer program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE14.
- EEE16. A method of processing an audio scene for three-dimensional audio rendering, wherein the method comprises: obtaining a scene configuration indicative of the audio scene, the scene configuration including a source position of an audio source, a given user position, and a scene description, the scene description including a voxel matrix for a voxelized representation of the audio scene, and associated occlusion and diffraction coefficients; obtaining a two-dimensional (2D) projection map related to the voxelized representation of the audio scene; determining a diffraction path between the audio source and the given user position based on the 2D projection map; and obtaining auralization data for rendering the audio scene based on a result of the determination.
- 2D two-dimensional
- EEE17 The method of EEE16, wherein the two-dimensional projection map is obtained by applying a projection operation to the voxelized representation of the audio scene, and wherein the diffraction path is determined based on one or more 2D path search algorithms using the 2D projection map for the voxelized representation of the audio scene.
- EEE18 The method of EEE16 or EEE17, wherein the 2D projection map is either received from a bitstream transmitted by an encoder or calculated at a Tenderer.
- EEE19 The method of any one of EEE16 to EEE18, wherein the 2D projection map is calculated at a Tenderer by applying filtering to the voxel matrix of the scene description based on the source position and/or the given user position.
- EEE20 The method of any one of EEE 16 to EEE 19, wherein the 2D projection map is obtained by selecting among at least one of: a horizontal projection relating to a diffraction path and a vertical projection relating to the diffraction path, the horizontal projection related to the voxelized representation by a horizontal projection operation that projects onto a horizontal plane, the vertical projection related to the voxelized representation by a vertical projection operation that projects onto a vertical plane.
- EEE21 The method of EEE20, wherein the selection is based on a route and/or a direction of the diffraction path.
- EEE22 The method of any one of EEE 16 to EEE21, wherein, for determining the diffraction path, to each voxel of the voxel matrix a set of diffraction coefficients is assigned based on occluding object geometry and/or respective material properties.
- EEE23 The method according to any one of EEE16 to EEE22, further comprising selecting a subset of voxels in the voxel matrix based on the determined diffraction path, the subset of voxels causing one or more changes in a diffracted sound direction.
- EEE24 The method according to EEE23 depending on EEE22, further comprising calculating coefficients of an EQ filter for a diffracted audio signal based on the selected subset of voxels, the calculation of the coefficients of the EQ filter depending on at least one of a direct distance between an audio object and a listener, a length of the diffraction path, the diffraction coefficients of the corresponding voxels, and an angle of a change in a direction of the diffraction path.
- EEE25 The method according to any one of EEE16 to EEE24, further comprising applying diffraction modelling to the voxelized representation of the audio scene.
- a method of processing an audio scene at a rendering device for three-dimensional audio rendering comprising: receiving a three-dimensional audio scene and information on a sound source at a source position; determining a rendering mode of the rendering device; and determining, for a given listener position, a virtual sound source at a virtual source position based on the source position to simulate an impact of acoustic diffraction by the three- dimensional audio scene on a source signal of the sound source at the source position, wherein the determination of a virtual sound source is based on one or more of: pre-computed data along a predefined user path, and real-time data computed using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling, depending on the given listener position and/or the determined rendering mode of the rendering device.
- EEE27 The method according to EEE26, further comprising determining that a scene state of the three-dimensional audio scene is a known state and in response, determining the virtual sound source based on the pre-computed data along the predefined user path.
- EEE28 The method according to EEE26 or EEE27, wherein the rendering mode comprises a complexity-reduced mode and/or a bandwidth-efficient mode.
- EEE29 when it is determined that the rendering mode is a complexity- reduced mode, further comprising: determining the virtual sound source based on the pre-computed data along the predefined user path or applying diffraction modeling to the three-dimensional audio scene using the precomputed data along the predefined user path; and/or determining the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling, depending on the given listener position.
- EEE30 The method of EEE 28, when it is determined that the rendering mode is a bandwidthefficient mode, further comprising: determining the virtual sound source based on the real-time data computed using the mesh-based diffraction modeling and/or the voxel-based diffraction modeling, depending on the given listener position.
- EEE31 The method of any one of EEE26 to EEE30, further comprising: determining a diffraction modeling order indicative of a complexity of a diffraction path for the given listener position; and determining whether to use the mesh-based diffraction modeling or the voxel-based diffraction modeling based on the determined diffraction modeling order for the given listener position.
- EEE32 when it is determined that the diffraction modeling order for the given listener position is higher than a predefined order of diffraction, further comprising determining the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling.
- EEE33 The method of any one of EEE26 to EEE32, when it is determined that, for the given listener position, the pre-computed data along the predefined user path is not available or no realtime data computed using the mesh-based diffraction modeling is available, further comprising determining, at the given listener position, the virtual sound source based on the real-time data computed using the voxel-based diffraction modeling.
- EEE34 A non-transitory computer program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE26 to EEE33.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Stereophonic System (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Human Computer Interaction (AREA)
- General Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363485731P | 2023-02-17 | 2023-02-17 | |
| PCT/EP2024/053833 WO2024170671A2 (en) | 2023-02-17 | 2024-02-15 | Methods, apparatus, and systems for processing audio scenes for audio rendering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4666595A2 true EP4666595A2 (en) | 2025-12-24 |
Family
ID=90038330
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24706949.5A Pending EP4666595A2 (en) | 2023-02-17 | 2024-02-15 | Methods, apparatus, and systems for processing audio scenes for audio rendering |
Country Status (5)
| Country | Link |
|---|---|
| EP (1) | EP4666595A2 (en) |
| JP (1) | JP2026505661A (en) |
| KR (1) | KR20250151466A (en) |
| CN (1) | CN120712795A (en) |
| WO (1) | WO2024170671A2 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11146905B2 (en) * | 2017-09-29 | 2021-10-12 | Apple Inc. | 3D audio rendering using volumetric audio rendering and scripted audio level-of-detail |
| KR20220162718A (en) * | 2020-04-03 | 2022-12-08 | 돌비 인터네셔널 에이비 | Diffraction modeling based on grating path finding |
-
2024
- 2024-02-15 EP EP24706949.5A patent/EP4666595A2/en active Pending
- 2024-02-15 WO PCT/EP2024/053833 patent/WO2024170671A2/en not_active Ceased
- 2024-02-15 KR KR1020257030953A patent/KR20250151466A/en active Pending
- 2024-02-15 CN CN202480012768.2A patent/CN120712795A/en active Pending
- 2024-02-15 JP JP2025545796A patent/JP2026505661A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024170671A2 (en) | 2024-08-22 |
| CN120712795A (en) | 2025-09-26 |
| WO2024170671A3 (en) | 2024-10-03 |
| JP2026505661A (en) | 2026-02-17 |
| KR20250151466A (en) | 2025-10-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250203316A1 (en) | Methods, apparatus, and systems for processing audio scenes for audio rendering | |
| JP7467340B2 (en) | Method and system for handling local transitions between listening positions in a virtual reality environment - Patents.com | |
| KR20220162718A (en) | Diffraction modeling based on grating path finding | |
| KR101842411B1 (en) | System for adaptively streaming audio objects | |
| US12401963B2 (en) | Method and apparatus for fusion of virtual scene description and listener space description | |
| KR20200141438A (en) | Method, apparatus, and system for 6DoF audio rendering, and data representation and bitstream structure for 6DoF audio rendering | |
| KR102760260B1 (en) | Apparatus for Immersive Spatial Audio Modeling and Rendering | |
| CN118511547A (en) | Renderer, decoder, encoder, method and bit stream using spatially extended sound sources | |
| KR20240056791A (en) | Bitstream representing audio in an environment | |
| WO2024170671A2 (en) | Methods, apparatus, and systems for processing audio scenes for audio rendering | |
| HK40118970A (en) | Methods, apparatus, and systems for processing audio scenes for audio rendering | |
| JP2025517640A (en) | Method, apparatus and system for early reflection estimation for voxel-based geometry representations - Patents.com | |
| KR20250022845A (en) | Method, system and device for acoustic 3D range modeling on voxel-based geometric representations | |
| WO2024256238A1 (en) | Methods, apparatus, and systems for processing audio scene information | |
| RU2832227C1 (en) | Diffraction simulation based on finding path on grid | |
| EP4674140A1 (en) | Multi-directional audio diffraction modeling for voxel-based audio scene representations | |
| WO2025056788A1 (en) | Methods and apparatus for processing voxel-based scene representations | |
| TWI852358B (en) | Diffraction modelling based on grid pathfinding | |
| EP4539509A1 (en) | Generating an audio data signal | |
| HK40116205A (en) | Methods, apparatus, and systems for early reflection estimation for voxel-based geometry representation(s) | |
| HK40112770A (en) | Diffraction modelling based on grid pathfinding | |
| KR20240004337A (en) | Method, apparatus and system for modeling audio objects with range |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250903 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| REG | Reference to a national code |
Ref country code: HK Ref legal event code: DE Ref document number: 40128381 Country of ref document: HK |
|
| P01 | Opt-out of the competence of the unified patent court (upc) registered |
Free format text: CASE NUMBER: UPC_APP_0004279_4666595/2026 Effective date: 20260206 |