EP4635205A1 - Verfahren und vorrichtung zur effizienten audiowiedergabe - Google Patents
Verfahren und vorrichtung zur effizienten audiowiedergabeInfo
- Publication number
- EP4635205A1 EP4635205A1 EP23833615.0A EP23833615A EP4635205A1 EP 4635205 A1 EP4635205 A1 EP 4635205A1 EP 23833615 A EP23833615 A EP 23833615A EP 4635205 A1 EP4635205 A1 EP 4635205A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio sources
- audio
- sources
- rendering
- control value
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
Definitions
- the present disclosure relates generally to a method of rendering audio sources.
- the present disclosure relates to controlling the rendering of audio sources for efficient audio rendering.
- Audio rendering in particular of high quality immersive content, is computationally expensive. To render immersive audio, especially on power limited devices, within reasonable operation times, allows only very complexity constrained numerical operations on the processors included in them. Thus, the output audio is often degraded in quality and the listener experience is poor.
- a method of rendering audio sources may comprise determining, for each of a plurality of audio sources, a rendering control value based on one or more parameters indicative of a perceptual relevance of the respective audio source.
- the method may further comprise comparing, for each of the plurality of audio sources, the rendering control value to a perceptibility threshold to determine whether a respective audio source fulfills a perceptual irrelevance criterion.
- the method may further comprise modifying audio sources, out of the plurality of audio sources, for which the rendering control value satisfies the threshold to obtain modified audio sources.
- the method may comprise rendering unmodified audio sources, out of the plurality of audio sources, relative to the modified audio sources.
- the one or more parameters may include a loudness of a respective audio source.
- the rendering control value may then be determined based on a loudness value.
- the one or more parameters may include a relationship between the loudness of a single audio source and the loudness of audio output resulting from rendering the plurality of audio sources excluding said single audio source.
- the one or more parameters may include a relationship between a short-term spectral energy of a single audio source and a spectral energy of audio output resulting from rendering the plurality of audio sources.
- the rendered audio output may be associated with a psychoacousticmasking model.
- the modifying the audio sources for which the rendering control value satisfies the threshold may include culling the audio sources.
- the modified audio sources may then correspond to a culling sub-set of the plurality of audio sources.
- the rendering the unmodified audio sources, out of the plurality of audio sources, relative to the modified audio sources may include not to render the audio sources included in the culling sub-set.
- the one or more parameters may include a positional proximity of an audio source in respect to a listener position.
- the rendering control value may be determined based on a distance of a respective audio source position of the audio source for which the rendering control value is determined to another audio source, and wherein the rendering control value may be compared with a fraction of a distance between the listener position and a closest source position.
- the one or more parameters may include a directional proximity with respect to a listener position.
- the rendering control value may be determined based on at least one of azimuth or elevation between the respective audio source position of the audio source for which the rendering control value is determined and the listener position.
- the modifying the audio sources may include clustering the audio sources for which the rendering control value satisfies the threshold. The modified audio sources may then correspond to cluster audio sources.
- the rendering the unmodified audio sources, out of the plurality of audio sources, relative to the modified audio sources may include substituting the cluster audio sources by a smaller number of audio sources and rendering the smaller number of audio sources as the modified audio sources.
- the one or more parameters may include one or more of acoustic occlusion, optical occlusion, or abstraction.
- the rendering control value may then be determined based on a level of occlusion or abstraction of the respective audio source.
- the modifying the audio sources may include converting a type of the audio sources for which the rendering control value satisfies the threshold.
- the modified audio sources may then correspond to typecasted audio sources.
- the rendering the unmodified audio sources, out of the plurality of audio sources, relative to the modified audio sources may include rendering the typecasted audio sources as the modified audio sources.
- a directivity calculation may not be performed, if the rendering control value satisfies the threshold.
- the method may further include receiving control information indicative of whether to perform the modifying of some or all of the audio sources for which the rendering control value satisfies the threshold.
- control information may be one of activation or deactivation parameters, grouping data, prioritization information, and/or visibility information.
- an apparatus for rendering audio sources may include one or more processors configured to implement a method including: determining, for each of a plurality of audio sources, a rendering control value based on one or more parameters indicative of a perceptual relevance of the respective audio source; comparing, for each of the plurality of audio sources, the rendering control value to a perceptibility threshold to determine whether a respective audio source fulfills a perceptual irrelevance criterion; modifying audio sources, out of the plurality of audio sources, for which the rendering control value satisfies the threshold to obtain modified audio sources; and rendering unmodified audio sources, out of the plurality of audio sources, relative to the modified audio sources.
- an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method described herein.
- a computer-readable storage medium may store the program.
- FIG. 1 illustrates an example of a method of rendering audio sources according to an embodiment of the disclosure.
- FIG. 2 illustrates schematically an example of culling audio sources according to an embodiment of the disclosure.
- FIG. 3 illustrates schematically an example of clustering audio sources according to an embodiment of the disclosure.
- FIG. 4 illustrates schematically another example of clustering audio sources according to an embodiment of the disclosure.
- FIG. 5 illustrates schematically an example of typecasting audio sources according to an embodiment of the disclosure.
- FIG. 6 illustrates an example of a method of processing audio according to an embodiment of the disclosure.
- FIG. 7 illustrates another example of a method of processing audio according to an embodiment of the disclosure.
- FIG. 8 illustrates yet another example of a method of processing audio according to an embodiment of the disclosure.
- FIG. 9 illustrates an example of using control information according to an embodiment of the disclosure.
- FIG. 10 illustrates schematically an example of control information parameters according to an embodiment of the disclosure.
- FIG. 11 illustrates an example of an apparatus comprising a processor and a memory coupled to the processor according to an embodiment of the disclosure.
- connecting elements such as solid or dashed lines or arrows
- connecting elements such as solid or dashed lines or arrows
- the absence of any such connecting elements is not meant to imply that no connection, relationship, or association can exist.
- some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure.
- a single connecting element is used to represent multiple connections, relationships or associations between elements.
- a connecting element represents a communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or multiple signal paths, as may be needed, to affect the communication.
- Methods and apparatus as described herein provide the technical benefits and advantages of decreasing computational and memory system workload without perceptual quality degradation of the resulting rendered audio output.
- An absence of perceptual quality degradation may be understood as providing a similar (or better) listener subjective experience.
- the present disclosure provides for a method that modifies audio output to achieve this benefit, as opposed to un-modified audio output.
- the present disclosure provides the additional benefit of workload reduction. This reduction may be achieved by modifying and/or deactivating of selected audio source digital signal processing, DSP, instances and optionally associated metadata processing.
- Parameters, downmix matrices and/or related information for methods described herein can be defined and signaled by an encoder, decoder/renderer, and/or application.
- an encoder can operate in an “encoder-assisted” mode where the encoder can manually allow a content-creator to have an influence on an otherwise automatic method. This can provide the benefit of avoiding non-sensical culling as described below.
- this information can be provided in a decoder and/or Tenderer (e.g., automatically being provided in a “default” mode).
- the information can be provided by an application (e.g., in a “system-assisted” mode being provided automatically by application or system middleware).
- audio sources may be any type of audio input sources, such as, for example, channel (static) sources, audio object(s), Higher Order Ambisonic(s), HOA, First Order Ambisonics, FOA, and/or B-format ambisonics.
- channel (static) sources such as, for example, channel (static) sources, audio object(s), Higher Order Ambisonic(s), HOA, First Order Ambisonics, FOA, and/or B-format ambisonics.
- step S101 of the method for each of a plurality of audio sources, a rendering control value is determined based on one or more parameters indicative of a perceptual relevance of the respective audio source.
- step S102 of the method for each of the plurality of audio sources, the rendering control value is compared to a perceptibility threshold to determine whether a respective audio source fulfills a perceptual irrelevance criterion.
- step S103 audio sources, out of the plurality of audio sources, for which the rendering control value satisfies the threshold are modified to obtain modified audio sources.
- step SI 04 unmodified audio sources, out of the plurality of audio sources, are rendered relative to the modified audio sources.
- the one or more parameters may include a loudness of a respective (single) audio source.
- the rendering control value may then be determined based on a respective loudness value.
- loudness values may be obtained from:
- Metadata from existing audio standards set by the MPEG audio group including the MPEG-H 3D audio (ISO_IEC_23008-3) or MPEG-D part 4 (Dynamic Range Control).
- the metadata may be presented for a legacy content;
- the loudness information may be estimated and transmitted via an MPEG-H Audio Stream (MHAS) packet payload;
- MHAS MPEG-H Audio Stream
- Information by an MPEG-I Tenderer For example, information estimated in real time using the audio content and rendering gains.
- loudness data sources may be applied alone or in combination; the latter, for example, for content without or with non-reliable loudness information, social virtual reality, VR, content, etc.
- the loudness may be signaled and treated differently depending on different loudness definitions and measurement methods, for example, short term, momentary, level gated, IBU defined etc.
- the one or more parameters may include a relationship between the loudness of a single audio source and the loudness of audio output resulting from rendering the plurality of audio sources excluding said single audio source.
- a relationship may be, for example, in a non-limiting manner a signal to noise ratio.
- Other measures may also be conceivable as well using, for example, spatial masking.
- the one or more parameters may include a relationship between a short-term spectral energy of a single audio source and a spectral energy of audio output resulting from rendering the plurality of audio sources.
- the rendererd audio output may be associated with a psychoacoustic-masking model.
- the above described parameters may be used alone or in combination to determine the respective rendering control value for each of the plurality of audio sources.
- the rendering control value may then be said to be indicative of a perceptual relevance (perceptibility) of the respective audio source.
- the rendering control value to a (predetermined) perceptibility threshold may then indicate as to whether the respective audio source is perceptually relevant or irrelevant, that is, as to whether the respective audio source fulfills a perceptual irrelevance criterion. For example, if the rendering control value is equal to or above a certain threshold value, the respective audio source may be considered to be perceptually relevant. If the rendering control value is below a certain threshold value, the respective audio source may be considered to be perceptually irrelevant, that is, the respective audio source fulfills/satisfies the perceptual irrelevance criterion.
- an audio source associated to a bird singing audio may become perceptually irrelevant if a listener would start to experience heavy rain noise.
- the rendering control value would satisfy the respective (perceptual relevance) threshold, the respective audio source would fulfill the perceptual irrelevance criterion.
- the modifying of the audio sources for which the rendering control value satisfies the threshold may include culling the audio sources.
- the modified audio sources may then correspond to a culling sub-set of the plurality of audio sources.
- the culling of audio sources is illustrated schematically by forming a respective culling sub-set.
- Figure 2 illustrates a plurality of audio sources 201, 202 associated to a respective audio scene 200.
- the audio sources 202 may be associated to a bird singing
- the audio sources 201 may be associated with a light rain. If a listener now starts to experience heavy rain noise 201, the audio sources associated with the bird singing 202 may become perceptually irrelevant, i.e. the rendering control value may satisfy the perceptibility threshold.
- the audio sources 202 are modified in that the audio sources may be culled.
- the culled audio sources may then correspond to a respective culling sub-set 203.
- Culling in this context, may be said to refer to removing the respective perceptually irrelevant audio sources. That is, in an embodiment, the rendering the unmodified audio sources, out of the plurality of audio sources, (for which the rendering control value may not satisfy the perceptibility threshold) relative to the modified audio sources may include not to render the audio sources included in the culling sub-set.
- the one or more parameters may alternatively, or additionally, include a positional proximity (close source location) of an audio source in respect to a listener position.
- the rendering control value may then be determined based on a distance of a respective audio source position of the audio source for which the rendering control value is determined to another audio source, and the rendering control value may be compared with a fraction of a distance between the listener position and a closest source position. For example, a distance between N audio source locations may be smaller than Tdistance o of the distance between the listener and the closest source.
- Tdistance % refers to a variable denoting a percentage value, (e.g., 5%).
- the distance between the listener to a closest source can be calculated. This distance to this closest source can be referred to as a variable Tc. Additionally, the distance between each of the remaining audio sources in that scene can also be calculated. Each of these distances is compared against Tc. If a distance between two sources is less than Tdistance o of Tc, then these sources are clustered (grouped together). The cluster may add to or remove from it more members (audio sources) since the comparison is done on all audio sources.
- Figure 3 illustrates an example of abstract visualization of clustering based on positional proximity.
- cluster A 301 contains the audio sources 301 A closest to one another in terms of distance.
- Cluster B 302 contains different sources 302B that, likewise, are closest to one another in terms of distance.
- source X 303 is located further away from any audio source 301 A, 302B in both clusters A 301 and B 302, and hence is not assigned to either cluster.
- the one or more parameters may include a directional proximity (close source azimuth and elevation) with respect to a listener position.
- the rendering control value may then be determined based on at least one of azimuth or elevation between the respective audio source position of the audio source for which the rendering control value is determined and the listener position.
- a listener position related azimuth and elevation of N audio sources may be smaller than Tazimuth and Teievation degrees, respectively.
- Figure 4 illustrates an example of abstract visualization of clustering based on directional and/or positional/distance proximity.
- the above described parameters may be used alone or in combination to determine the respective rendering control value for each of the plurality of audio sources.
- the rendering control value may then be said to be indicative of a perceptual relevance of the respective audio source.
- the rendering control value to the (predetermined) perceptibility threshold may then indicate as to whether the respective audio source is perceptually relevant, that is, whether the respective audio source fulfills a perceptual irrelevance criterion. For example, if the rendering control value is equal to or above a certain threshold value, the respective audio source may be considered to be perceptually relevant.
- the respective audio source may be considered to be perceptually irrelevant, that is, the respective audio source fulfills/satisfies the perceptual irrelevance criterion.
- the above-described parameters refer to perceptual localization related aspects for the estimation of a perceptual rel evance/irrel evance .
- the modifying the audio sources may include clustering the audio sources for which the rendering control value satisfies the perceptibility threshold.
- the modified audio sources may then correspond to cluster audio sources.
- the rendering the unmodified audio sources, out of the plurality of audio sources, (for which the rendering control value may not satisfy the perceptibility threshold) relative to the modified audio sources may include substituting the cluster audio sources by a smaller number of audio sources and rendering the smaller number of audio sources as the modified audio sources.
- AT may be the number of modified audio sources to be rendered instead of N. Hence, AT audio sources still represent the same cluster.
- the one or more parameters may include one or more of acoustic occlusion, optical occlusion, or abstraction.
- the rendering control value may then be determined based on a level of occlusion or abstraction of the respective audio source.
- the modifying the audio sources may include converting a type of the audio sources for which the rendering control value satisfies the threshold. That is, modifying the audio sources may include converting a respective audio source from one type to another. The resulting type may be rendered more computationally efficient. This is schematically illustrated in the example of Figure 5.
- Figure 5 illustrates a plurality of audio sources 501, associated to a respective audio scene 500.
- those audio sources 501 can be substituted by one audio object 503 (e.g., distant waterfall point-source); or many audio point-sources (e.g., rain droplets) can be substituted by one Higher Order Ambisonics, HO A, (e.g., rain noise) signal.
- HO A Higher Order Ambisonics
- the modified audio sources may then correspond to typecasted audio sources.
- the parameters to determine the respective rendering control condition (rendering control value compared to the threshold) for typecasting may, in particular, dependent on the positional and/or directional proximity and/or the level of acoustic/optical occlusion/abstraction as described above. For example, if several audio sources are occluded and still acoustically important, then those sources can simply be replaced by a single point source audio or an FOA signal.
- the rendering the unmodified audio sources, out of the plurality of audio sources, (for which the rendering control value may not satisfy the perceptibility threshold) relative to the modified audio sources may then include rendering the typecasted audio sources as the modified audio sources.
- a directivity calculation may not be performed, if the rendering control value satisfies the threshold.
- An example may be a motor engine sound of a car that has a directivity. Assuming that a car is running on a circular track and a listener is standing on a particular location/ stage as in a racing car scenario, the effect of a motor engine sound with directivity will be perceived when the car is passing the listener by. But when the car is at a distant location, the directivity effect is neglectable as the listener only hears a faint sound of the motor engine, making it perceptually irrelevant, and therefore it is not necessary to perform the directivity calculation.
- the source may even be further considered to be included in the culling subset as well if there is a significant ob stacl e/occluder in the middle of the circular track such as a hill or a bunch of solid rocks.
- the directivity calculation may be skipped by additionally considering the so-called worst-case directivity gain estimate parameter.
- This worst-case estimate is straightforward to compute, e.g., by simply taking the maximum directivity gain within a given directivity pattern of an audio source. This estimate may be used in combination with the other existing rendering stages threshold, e.g., 60 dB culling threshold, so that if an audio source rendering gain satisfies the “Culling” stage gain threshold, the directivity calculation/stage may be skipped.
- the method may further include receiving control information indicative of whether to perform the modifying of some or all of the audio sources for which the rendering control value satisfies the threshold.
- (De-)activation parameters may be included in the control information. Such parameters may, for example, allow to disable the modifying for narrator speech signals, even if respective audio sources may have rendering control values that satisfy the threshold, i.e. fulfill the perceptual irrelevance criterion.
- grouping data for culling/clustering/typecasting may be included in the control information. Such grouping data may, for example, include means for defining a set of audio sources within a certain area (e.g., sound emitting parts of a car) or with certain signal categories (e.g., ambience sound of rain and wind).
- the encoder may define a culling and clustering strategy to either combine or remove objects which may be included in the control information. For example: o if all objects in a group are below a certain threshold loudness, the group may be combined into one object (loudness-based culling); o if azimuth/elevation of all objects are within a certain range, the group may be combined into one object (directional -based clustering); o if one or many objects of the group are acoustically occluded, the group may be combined into one object.
- a prioritization in respect to signal perceptual relevance (and its informational importance) to the listener (and listener’s focus of attention) may be included in the control information.
- culling of non-relevant audio sources may be applied (instead of their leveling/EQing) for the “cocktail party effect” reproduction to improve speech intelligibility.
- a visibility of a graphical representation of audio source(s) to the listener i.e., whether a visual representation of audio source is visible to the user or not
- an audio source may be associated with visual representation, which may be: - not in the video rendered viewport (behind the listener);
- Such audio sources may be treated differently depending on their nature and application scenario; e.g., thresholds for positional or directional proximity conditions can differ depending on a virtual object visibility.
- time thresholds may be determined based on measured velocities.
- parameters for all rendering stages can be defined per (virtual) object. All these rendering stages may be implemented as part of the MPEG-I rendering process or as the MPEG-I Audio “External Tools”. Transition from the “culling-”, “clustering-” and “typecasting-” state to the original one may be initiated (when the corresponding condition is met) and the sound source(s) energy is(are) low. That is, in a further embodiment, the above described method may be implemented by a respective two step process for determining the ‘right’ frame for applying the complexity reduction measure. This two step may include, in a first step, determining that a measure for complexitiy reduction (i.e. modification of audio sources as described) is to be applied, and in a second step, determining a point in time or time period for applying the respective measure.
- a measure for complexitiy reduction i.e. modification of audio sources as described
- the above described method may be implemented by a respective apparatus including one or more processors.
- the above described method may be implemented in the form of a respective program comprising instructions that, when executed by a processor, cause the processor to carry out the method.
- the program may be stored on a computer-readable storage medium.
- a method according to the present disclosure is directed towards the selection (i.e., “culling”) of such audio sources.
- the present disclosure is directed towards a method of culling of audio sources.
- the culling may be performed either during the core audio decoding stage or during the audio rendering stage.
- the “culling” method is directed towards removing perceptually irrelevant audio sources. For example, when a listener starts experiencing heavy rain noise, the culling process would remove an audio source associated to a bird singing audio from the rendering pipeline.
- existing techniques are not capable of detecting a subclass of perceptually irrelevant audio sources.
- the subclass contains sources which are inaudible, but having the corresponding distance (gains) values higher/lower than a culling threshold.
- the threshold values may be based on the distance and gain. For example, if the source distance is bigger than a distance threshold value D, the source is culled. Alternatively, if the source gain value is lower than a gain threshold G, the source is culled.
- some audio sources may end up to be categorized as “unculled” but they may be perceptually irrelevant or inaudible.
- the present disclosure is directed towards performing culling of audio sources based on additional information, such as information regarding perceptual energy /loudness-related aspects. This information allows for estimation of perceptual relevance and improvement in setting the “culling” application condition.
- the relevant information is the loudness of a single audio source.
- the loudness information i.e., values
- Metadata from existing audio standards set by the MPEG audio group including the MPEG-H 3D audio (ISO_IEC_23008-3) or MPEG-D part 4 (Dynamic Range Control).
- the metadata may be presented for a legacy content.
- the loudness information may be estimated and transmitted via an MPEG-H Audio Stream (MHAS) packet payload.
- MHAS MPEG-H Audio Stream
- Information by an MPEG-I Tenderer For example, information estimated in real time using the audio content and rendering gains.
- Each of these sources may be considered individually or in combination.
- combination of different loudness data sources (1), (2), (3) may be used for certain type of applications (e.g., for content without or with not-reliable loudness information, social VR audio content).
- Another aspect of the present disclosure considers information regarding a relationship between the loudness of a single audio source and the loudness of the overall rendered audio output. Some implementations may exclude this source, e.g., in case of a SNR-based relationship. Loudness information may be signaled and treated differently depending on the different loudness definitions and measurement methods (e.g., short term, momentary, EBU R128 defined).
- a further aspect of the present disclosure considers information regarding a relationship between the short term spectral energy of a single audio source and the spectral energy of the final rendered audio output associated with a psychoacoustic-masking model.
- Figure 6 illustrates an exemplary method of culling audio sources according to the present disclosure.
- the method includes a first step S601.
- step S601 for a plurality of audio sources, one or more values and/or relationship data related to the perceptually relevance information may be obtained.
- the audio sources may be received and/or predetermined.
- the value(s) indicate loudness of each of the audio sources.
- the relationship data can indicate a relationship between the loudness of a single audio source and the loudness of the overall rendered audio output (excluding the single source).
- the single audio source would be one of the plurality of audio sources.
- the relationship data would relate to a relationship between the short-term spectral energy of a single audio source and spectral energy of the final rendered audio output associated with a psychoacoustic-masking model.
- the single audio source would likewise be one of the plurality of audio sources.
- each of the values/data from S601 would be verified whether they satisfy the corresponding predefined condition threshold(s).
- An example will be an “SNR” value, where a single audio source is considered as the “useful Signal” and the rest of audio sources are considered “Noise”.
- the SNR value is then compared to a threshold.
- the threshold value may be preset or may be determined dynamically.
- An audio source having a value satisfying the corresponding condition threshold is considered as perceptually irrelevant.
- a subset of perceptually irrelevant sources is then chosen from the plurality of sources from S601.
- Step S603 audio culling is performed on the perceptually irrelevant audio sources from S602 that satisfied the condition threshold.
- Step S603 outputs information that relates to the state of an audio source (culled or unculled).
- the state of an audio source which can be “culled” or “unculled” can be expressed by a Boolean variable indicating “true” or “false”.
- the audio Tenderer uses the culling information in conjunction with the audio sources to determine which audio sources would be rendered.
- the present disclosure is further directed towards performing clustering of audio sources.
- the “clustering” stage should substitute several audio sources by a smaller number of modified sources.
- Clustering refers to grouping based on certain conditions, (e.g., distance proximity of audio sources, such as sources that are close to each other). The clustering may be performed either during the core audio decoding stage or during the audio rendering stage.
- the present disclosure is directed to clustering of audio sources that is based on additional perceptual localization-related aspects for estimation of perceptual relevance and setting the “clustering” application condition.
- the AT resulting audio sources should contain weighted downmixed signal(s) or multichannel audio signals.
- the Tdistance % refers to a variable denoting a percentage value, (e.g., 5%).
- Tc the distance between the listener to a closest source
- Tc the distance between each of the remaining audio sources in that scene
- Each of these distances is compared against Tc. If a distance between two sources is less than Tdistance% of Tc, then these sources are clustered (grouped together). The cluster may add to it more members (audio sources) since the comparison is done on all audio sources. Assuming that a total of N audio sources is being clustered, these sources are then substituted by AT audio sources. M is the number of audio sources to be rendered instead of N. Hence, AT audio sources still represent the same cluster.
- Figure 3 illustrates an example of abstract visualization of clustering based on positional proximity.
- cluster A 301 contains the audio sources 301 A closest to one another in terms of distance.
- Cluster B 302 contains different sources 302B that, likewise, are closest to one another in terms of distance.
- source X 303 is located further away from any audio source 301 A, 302B in both clusters A 301 and B 302, and hence is not assigned to either cluster.
- a further aspect of the present disclosure is determining clustering based on directional proximity (i.e., close source azimuth and elevation) in respect to a listener position.
- perceptual localization ability of the listener has finer resolution on azimuth rather than on elevation.
- the Tazimuth and Teievation are each threshold values in degrees, e.g., 5 degrees, for azimuth and elevation, respectively, representing the directional proximity.
- the angle/direction difference between any two sources as seen from a listener pose can be calculated (based on azimuth and elevation). If the angle/direction difference between two sources is less than Tazimuth and Teievation, then these sources are clustered (grouped together). The cluster may add to or remove from it more members (audio sources) since the comparison is done on all audio sources.
- Figure 4 illustrates an example of abstract visualization of clustering based on directional and/or positional/di stance proximity.
- any two sources e.g., A2 and An
- the Listener L is smaller than Tazimuth and Teievation degrees, respectively.
- These sources are then clustered together. It could also be assumed that since the distance between those audio sources (A, Al, A2, A3 and An) are so close to each other so they are clustered based on the positional or distance proximity.
- Figure 7 illustrates an exemplary method of clustering audio sources according to the present disclosure.
- the method includes a first step S701.
- step S701 for a plurality of audio sources, one or more perceptual relevance information values for each of the audio sources may be determined.
- the perceptual relevance information may be one of i) positional proximity and/or ii) directional proximity.
- the values from S701 are each verified to determine whether they satisfy a corresponding predefined condition threshold(s).
- An audio source having a value satisfying the corresponding condition threshold(s) is added into a cluster.
- all sources e.g., all audio sources of a number A
- a lower number of audio sources e.g., AT where M ⁇ N
- the substitution may involve a weighted downmix operation of the whole set or a subset of N audio sources. It may also be obtained by simply culling some of the sources.
- the output of step S703 may be a set of M audio sources to be rendered. These AT audio sources are substituting the original N audio sources without perceptual quality degradation when rendered.
- the present disclosure is directed to “typecasting” stage which converts one or several audio source(s) from one type to another.
- the resulting type is expected to be rendered more computationally efficient.
- the “typecasting” may be performed either during the core audio decoding stage or during the audio rendering stage.
- the “typecasting” method would determine if the distance between a listener and each of the several audio sources is sufficiently large (e.g., by comparing the distance between the listener and the audio source to a threshold). If the distance is large enough, then those audio sources (e.g., waterfall extent sound) can be substituted by an audio object/source (e.g., distant waterfall point-source). Alternatively, those audio point-sources (e.g., rain droplets) can be substituted by one HOA (e.g., rain noise) signal.
- HOA e.g., rain noise
- the present disclosure is directed to typecasting of audio sources that is based on one or more of these conditions: positional and directional proximity (see clustering discussion above for more detail regarding measurement of these parameters) level of acoustic/optical occlusion/abstraction
- the level of acoustic/optical occlusion/abstraction can be measured by factoring in several aspects such as the acoustic property (e.g., transmission, absorption, reflection coefficients) of an occluder that blocks the direct view from the listener and the distance of an audio source from the listener.
- acoustic property e.g., transmission, absorption, reflection coefficients
- those sources can simply be replaced by a single point source audio or an FOA.
- a source is considered occluded if the source is obstructed by an occluding material which makes it not visible to the listener.
- An occluding material is described by its acoustical properties such as transmission, reflection and absorption coefficients.
- the occluded multichannel audio sources may be replaced by a mono audio source obtained from the dominant channel of the multichannel audio or from a weighted mix of the audio sources.
- the occluded multichannel audio sources may be replaced by a stereo audio source.
- Figure 8 illustrates an exemplary method of typecasting audio sources according to the present disclosure.
- each audio source of a plurality of audio sources may be evaluated to obtain one or more values for the respective audio source, where the value(s) indicate perceptual relevance information.
- the value(s) can indicate (i) positional proximity; (ii) directional proximity; and/or (iii) level of acoustic/optical occlusion/abstraction.
- each of the value(s) is evaluated to determine whether it satisfy a corresponding predefined condition threshold(s). For example, the threshold is 5 degrees. If the obtained value is lower than 5 degrees, then it is considered to be in the near proximity in terms of directional proximity criterion.
- the resulting sources that satisfy the threshold are categorized as typecasting sources.
- An audio type may be a multichannel audio (e.g., stereo, 5.0, 7.0, 11.0), mono audio, Ambisonics (FOA, HO A)).
- Each of these methods can be implemented in the decoding and/or rendering stages of an audio decoder.
- the audio decoder and/or decoding method can be compatible with standards set by the audio group of MPEG of the ISO/IEC organization, such as the MPEG-I immersive audio standard.
- each of these stages alone or in combination can be implemented as part of the MPEG-I rendering process or as external tools for the MPEG-I immersive audio standard.
- Each method can be (de-)/activated by control information.
- the control information can be provided as conditions.
- the control information may be in the form of a parameter to be processed by a system and/or device.
- the param eter(s) can be defined per (virtual) object
- the application of the inventive methods may result in the change of audio source state.
- an audio source state is changed from “unculled” to “culled” or the other way around. If such a change happens then the actual change in the rendering should ideally be done when the corresponding audio source loudness is low. This is to avoid an abrupt change of a loud signal.
- the control information could be the loudness or energy threshold which sets a threshold of when such a change is triggered. Additionally, the control information could include the maximum allowed changes within a certain period to avoid, for example, highly repetitive changes within a short period. One example, only one change is allowed within 1 minute for speech signal. Alternatively, the control information may be set by a listener or an application at the decoder/ render er side.
- control information parameters that can be used to control application of perceptual audio culling, clustering and typecasting rendering stages include activation or deactivation parameters.
- activation or deactivation parameters can indicate when certain audio source(s) is disabled (e.g., signaling to disable a narrator speech signal).
- An activation example may be applied to certain audio sources in a scene such as car audio sources. Normally, a car sound may be active (e.g., engine started, car moving) or inactive (e.g., engine off, car parked) in a scene. In this case, the activation implies that the inventive methods (i.e., culling, clustering, typecasting) may be applied to car audio sources. This is not applicable for the narrator speech signal for example, where it is expected to always render the speech signal, hence the deactivation control.
- control information parameters is grouping data.
- the grouping data can be a set of audio source IDs that belong together such as parts of a car that emit sound (tyres, engine, horn). These audio sources can be grouped together if the directional differences between those sources relative to a listener position is less than 5 degrees (directional proximity).
- the grouping data can be used for controlling each of the culling/clustering/typecasting.
- the grouping data can provide means for definition of a set of audio sources within a certain area (e.g., sound emitting parts of a car) or with certain signal categories (e.g., ambience sound of rain and wind).
- an encoder, decoder, and/or Tenderer can define a culling and clustering strategy to either combine or remove audio sources (e.g., objects). For example, if all audio sources in the group are below a certain threshold loudness, the group is combined into one object (loudness-based culling). Alternatively, if azimuth/elevation of all objects of a group are each within a certain respective range, the group can be combined into one object (directional-based clustering). Alternatively, if one or many objects/sources of the group are acoustically occluded, the group is combined into one object/source.
- audio sources e.g., objects
- control information parameters are related to prioritization.
- the control information can provide for prioritization in respect to signal perceptual relevance (and its informational importance) to the listener (and listener’s focus of attention).
- the prioritization parameters can control when the culling of non-relevant audio sources is applied (instead of their leveling/EQing) for a “cocktail party effect” reproduction to improve speech intelligibility.
- control information parameters is visibility of graphical representation of audio sources to the listener (i.e., whether a visual representation of audio source is visible to the user or not).
- an audio source can be associated with visual representation, which is (i)not in the video rendered viewport (behind the listener) and (ii) obstructed by an optically non-transparent occluder.
- Such audio source can be treated differently depending on their nature and application scenario; e.g., thresholds for positional or directional proximity conditions can differ depending on a virtual object visibility.
- the control information parameters would define one or more sets of predefined condition thresholds for object visibility.
- Figure 9 illustrates exemplary uses of control information to control one or more the inventive culling, clustering and/or typecasting methods.
- listener pose (i.e., location) information and a plurality of audio sources P may be received.
- control information A can be received.
- IB control information B can be received.
- control information C can be received.
- 901 A, 901B, and 901C can each be performed individually, or in combination.
- Control information A illustrates an example of when the information controls both culling and clustering methods.
- Control information B and C are used to control only the clustering and typecasting methods, respectively.
- Figure 10 illustrates the control information parameters that would define one or more sets of predefined condition thresholds for object/source visibility.
- source A is clearly visible from the listener location (not obstructed by an occluder) and therefore a set of condition thresholds X is applied to perform the inventive audio culling, clustering, typecasting.
- source B is obstructed by an occluder and therefore another set of condition thresholds Y is applied to perform the inventive audio culling, clustering, typecasting.
- the control information here is whether a visual representation of audio source is visible to the user or not.
- apparatus 1100 comprises a processor 1110 and a memory 1120 coupled to the processor 1110.
- the memory 1120 may store instructions for the processor 1110.
- the processor 1110 may also receive, among others, suitable input data 1130, depending on use cases and/or implementations.
- the processor 1110 may be adapted to carry out the methods/techniques described throughout the present disclosure and to generate corresponding output data 1140 depending on use cases and/or implementations.
- a computing device implementing the techniques described above can have the following example architecture.
- Other architectures are possible, including architectures with more or fewer components.
- the example architecture includes one or more processors (e.g., dual-core Intel® Processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display) and one or more computer-readable mediums (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.).
- These components can exchange communications and data over one or more communication channels (e.g., buses), which can utilize various hardware and software for facilitating the transfer of data and control signals between components.
- computer-readable medium refers to a medium that participates in providing instructions to processor for execution, including without limitation, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory) and transmission media.
- Transmission media includes, without limitation, coaxial cables, copper wire and fiber optics.
- Computer-readable medium can further include operating system (e.g., a Linux® operating system), network communication module, audio interface manager, audio processing manager and live content distributor.
- Operating system can be multi-user, multiprocessing, multitasking, multithreading, real time, etc.
- Operating system performs basic tasks, including but not limited to: recognizing input from and providing output to network interfaces and/or devices; keeping track and managing files and directories on computer-readable mediums (e.g., memory or a storage device); controlling peripheral devices; and managing traffic on the one or more communication channels.
- Network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols, such as TCP/IP, HTTP, etc.).
- Architecture can be implemented in a parallel processing or peer-to-peer infrastructure or on a single device with one or more processors.
- Software can include multiple software components or can be a single body of code.
- the described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device.
- a computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result.
- a computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, a browser-based web application, or other unit suitable for use in a computing environment.
- Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer.
- a processor will receive instructions and data from a read-only memory or a random access memory or both.
- the essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data.
- a computer will also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks.
- Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto- optical disks; and CD-ROM and DVD-ROM disks.
- semiconductor memory devices such as EPROM, EEPROM, and flash memory devices
- magnetic disks such as internal hard disks and removable disks
- magneto- optical disks and CD-ROM and DVD-ROM disks.
- the processor and the memory can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).
- ASICs application-specific integrated circuits
- the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retina display device for displaying information to the user.
- the computer can have a touch surface input device (e.g., a touch screen) or a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer.
- the computer can have a voice input device for receiving voice commands from the user.
- the features can be implemented in a computer system that includes a back-end component, such as a data server, or that includes a middleware component, such as an application server or an Internet server, or that includes a front-end component, such as a client computer having a graphical user interface or an Internet browser, or any combination of them.
- the components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, e.g., a LAN, a WAN, and the computers and networks forming the Internet.
- the computing system can include clients and servers.
- a client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
- a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device).
- client device e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device.
- Data generated at the client device e.g., a result of the user interaction
- a system of one or more computers can be configured to perform particular actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions.
- One or more computer programs can be configured to perform particular actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
- any one of the terms comprising, comprised of or which comprises is an open term that means including at least the elements/features that follow, but not excluding others.
- the term comprising, when used in the claims should not be interpreted as being limitative to the means or elements or steps listed thereafter.
- the scope of the expression a device comprising A and B should not be limited to devices consisting only of elements A and B.
- Any one of the terms including or which includes or that includes as used herein is also an open term that also means including at least the elements/features that follow the term, but not excluding others. Thus, including is synonymous with and means comprising.
- a method for processing a plurality of audio sources comprising: determining, for each of the plurality of audio sources, a respective culling value; comparing, for each of the plurality of audio sources, the respective culling value of each of the plurality of sources to a threshold to determine whether this is a culling source; outputting information identifying a culling subset of the plurality of audio sources, wherein the culling subset includes the culling source(s) identified by the comparison.
- EEE2 The method of EEE1, wherein the culling value is based on at least one of: a loudness value, a relationship between loudness of a single source and overall loudness of overall rendered audio, and a relationship between short-term spectral energy of a single audio source an overall spectral energy of final rendered audio.
- EEE3 The method of EEE1, further comprising rendering the plurality of audio sources, wherein the rendering only renders audio sources that are not part of the culling subset of the plurality of audio sources.
- a method for processing a plurality of audio sources comprising: determining, for each of the plurality of audio sources, a respective clustering value; comparing, for each of the plurality of audio sources, the respective clustering value of each of the plurality of sources to a threshold to determine whether this is a cluster source; outputting information identifying a cluster subset of the plurality of audio sources based on the comparison.
- EEE5. The method of EEE4, wherein the clustering value indicates positional proximity.
- EEE6 The method of EEE5, wherein the threshold value is a fraction of a distance between a listener and a closest source.
- EEE7 The method of EEE6, wherein the comparison compares a distance between an audio source and the listener to the threshold value.
- EEE8 The method of EEE4, wherein the clustering value indicates directional proximity.
- EEE9 The method of EEE8, where in the threshold value is at least one of azimuth angle or elevation between sources with respect to a listener position.
- EEE 10 The method of EEE4, further comprising a set of clusters, and rendering the set of clusters by an audio Tenderer.
- a method for processing a plurality of audio sources comprising: determining, for each of the plurality of audio sources, a respective typecasting value; comparing, for each of the plurality of audio sources, the respective typecasting value of each of the plurality of sources to a threshold to determine whether this is a typecasting source; outputting information identifying a new type of source replacing the original source type.
- EEE12 A method of processing control information, wherein the control information determines whether to perform one or more of methods of EEEs 1, 4 and/or 11.
- EEE13 The method of EEE12, wherein the control information is one of: activation or deactivation parameters, grouping data, prioritization information, and/or visibility information.
- EEE 14 The method of any of EEEs 1-13, wherein the method is performed in accordance to a standard set by the MPEG audio group of ISO/IEC.
- EEE15 The method of EEE14, wherein the standard is the MPEG-I Immersive Audio standard.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Stereophonic System (AREA)
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263431822P | 2022-12-12 | 2022-12-12 | |
| US202363491258P | 2023-03-20 | 2023-03-20 | |
| PCT/EP2023/085408 WO2024126511A1 (en) | 2022-12-12 | 2023-12-12 | Method and apparatus for efficient audio rendering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4635205A1 true EP4635205A1 (de) | 2025-10-22 |
Family
ID=89452572
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23833615.0A Pending EP4635205A1 (de) | 2022-12-12 | 2023-12-12 | Verfahren und vorrichtung zur effizienten audiowiedergabe |
Country Status (7)
| Country | Link |
|---|---|
| EP (1) | EP4635205A1 (de) |
| JP (1) | JP2026500108A (de) |
| KR (1) | KR20250123162A (de) |
| CN (1) | CN120359766A (de) |
| AU (1) | AU2023393729A1 (de) |
| MX (1) | MX2025006557A (de) |
| WO (1) | WO2024126511A1 (de) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2026006172A1 (en) * | 2024-06-25 | 2026-01-02 | Dolby Laboratories Licensing Corporation | Audio object clustering system |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9564138B2 (en) * | 2012-07-31 | 2017-02-07 | Intellectual Discovery Co., Ltd. | Method and device for processing audio signal |
| CN115244501A (zh) * | 2020-03-10 | 2022-10-25 | 瑞典爱立信有限公司 | 音频对象的表示和渲染 |
| US20240155304A1 (en) * | 2021-05-17 | 2024-05-09 | Dolby International Ab | Method and system for controlling directivity of an audio source in a virtual reality environment |
-
2023
- 2023-12-12 JP JP2025530499A patent/JP2026500108A/ja active Pending
- 2023-12-12 KR KR1020257023046A patent/KR20250123162A/ko active Pending
- 2023-12-12 WO PCT/EP2023/085408 patent/WO2024126511A1/en not_active Ceased
- 2023-12-12 CN CN202380084899.7A patent/CN120359766A/zh active Pending
- 2023-12-12 AU AU2023393729A patent/AU2023393729A1/en active Pending
- 2023-12-12 EP EP23833615.0A patent/EP4635205A1/de active Pending
-
2025
- 2025-06-05 MX MX2025006557A patent/MX2025006557A/es unknown
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024126511A1 (en) | 2024-06-20 |
| JP2026500108A (ja) | 2026-01-06 |
| KR20250123162A (ko) | 2025-08-14 |
| MX2025006557A (es) | 2025-07-01 |
| AU2023393729A1 (en) | 2025-06-19 |
| CN120359766A (zh) | 2025-07-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP5625032B2 (ja) | マルチチャネルシンセサイザ制御信号を発生するための装置および方法並びにマルチチャネル合成のための装置および方法 | |
| US12431152B2 (en) | Apparatus and method for audio encoding | |
| EP2936485B1 (de) | Objektzusammenlegung für die auf perzeptiven kriterien beruhende wiedergabe objektbasierter audio-inhalte | |
| US9761229B2 (en) | Systems, methods, apparatus, and computer-readable media for audio object clustering | |
| US9516446B2 (en) | Scalable downmix design for object-based surround codec with cluster analysis by synthesis | |
| KR101450414B1 (ko) | 멀티-채널 오디오 프로세싱 | |
| RU2635884C2 (ru) | Устройство и способ для предоставления улучшенных характеристик направленного понижающего микширования для трехмерного аудио | |
| WO2012158705A1 (en) | Adaptive audio processing based on forensic detection of media processing history | |
| EP1991984A1 (de) | Verfahren, medium und system zum synthetisieren eines stereosignals | |
| JP6843992B2 (ja) | 相関分離フィルタの適応制御のための方法および装置 | |
| WO2024126511A1 (en) | Method and apparatus for efficient audio rendering | |
| RU2823537C1 (ru) | Устройство и способ кодирования аудио | |
| KR20240097694A (ko) | 임펄스 응답 결정 방법 및 상기 방법을 수행하는 전자 장치 | |
| TW202411984A (zh) | 用於具有元資料之參數化經寫碼獨立串流之不連續傳輸的編碼器及編碼方法 | |
| TW202429446A (zh) | 用於具有元資料之參數化經寫碼獨立串流之不連續傳輸的解碼器及解碼方法 | |
| KR20230139766A (ko) | 객체 오디오 렌더링 방법 및 상기 방법을 수행하는 전자 장치 | |
| HK1095195B (en) | Apparatus and method for generating multi-channel synthesizer control signal and apparatus and method for multi-channel synthesizing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250616 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| P01 | Opt-out of the competence of the unified patent court (upc) registered |
Free format text: CASE NUMBER: UPC_APP_0012839_4635205/2025 Effective date: 20251111 |
|
| REG | Reference to a national code |
Ref country code: HK Ref legal event code: DE Ref document number: 40126795 Country of ref document: HK |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |