EP4670156A1 - DIFFUSE PRESERVATION OF THE MERGING OF MASA AND ISM METADATA - Google Patents
DIFFUSE PRESERVATION OF THE MERGING OF MASA AND ISM METADATAInfo
- Publication number
- EP4670156A1 EP4670156A1 EP24703704.7A EP24703704A EP4670156A1 EP 4670156 A1 EP4670156 A1 EP 4670156A1 EP 24703704 A EP24703704 A EP 24703704A EP 4670156 A1 EP4670156 A1 EP 4670156A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- total ratio
- direct
- diffuse
- parameter
- ratio parameter
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/167—Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S3/00—Systems employing more than two channels, e.g. quadraphonic
- H04S3/008—Systems employing more than two channels, e.g. quadraphonic in which the audio signals are in digital form, i.e. employing more than two discrete digital channels
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
Definitions
- the present application relates to apparatus and methods for merging of MASA and ISM metadata aiming to preserve diffuse parameters, not exclusively for audio encoding.
- Background Parametric spatial audio capture from inputs such as microphone arrays and other sources, is a typical and an effective choice to estimate from the input (microphone array signals) a set of parameters such as directions of the sound in frequency bands, and the ratios between the directional and non-directional parts of the captured sound in frequency bands. These parameters are known to well describe the perceptual spatial properties of the captured sound at the position of the microphone array.
- a parameter set consisting of a direction parameter in frequency bands and an energy ratio parameter in frequency bands (indicating the directionality of the sound) can be also utilized as the spatial metadata (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance etc) for an audio codec.
- these parameters can be estimated from microphone-array captured audio signals, and, for example, a stereo or mono signal can be generated from the microphone array signals to be conveyed with the spatial metadata.
- Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency.
- An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec which is being designed to be suitable for use over a communications network such as a 3GPP 4G/5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR).
- IVAS Immersive Voice and Audio Services
- This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio, object-based audio, and scene-based audio inputs including spatial information about the sound field and sound sources.
- the codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.
- the stereo signal could be encoded, for example, using an IVAS audio core codec, or with an AAC (Advanced Audio Coding) or EVS (Enhanced Voice Services) encoder.
- a decoder can decode the audio signals into PCM (Pulse code modulation) signals and process the sound in frequency bands (using the spatial metadata) to obtain the spatial output, for example, a binaural output.
- PCM Pulse code modulation
- the aforementioned immersive audio codecs are particularly suitable for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, stand-alone microphone arrays).
- such an encoder can have other input types, for example, loudspeaker signals, audio object signals, Ambisonic signals.
- an apparatus comprising means for: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct- to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter.
- the means for selecting the diffuse-energy compensated direct-to-total ratio parameter when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.
- the means for generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may be for generating the at least one first weighting value based on a multiplication of the at least one first direct-to-total ratio parameter and the first signal energy parameter.
- the at least one first direct-to-total ratio may comprise at least two first direct-to-total ratios, wherein the means for generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may be for generating first weighting values associated with each of the at least two first direct-to-total ratio parameters.
- the means for generating first weighting values associated with each of the at least two first direct-to-total ratio parameters may be for generating each first weighting value based on a multiplication of each of the at least one first direct-to- total ratio parameter and the first signal energy parameter.
- the means for selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct- to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter may be further for: selecting the at least two first direct-to-total ratio parameters when the first weighting values are both greater than the second weighting value; and selecting the diffuse-energy compensated direct-to-total ratio parameter and one of the at least two first direct-to-total ratio parameters when the second weighting value is greater than one of the at least two first weighting values.
- the means for generating a second weighting value based, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter may be for one of: generating the at least one second weighting value based on a multiplication of the at least one second direct-to-total ratio parameter and the second signal energy parameter; generating the at least one second weighting value based on a multiplication of the diffuse-energy compensated direct-to-total ratio parameter and the second signal energy parameter; and generating the at least one second weighting value based on an average of the diffuse-energy compensated direct-to- total ratio parameter and the at least one second direct-to-total ratio parameter multiplied by the second signal energy parameter.
- the means for generating a diffuse-energy compensated direct-to-total ratio parameter may be for generating at least one merged diffuse-to-total ratio based at least in part on the first signal energy parameter, the second signal energy parameter and a diffuse signal energy parameter.
- the means for may be further for: obtaining, for the first audio stream at least one first diffuse-to-total ratio parameter; and obtaining the diffuse signal energy parameter based, in part, on the first signal energy parameter and the first diffuse- to-total ratio parameter.
- the means for generating a diffuse-energy compensated direct-to-total ratio parameter may be for generating the at least one diffuse-energy compensated direct-to-total ratio parameter based on the at least one merged diffuse-to-total ratio.
- the means for may be further for obtaining, for the second audio stream at least one second diffuse-to-total ratio parameter, wherein the means for generating at least one merged diffuse-to-total ratio may be further for generating the at least one merged diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.
- the means for generating the diffuse-energy compensated direct-to-total ratio parameter may be further for: generating at least one prospective diffuse- energy compensated direct-to-total ratio parameter based on the at least one merged diffuse-to-total ratio; comparing the at least one prospective diffuse-energy compensated direct-to-total ratio parameter to the second direct-to-total ratio parameter; selecting the diffuse-energy compensated direct-to-total ratio parameter from the at least one prospective diffuse-energy compensated direct-to- total ratio parameter and the second direct-to-total ratio parameter based on the comparing, such that the diffuse-energy compensated direct-to-total ratio parameter is the smaller of the at least one prospective diffuse-energy compensated direct-to-total ratio parameter and the second direct-to-total ratio parameter.
- the means may be further for encoding the selected one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to- total ratio parameter.
- the means for may be further for merging the first audio stream and the second audio stream.
- a method comprising: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and selecting
- Selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter may further comprise: selecting one of the at least one first direct-to-total ratio parameter when the at least one first weighting value is greater than the second weighting value; and selecting the diffuse-energy compensated direct-to-total ratio parameter when the second weighting value is greater than the at least one first weighting value. Selecting one of the at least one first direct-to-total ratio parameter when the at least one first weighting value is greater than the second weighting value may be when the at least one first weighting value is strictly greater than the second weighting value.
- Selecting the diffuse-energy compensated direct-to-total ratio parameter when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.
- Generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may comprise generating the at least one first weighting value based on a multiplication of the at least one first direct-to-total ratio parameter and the first signal energy parameter.
- the at least one first direct-to-total ratio may comprise at least two first direct-to-total ratios, wherein generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may comprise generating first weighting values associated with each of the at least two first direct-to-total ratio parameters. Generating first weighting values associated with each of the at least two first direct-to-total ratio parameters may comprise generating each first weighting value based on a multiplication of each of the at least one first direct-to-total ratio parameter and the first signal energy parameter.
- Selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter may comprise: selecting the at least two first direct-to-total ratio parameters when the first weighting values are both greater than the second weighting value; and selecting the diffuse-energy compensated direct-to-total ratio parameter and one of the at least two first direct-to-total ratio parameters when the second weighting value is greater than one of the at least two first weighting values.
- Generating a second weighting value based, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to- total ratio parameter; and the second signal energy parameter may comprise one of: generating the at least one second weighting value based on a multiplication of the at least one second direct-to-total ratio parameter and the second signal energy parameter; generating the at least one second weighting value based on a multiplication of the diffuse-energy compensated direct-to-total ratio parameter and the second signal energy parameter; and generating the at least one second weighting value based on an average of the diffuse-energy compensated direct-to- total ratio parameter and the at least one second direct-to-total ratio parameter multiplied by the second signal energy parameter.
- Generating a diffuse-energy compensated direct-to-total ratio parameter may comprise generating at least one merged diffuse-to-total ratio based at least in part on the first signal energy parameter, the second signal energy parameter and a diffuse signal energy parameter.
- the method may further comprise: obtaining, for the first audio stream at least one first diffuse-to-total ratio parameter; and obtaining the diffuse signal energy parameter based, in part, on the first signal energy parameter and the first diffuse-to-total ratio parameter.
- Generating a diffuse-energy compensated direct-to-total ratio parameter may comprise generating the at least one diffuse-energy compensated direct-to- total ratio parameter based on the at least one merged diffuse-to-total ratio.
- the method may further comprise obtaining, for the second audio stream at least one second diffuse-to-total ratio parameter, wherein generating at least one merged diffuse-to-total ratio may further comprise generating the at least one merged diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.
- Generating the diffuse-energy compensated direct-to-total ratio parameter may further comprise: generating at least one prospective diffuse-energy compensated direct-to-total ratio parameter based on the at least one merged diffuse-to-total ratio; comparing the at least one prospective diffuse-energy compensated direct-to-total ratio parameter to the second direct-to-total ratio parameter; selecting the diffuse-energy compensated direct-to-total ratio parameter from the at least one prospective diffuse-energy compensated direct-to- total ratio parameter and the second direct-to-total ratio parameter based on the comparing, such that the diffuse-energy compensated direct-to-total ratio parameter is the smaller of the at least one prospective diffuse-energy compensated direct-to-total ratio parameter and the second direct-to-total ratio parameter.
- the method may further comprises encoding the selected one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct- to-total ratio parameter.
- the method may further comprises merging the first audio stream and the second audio stream.
- an apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio
- the apparatus caused to perform selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct- to-total ratio parameter may be further caused to perform: selecting one of the at least one first direct-to-total ratio parameter when the at least one first weighting value is greater than the second weighting value; and selecting the diffuse-energy compensated direct-to-total ratio parameter when the second weighting value is greater than the at least one first weighting value.
- the apparatus caused to perform selecting one of the at least one first direct- to-total ratio parameter when the at least one first weighting value is greater than the second weighting value may be when the at least one first weighting value is strictly greater than the second weighting value.
- the apparatus caused to perform selecting the diffuse-energy compensated direct-to-total ratio parameter when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.
- the apparatus caused to perform generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may be further caused to perform generating the at least one first weighting value based on a multiplication of the at least one first direct-to- total ratio parameter and the first signal energy parameter.
- the at least one first direct-to-total ratio may comprise at least two first direct-to-total ratios, wherein the apparatus caused to perform generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may be further caused to perform generating first weighting values associated with each of the at least two first direct- to-total ratio parameters.
- the apparatus caused to perform generating first weighting values associated with each of the at least two first direct-to-total ratio parameters may be further caused to perform generating each first weighting value based on a multiplication of each of the at least one first direct-to-total ratio parameter and the first signal energy parameter.
- the apparatus caused to perform selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct- to-total ratio parameter may be further caused to perform: selecting the at least two first direct-to-total ratio parameters when the first weighting values are both greater than the second weighting value; and selecting the diffuse-energy compensated direct-to-total ratio parameter and one of the at least two first direct-to-total ratio parameters when the second weighting value is greater than one of the at least two first weighting values.
- the apparatus caused to perform generating a second weighting value based, in part, at least one of: the second direct-to-total ratio parameter; the diffuse- energy compensated direct-to-total ratio parameter; and the second signal energy parameter may be further caused to perform one of: generating the at least one second weighting value based on a multiplication of the at least one second direct- to-total ratio parameter and the second signal energy parameter; generating the at least one second weighting value based on a multiplication of the diffuse-energy compensated direct-to-total ratio parameter and the second signal energy parameter; and generating the at least one second weighting value based on an average of the diffuse-energy compensated direct-to-total ratio parameter and the at least one second direct-to-total ratio parameter multiplied by the second signal energy parameter.
- the apparatus caused to perform generating a diffuse-energy compensated direct-to-total ratio parameter may be further caused to perform generating at least one merged diffuse-to-total ratio based at least in part on the first signal energy parameter, the second signal energy parameter and a diffuse signal energy parameter.
- the apparatus may be further caused to perform: obtaining, for the first audio stream at least one first diffuse-to-total ratio parameter; and obtaining the diffuse signal energy parameter based, in part, on the first signal energy parameter and the first diffuse-to-total ratio parameter.
- the apparatus caused to perform generating a diffuse-energy compensated direct-to-total ratio parameter may be further caused to perform generating the at least one diffuse-energy compensated direct-to-total ratio parameter based on the at least one merged diffuse-to-total ratio.
- the apparatus may be further caused to perform obtaining, for the second audio stream at least one second diffuse-to-total ratio parameter, wherein the apparatus caused to perform generating at least one merged diffuse-to-total ratio may be further caused to perform generating the at least one merged diffuse-to- total ratio based on the second diffuse-to-total ratio parameter.
- the apparatus caused to perform generating the diffuse-energy compensated direct-to-total ratio parameter may be further caused to perform: generating at least one prospective diffuse-energy compensated direct-to-total ratio parameter based on the at least one merged diffuse-to-total ratio; comparing the at least one prospective diffuse-energy compensated direct-to-total ratio parameter to the second direct-to-total ratio parameter; selecting the diffuse-energy compensated direct-to-total ratio parameter from the at least one prospective diffuse-energy compensated direct-to-total ratio parameter and the second direct- to-total ratio parameter based on the comparing, such that the diffuse-energy compensated direct-to-total ratio parameter is the smaller of the at least one prospective diffuse-energy compensated direct-to-total ratio parameter and the second direct-to-total ratio parameter.
- the apparatus may be further caused to perform encoding the selected one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter.
- the apparatus may be further caused to perform merging the first audio stream and the second audio stream.
- an apparatus comprising: obtaining circuitry configured to obtain, for a first audio stream at least one first direct-to-total ratio parameter; obtaining circuitry configured to obtain, for the first audio stream, a first signal energy parameter; generating circuitry configured to generate at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining circuitry configured to obtain, for a second audio stream a second direct-to-total ratio parameter; obtaining circuitry configured to obtain, for the second audio stream a second signal energy parameter; generating circuitry configured to generate a diffuse-energy compensated direct-to-total ratio parameter; generating circuitry configured to generate a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and selecting circuitry configured to select based on a comparison of the at least one first weighting value and the second weighting value one
- a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct- to-total ratio parameter; and the second signal energy parameter; and selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio
- a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-
- an apparatus comprising: means for obtaining, for a first audio stream at least one first direct-to-total ratio parameter; means for obtaining, for the first audio stream, a first signal energy parameter; means for generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; means for obtaining, for a second audio stream a second direct-to-total ratio parameter; means for obtaining, for the second audio stream a second signal energy parameter; means for generating a diffuse-energy compensated direct-to- total ratio parameter; means for generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and means for selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct- to-total ratio parameter and the diffuse-energy compensated direct-to-to-
- a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining, for a first audio stream at least one first direct-to-total ratio parameter; obtaining, for the first audio stream, a first signal energy parameter; generating at least one first weighting value based on the at least one first direct- to-total ratio parameter and the first signal energy parameter; obtaining, for a second audio stream a second direct-to-total ratio parameter; obtaining, for the second audio stream a second signal energy parameter; generating a diffuse- energy compensated direct-to-total ratio parameter; generating a second weighting value based on, in part, at least one of: the second direct-to-total ratio parameter; the diffuse-energy compensated direct-to-total ratio parameter; and the second signal energy parameter; and selecting based on a comparison of the at least one first weighting value and the second weighting value one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to
- An apparatus comprising means for performing the actions of the method as described above.
- An apparatus configured to perform the actions of the method as described above.
- a computer program comprising program instructions for causing a computer to perform the method as described above.
- a computer program product stored on a medium may cause an apparatus to perform the method as described herein.
- An electronic device may comprise apparatus as described herein.
- a chipset may comprise apparatus as described herein.
- Figure 1 shows schematically an apparatus for MASA metadata extraction
- Figure 2 shows schematically an example MASA metadata stream merger
- Figure 3 shows schematically an example MASA stream merger as shown in Figure 2 in further detail
- Figure 4 shows schematically an example OMASA stream merger, for merging a MASA stream and ISM stream into a combined stream
- Figure 5 shows schematically an example system of apparatus suitable for implementing some embodiments
- Figure 6 shows schematically an example merger suitable for employing as part of the encoder as shown in Figure 5, configured to implement merging according to some embodiments.
- Figure 7 shows a flow diagram of the operation of the example metadata merger as shown in Figure 6 according to some embodiments
- Figure 8 shows schematically an additional example metadata merger suitable for employing as part of the encoder as shown in Figure 5, configured to implement merging according to some embodiments
- Figure 9 shows a flow diagram of the operation of the example additional example metadata merger as shown in Figure 8 according to some embodiments
- Figure 10 shows schematically another example metadata merger suitable for employing as part of the encoder as shown in Figure 5, configured to implement merging according to some embodiments
- Figure 11 shows schematically a multi-directional metadata merger suitable for employing as part of the encoder as shown in Figure 5, configured to implement merging according to some embodiments
- Figure 12 shows an example device suitable for implementing the apparatus shown in previous figures.
- Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the encoding of parametric spatial audio signals comprising transport audio signals and spatial metadata.
- immersive audio codecs such as 3GPP IVAS
- immersive audio codecs are being planned which support a multitude of operating points ranging from a low bit rate operation to transparency. It is expected to support channel-based audio, object-based audio, and scene-based audio inputs including spatial information about the sound field and sound sources.
- the example codec is configured to be able to receive multiple input formats.
- the codec is configured to obtain or receive a multi audio signal (for example, received from a microphone array, or as a multichannel audio format input, or an Ambisonics format input) and an audio object signal (these can also be called an Independent Stream with Metadata – ISM format).
- a multi audio signal for example, received from a microphone array, or as a multichannel audio format input, or an Ambisonics format input
- an audio object signal these can also be called an Independent Stream with Metadata – ISM format.
- This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats.
- An example of two different audio input formats being currently considered is the combination of the MASA format with audio object format (ISM format).
- Metadata-Assisted Spatial Audio is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.
- spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, spread coherence, distance, etc.) per time-frequency tile.
- the spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene.
- a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency subframe (and associated with each direction direct-to- total ratios, spread coherence, distance values etc) are determined.
- Figure 1 is shown an example MASA analyser 101.
- the MASA analyser 101 is configured to receive the input audio signal(s) 100 and analyse the input audio signals to generate transport audio signal(s) 102 and spatial metadata 104.
- MASA spatial metadata is presented in the following table. These values are available for each time-frequency tile.
- a frame is subdivided into 24 frequency bands and 4 temporal sub-frames. In other implementations other divisions of frequency and time can be employed.
- a frame size (for example as implemented in IVAS) is 20 ms (and thus the temporal sub-frame is 5 ms).
- the MASA analyser is configured to determine 1 or 2 directions for each time- frequency tile (i.e., there are 1 or 2 direction index, direct-to-total energy ratio, and spread coherence parameters for each time-frequency tile).
- the analyser is configured to generate more than 2 directions for a time-frequency tile.
- Field bits Description Direction index 16 Direction of arrival of the sound at a time-frequency parameter interval.
- Direct-to-total 8 Energy ratio for the direction index i.e., time-frequency energy ratio subframe). Calculated as energy in direction / total energy. Range of values: [0.0, 1.0] Values stored as 8-bit unsigned integers with uniform spacing of mapped values.
- Spread coherence 8 Spread of energy for the direction index i.e., time- frequency subframe).
- the MASA stream can be rendered to various outputs, such as multichannel loudspeaker signals (e.g., 5.1) or binaural signals.
- Another input format to be supported by IVAS is, as discussed above, the ISM (Independent Streams with Metadata) format.
- the ISM format is intended for representing individual sound sources (audio objects) in a sound scene with the associated metadata describing how the rendering of the audio signal may be implemented.
- Example metadata for the ISM format may comprise the following parameters: Position (e.g., as azimuth and elevation; or as azimuth, elevation and radius); Orientation; Extent/spread; Distance attenuation; and Directivity pattern.
- Position e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.g., as azimuth and elevation; or as azimuth, elevation and radius
- Orientation e.
- an ISM format input can be obtained by capturing individual sources in a scene using, for example, close or lavalier microphones (located on or near to an individual source). Examples of such individual sources are: separate speakers in a teleconference, a singer, or individual instruments. Metadata can then be associated to these ISM format signals automatically by, for example, position trackers, or manually by a mixing professional. Alternatively, ISM format signals may be generated by mixing and creating suitable associated metadata. An example of fully generated sound sources is in game audio engines. As indicated above research is being carried out into enabling the IVAS codec to support combined coding of multiple audio formats.
- This combined format can be referred to as an “Objects and MASA”-format (OMASA), which combines MASA format and ISM format into single combined format.
- OMASA Objects and MASA
- GB2217905.5 describes bitrate-dependent operation modes of OMASA coding.
- the following examples feature the lowest bitrate operation modes at which the ISM format input is merged with the MASA format input and where they are encoded as a MASA format stream.
- embodiments can use other modes or can be employed in other OMASA coding situations.
- a format merging method is required to merge the format together into a form more suitable for encoding and transmitting through the tight bitrate limitation.
- An example merging method has been described in GB2574238.
- This merging method shows two parametric spatial audio format inputs such as MASA format which are merged into a single output parametric spatial audio format.
- MASA format which are merged into a single output parametric spatial audio format.
- Figure 2 shows (metadata) stream merger 201 which is configured to receive a MASA stream 1200 and a MASA stream 2202 and generate a (MASA) combined stream 204.
- Figure 3 furthermore shows the example (metadata) stream merger 201 in further detail and with respect to metadata merging.
- the (metadata) stream merger 201 is configured to receive the MASA stream 1200 and MASA stream 2202.
- the example (metadata) stream merger 201 comprises an energy determiner (reference 311 for stream 1 and reference 313 for stream 2) which is configured to determine stream 1 energy 300 and stream 2 energy 302.
- the energy determiner can, for example, compute the energy of the MASA transport signal in each parameter TF-tile ( ⁇ , ⁇ ) separately for both MASA format streams:
- ⁇ ⁇ ⁇ ( ⁇ , ⁇ ) is the CLDFB-domain representation of ⁇ th MASA stream in ⁇ th transport channel in CLDFB bin ⁇ and slot ⁇ that are grouped into the parameter band ⁇ and sub-frame ⁇ .
- the sub-frame grouping is employed since the values are usually computed per 5 ms (and the CLDFB representation has temporal resolution of 1.25 ms), but in some embodiments can be determined on a per frame basis, in other words, every 20 ms.
- the indices ( ⁇ , ⁇ ) are omitted in the following text for clarity, but the operations are done for each TF-tile. This computation is done for both streams to be merged and are represented by stream 1 energy 300 and stream 2 energy 302.
- the example (metadata) stream merger 201 further comprises a ratio determiner (reference 301 for stream 1 and reference 303 for stream 2). The ratio determiner can obtain the ratios from the MASA streams and pass these to the weight determiner.
- the example (metadata) stream merger 201 can comprise a weight determiner (reference 305 for stream 1 and reference 307 for stream 2).
- the example (metadata) stream merger 201 further comprises a weight comparator 309 configured to compare the weights for each TF-tile and control a metadata selector 321.
- the example (metadata) stream merger 201 comprises a metadata selector 321 which is configured to receive the metadata from the MASA streams and the control from the weight comparator 309.
- the metadata selector 321 can be configured such that If ⁇ ⁇ > ⁇ ⁇ for the TF-tile, then select metadata from MASA stream 1 for that TF-tile. Otherwise, select metadata from MASA stream 2 for that TF-tile.
- the selected metadata can then be output as the merged metadata within the (MASA) combined stream 204.
- the comparison “greater than” > is the comparison “greater than or equal to” ⁇ . This option can be applied throughout the following examples.
- the comparison can contain a bias factor (which can be an additive and/or multiplicative factor) configured to bias the decision in a direction.
- a bias factor which can be an additive and/or multiplicative factor
- This can be furthermore extended to input formats other than MASA input formats.
- An example of which is shown in Figure 4 where an ISM stream and MASA stream are merged and encoded to generate a bitstream.
- Figure 4 shows a merger/encoder comprising an audio stream merger 401 configured to generate a combined audio stream 412 from the stream 1 (MASA) input audio signals 400 and the stream 2 (ISM) input audio signals 402.
- the merger/encoder further comprises an audio stream energy determiner 403 which is configured to determine the audio signal energies 408, for example in a manner similar to that described above with respect to the energy determiner 311, 313 as shown in Figure 3.
- the merger/encoder can comprise an ISM-to-MASA converter 405 which is configured to receive the (ISM) stream 2406 and generate a MASA stream 2202. As the metadata of an ISM stream contains values that are compatible with the values of a MASA stream the ISM-to-MASA converter 405 is configured to generate or construct a MASA stream from an ISM stream.
- the MASA stream can, for example, be generated by converting the position of an ISM stream into direction of a MASA stream and making the constructed MASA stream completely directional. In other words, assigning direct-to-total ratio value of 1 and diffuse-to-total ratio value of 0 for the constructed MASA format.
- the definition of the ISM to be a completely directive sound source and assign it with a direct-to-total ratio of 1 and diffuse-to-total ratio of 0 can be justified by the envisioned application in which the ISM is obtained as single audio objects, for example, a lavalier microphone.
- the converter 405 is configured to convert the ISM into first order Ambisonics (FOA) representations and analyze the source direction from representation.
- the direct-to-total energy ratio can be obtained by estimating the diffuse-to-total energy ratio of this FOA signal and determining the direct-to-total energy ratio from this.
- Any other suitable method may be used for converting the ISM stream to a MASA stream (with the above providing two specific examples).
- the two MASA format stream metadata can then be merged.
- the merger/encoder for example can comprise a metadata stream merger 407 which is configured to receive the (MASA) stream 1200 and MASA stream 2 202 and from these inputs merge the metadata in a manner similar to that described by the examples shown with respect to Figure 3.
- the combined MASA stream 410 generated by the metadata stream merger 407 can then be output to the spatial metadata encoder 409.
- the merger/encoder can then comprise a spatial metadata encoder 409 configured to encode the combined MASA stream 410 based on a suitable MASA metadata encoding method and output an encoded metadata to a bitstream creator 413.
- the merger/encoder can furthermore comprise a bitstream creator 413 configured to receive the encoded audio 414 from the audio encoder 411 and the encoded metadata 416 from the spatial metadata encoder 409 and generate a bitstream 418 to be output for storage and/or transmission.
- a bitstream creator 413 configured to receive the encoded audio 414 from the audio encoder 411 and the encoded metadata 416 from the spatial metadata encoder 409 and generate a bitstream 418 to be output for storage and/or transmission.
- FIG. 5 is shown an example system of apparatus suitable for implementing some embodiments.
- the system of apparatus is configured to receive the stream 1 input audio signal(s) 500, stream 1 spatial metadata (in format 1) 504 and stream 2 input audio signal(s) 502, stream 2 spatial metadata (in format 2) 506. These are received by the merger/encoder 501 which is configured to merge the inputs and generate a bitstream 510.
- the bitstream is then passed to the decoder 503 which then decodes the bitstream and generates the output audio signal(s) 512.
- the concept is to improve on the examples above merging MASA streams and MASA streams and ISM-converted-to-MASA streams in general.
- the above methods can suffer a loss in quality especially when time-frequency resolution and assigned bitrate of the encoded and transmitted merged MASA format is restricted, e.g., 24.4 kbps. Although the loss in quality is more notable at the lowest bitrates, where the time-frequency resolution is low, there can be loss even at higher bitrates, and therefore it is the aim to improve quality for a range of bitrates.
- the reason for the loss in quality is that the direct-to-total ratio of close to 1 of the ISM stream combined with usually a significant amount of signal energy emphasizes the presence of the ISM stream in the produced merged and encoded MASA format. This can be perceived as disturbing instability in the reproduced sound scene as the surrounding diffuse sound present in the MASA stream can disappear almost completely when the ISM stream is active.
- coding artifacts e.g., heavily quantized directions
- the concept as discussed hereafter with respect to some embodiments and examples is apparatus and methods that aim to preserve the overall diffuse properties of the MASA format source stream in the merged MASA format streams while being able to merge a MASA stream and an ISM-converted-to-MASA stream.
- the embodiments and example shown herein relate to merging two parametric spatial audio (i.e., audio signal(s) associated with spatial metadata) streams, where the first stream is a MASA format stream and the second stream is an ISM stream.
- the apparatus and methods propose adjusting the merging of the parameters of these two spatial metadata streams such that the perceived overall quality of the merged stream is improved by preserving the MASA format scene envelopment through preserving its diffuse energy.
- the selection can be “MASA stream” parameters for the TF-tile in the MASA stream where the MASA stream related weights are greater than the ISM stream related weights, a diffuse-energy-compensated ISM stream parameters where the ISM stream related weights are greater than the MASA ones, or a selection of the MASA stream and diffuse-energy-compensated ISM stream parameters for the TF-tile based of the value of the ISM stream related weights being greater than at least one of the more than one MASA stream related weight.
- weight(s) and weighting values are interchangeable. As such the apparatus and method aim to provide improved merging.
- the examples presented herein are two parametric spatial audio streams when one of the streams is originally an ISM stream, i.e., one or more independent audio objects, and the other stream is a MASA format stream (or any similar parametric format), in some embodiments other stream formats can be merged based on the examples presented herein and without significant inventive adjustment.
- the metadata of one input format data stream shows similar properties as the ISM stream, in other words having high directivity and low diffuseness, then this input format stream is particularly suitable for merging using the examples described herein.
- the apparatus and methods as described herein can be implemented within a communication codec encoder such as 3GPP IVAS and a combined format coding can be applied at a low bitrate where full merging of the separate format streams is required for a good quality transmission.
- the embodiments as described herein can furthermore be implemented as part of a merger/encoder 501 such as shown in Figure 5.
- the apparatus is configured to receive as inputs two streams of parametric spatial audio (“Stream 1” and “Stream 2”).
- Each of the input streams contains audio signal(s) + metadata (“Stream 1 input audio signal(s)”, “Stream 1 spatial metadata (format 1)”, “Stream 2 input audio signal(s)”, “Stream 2 spatial metadata (format 2)”).
- the streams are a MASA format stream and an ISM stream.
- the merger/encoder may contain a pre-processor which converts captured audio signals into MASA format streams and ISM streams.
- the input streams are given to an encoder which quantizes and encodes the formats into a single bitstream.
- This bitstream is transmitted to a decoder which in turn decodes the bitstream and produces the output audio signal(s), or some embodiments, an output format (e.g., MASA format).
- An example merger shown in Figure 6, is suitable for implementing within the merger/encoder of Figure 5 the metadata merging operations. In these embodiments the merging or combining of audio signals and metadata is implemented using different methods.
- the combining or merging of the audio signals can be implemented using any suitable manner. For example, by mixing the input audio signals together to provide merged audio signals that are then encoded. As such the merging or combining of the audio signals is not described in further detail.
- the (metadata) merger or combiner according to some embodiments is shown in Figure 6.
- the merger is configured to receive the MASA stream 1 metadata 600, MASA stream 1 transport signal(s) 602, and ISM stream metadata/object audio 610.
- the merger in some embodiments comprises an ISM-to-MASA converter 607.
- the ISM-to-MASA converter 607 is configured to receive the ISM stream metadata/object audio 610 and generate the MASA stream 2 metadata 614, and the ISM (MASA) stream transport signal(s) 612.
- the MASA stream 2 metadata 614 is passed to the direct-to-total ratio obtainer 609.
- the ISM (MASA) stream transport signal(s) 612 is passed to the transport signal energy determiner 611 and to the audio stream merger 670.
- the merger comprises a direct-to-total ratio obtainer 601 which is configured to obtain the MASA stream 1 metadata 600 and extract or otherwise obtain the MASA ratio ⁇ ⁇ ⁇ ⁇ 604 which is output to the weight determiner 605.
- the metadata merger comprises a direct-to-total ratio obtainer 609 which is configured to obtain the MASA stream 2 metadata 614 and extract or otherwise obtain the ISM based ratio 616 which is output to the weight determiner 613.
- the metadata merger comprises a transport signal energy determiner 603 which is configured to estimate the energy of the MASA transport signal in each parameter TF-tile:
- ⁇ ⁇ ( ⁇ , ⁇ ) is the CLDFB-domain representation of ⁇ th transport channel in CLDFB bin ⁇ and slot ⁇ that are grouped into the parameter band ⁇ and frame ⁇ .
- the indices ( ⁇ , ⁇ ) are omitted for clarity, but the operations are done for each TF-tile.
- the output ⁇ ⁇ 606 is output to the weight determiner 605, the diffuse signal energy obtainer 653 and merged diffuse-to-total ratio determiner 655.
- the merger comprises a further transport signal energy determiner 611 is configured to compute a transport signal energy estimate for the ISM transport signal 612 in each parameter TF-tile:
- ⁇ ⁇ ( ⁇ , ⁇ ) is the CLDFB-domain representation of ⁇ th channel of the ISM transport signal.
- the indices ( ⁇ , ⁇ ) are hereafter omitted for clarity, but the operations are done for each TF-tile.).
- the energy of the ISM transport signal ⁇ ⁇ 618 is which is output to the weight determiner 613 and the merged diffuse-to-total ratio determiner 655.
- the merger further comprises a diffuse-to-total ratio obtainer 651 which is configured to obtain or extract the diffuse-to-total ratio ⁇ ⁇ ⁇ 654 from the MASA stream 1 metadata 600.
- the merger further comprises a diffuse signal energy obtainer 653 is configured to obtain the diffuse-to-total ratio ⁇ ⁇ ⁇ 654 and the energy of the MASA transport signal in each parameter TF-tile ⁇ ⁇ 606 and generate a diffuse energy value ⁇ ⁇ ⁇ 656.
- the diffuse signal energy is determined based on the following
- the merger comprises a merged diffuse-to-total ratio determiner 655 which is configured to compute a merged diffuse-to-total ratio of the merged MASA stream 1 and ISM stream with: This value describes the amount of diffuse energy in the signal after merging the two streams.
- the diffuse signal energy ⁇ ⁇ 658 is passed to a merged direct- to-total ratio determiner 657.
- the metadata merger comprises a weight comparator 615 configured to compare the weights and generate a selection control to the metadata selector 617.
- the metadata merger furthermore comprises a metadata selector 617 configured to select and output selected metadata 622 based on the selection control from the weight comparator.
- the selector can be configured to implement the following selection operations: If ⁇ ⁇ > ⁇ ⁇ for the TF-tile, then select metadata from MASA stream 1600 for that TF-tile; Otherwise, select metadata from MASA stream 2614 (in other words, the metadata from the original “ISM stream”) for that TF-tile and use the new direct-to- total ratio ⁇ ⁇ 660.
- the determined new direct-to-total ratio ⁇ ⁇ 660 from the Merged direct-to-total ratio determiner 657 is selected (and effectively has replaced the possible selection of the original direct-to-total ratio 616 of the ISM stream).
- the merger comprises an audio stream merger 670 configured to receive the MASA stream 2 transport/audio signals 664 from the ISM-to-MASA converter 607 and the MASA stream 1 transport signal(s) 602 and from these generate a combined stream transport/audio signals 674 which can be output to the suitable audio encoder.
- the merged metadata 622 can then be passed to a suitable spatial metadata encoder and the result can be multiplexed together with the audio signals, for example, within a suitable bitstream creator to provide the encoded bitstream which can be transmitted to decoder (or stored for later decoding).
- the operations of the decoder are not discussed in further detail as the decoding of the bitstream and decoder itself can be implemented using any suitable decoder. In other words, no changes are required there to support the merged MASA stream.
- the decoder demultiplexes and decodes the bitstream into audio channels and metadata which are then either output as a MASA format or passed to a renderer which renders them into various output audio signals.
- FIG. 7 is shown an example flow diagram of the operations shown in Figure 6.
- obtaining MASA stream 1 metadata as indicated by 700.
- obtaining MASA stream 1 transport signal(s) by 702.
- obtaining ISM stream metadata/object audio by 721.
- the obtaining of the Direct-to-total ratio from the MASA stream 1 metadata is shown by 703.
- the determining of the transport signal energy from the MASA stream 1 transport signal(s) is shown by 705.
- the obtaining of the Diffuse-to-total ratio from the MASA stream 1 metadata is shown by 709.
- the obtaining of the diffuse signal energy from the diffuse-to-total ratio and transport signal energy is shown by 711.
- the determining of stream 1 weights from the direct-to-total ratio and transport signal energy is shown by 707.
- the conversion of the ISM-to-MASA metadata and obtaining ISM stream (or MASA stream 2) audio is shown by 723.
- the obtaining of the direct-to-total ratio from the converted MASA metadata is shown by 725.
- the determining of transport signal energy from the ISM stream transport signal(s) is shown by 719.
- the determining of stream 2 weights from the direct-to-total ratio and transport signal energy (for the original ISM stream) is shown by 727.
- the determining of a merged diffuse-to-total ratio is shown by 713.
- the determining of the merged direct-to-total ratio from the merged-to- total ratio is shown by 715.
- the comparison of the weights is shown by 729.
- the selection of the metadata is shown based on the comparison is shown by 731.
- the outputting of the selected metadata is shown by 735.
- the merging of the audio streams is shown by 741.
- the combined or merged audio stream is output as shown by 743.
- This implementation aims to take better into account the overall diffuse energy of the MASA format which is important for conveying the perceptual envelopment of the sound scene.
- This aims to improve the known merging methods, when the metadata from the ISM format stream would be selected for use in a TF-tile, it would also often completely remove (or severely reduce) diffuse energy for that tile.
- the embodiments thus adjust the amount of diffuse energy by taking into account the increased total energy from the ISM stream.
- ISM stream energy As ISM stream energy is assumed to be dominantly direct, it will decrease diffuse energy for the tile but only by proportion of relative stream energies. This aims to significantly improve a quality of the merged stream as the diffuse field perception is preserved better.
- these example embodiments aim to provide a significant quality improvement, the decision criteria and/or the determined ratio values for the merged stream metadata may be further adjusted in some alternative embodiments in an attempt to further improve the quality.
- these further embodiments are discussed in the following. It should also be noted that these improvements can also be used together. For example, it may be possible to apply the improvement of the two following examples in some embodiments.
- the difference from the earlier described embodiment is that a slightly different weight is used for the stream decision, but when the decision is the same, the resulting metadata will be the same, too.
- the weight determiner 813 is configured to determine an average of the two direct-to-total ratios ⁇ ⁇ ⁇ ⁇ , ⁇ . This is shown in Figure 8 by the dashed line elements of the MASA stream 2 metadata 614 which is processed by the direct-to-total ratio determiner 609 which outputs the ISM direct-to-total ratio ⁇ ⁇ to the weight determiner 813.
- Figure 9 shows a flow diagram showing the operations of Figure 8 example apparatus which differs from Figure 7 in that the operation of determining the stream 2 weights differs in the elements used to generate the weight as shown by 927.
- the difference to the example shown with respect to Figures 6 and 8 is that a (slightly) different weight is used for the stream decision.
- the decision is the same, the resulting selected metadata will also be the same, too.
- using these adjusted decision criteria can preserve the sound scene stability and further preserve the experience of spaciousness in the scene, as the MASA stream is selected more often, especially in case when the ISM stream contains ratios close to 1.
- the conversion from ISM to MASA representation can be implemented by assuming that the signal energy is completely directional assigning the direct-to-total energy ratio to 1. While this is true in some situations, it is possible to do the conversion using other methods, for example, using a FOA representation. In such embodiments it is possible that the ISM contains some amount of diffuse energy.
- non-negligible diffuse- to-total energy ratio may be present in the ISM stream
- two or more ISM objects are active in the same TF-tile, leading to a direct-to-total energy ratio that may be significantly smaller than 1.
- Such a parametrization may be desirable for a more realistic perception of space in the reproduction. If the amount of the diffuse energy in the ISM is non-negligible, it may be useful to take this diffuse energy into account. Otherwise, the ratio values may unintentionally be increased for TF tiles in which the ISM metadata was selected to be used but the MASA stream had less diffuse energy than the ISM stream. This may be perceived as losing spaciousness in the merged stream.
- the accounting for the diffuse energy can employ many forms.
- the resulting direct-to-total energy ratio ⁇ ⁇ may be limited to not be larger than the original direct-to-total energy ratio, in case the ISM metadata was selected for the TF tile, i.e.,
- the modified direct-to-total energy ratio ⁇ ⁇ , ⁇ 1010 is then passed to the metadata selector 709.
- the direct-to-total/diffuse-to-total ratio obtainer 1009 configured to obtain not only the direct-to-total ratio but also the diffuse-to-total ratio 1016 which is passed to the merged diffuse-to-total determiner 1005.
- the direct-to-total ratio obtainer 1101 is configured to obtain two direct- to-total energy ratios ⁇ ⁇ and ⁇ ⁇ in the original MASA metadata stream, one for each direction and which are also designated ⁇ , ⁇ ⁇ 1104.
- weights are compared in the weight comparator 715 and then selected by the metadata selector based on: If ⁇ ⁇ , ⁇ > ⁇ ⁇ and ⁇ ⁇ , ⁇ > ⁇ ⁇ , use the “MASA stream 1” parameters; If ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ or ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ , use parameters from “MASA direction 2” and “ISM stream”.
- the diffuse signal energy obtainer 1103 being configured to determine the MASA’s diffuse signal energy
- the merged direct-to-total ratio determiner 1107 is then configured to scale the newly-assigned direct-to-total energy ratios ⁇ ⁇ , ⁇ ⁇ depending on this new diffuse-to-total energy ratio with Similar as with the previous examples, this two-direction embodiment can be extended to include the diffuse energy of the ISM-objects too in a manner similar to the embodiment described in Figures 8 and 10.
- the merging or combining of metadata is selected based on a determining of a measure of “object-likeness” from the spatial metadata of the streams to be merged.
- a weight is determined based on an object-likeness parameter.
- both streams have low object-likeness, e.g., they contain a non-negligible amount of diffuse energy, then the merging method as shown in Figure 6 is implemented. If both streams have high object-likeness, e.g., both have a negligible amount of diffuse energy, then the merging method as shown in Figure 6 is implemented.
- the merging method as shown in embodiments shown in any of Figures 9 onwards or described above can be implemented.
- the “object-likeness” parameter can be determined by inspecting the overall diffuse energy (objects are usually highly directive with minimal or no diffuse part) and inspecting direction variation over frequency as there usually is no variation over frequency with objects.
- CDFB Complex-Values Low-Delay Filter Bank
- STFT Short-time Fourier Transform
- QMF Quadrature Mirrored Filterbank
- the parametric format described above has been MASA format but the embodiments can be extended to other parametric formats such as parametric coding of Ambisonics or multi-channel mixes.
- an example electronic device which may be used as any of the apparatus parts of the system as described above.
- the device may be any suitable electronics device or apparatus.
- the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
- the device may for example be configured to implement the encoder and/or decoder or any functional block as described above.
- the device 1400 comprises at least one processor or central processing unit 1407.
- the processor 1407 can be configured to execute various program codes such as the methods such as described herein.
- the device 1400 comprises at least one memory 1411.
- the at least one processor 1407 is coupled to the memory 1411.
- the memory 1411 can be any suitable storage means.
- the memory 1411 comprises a program code section for storing program codes implementable upon the processor 1407.
- the memory 1411 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein.
- the implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1407 whenever needed via the memory-processor coupling.
- the device 1400 comprises a user interface 1405.
- the user interface 1405 can be coupled in some embodiments to the processor 1407.
- the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405.
- the user interface 1405 can enable a user to input commands to the device 1400, for example via a keypad.
- the user interface 1405 can enable the user to obtain information from the device 1400.
- the user interface 1405 may comprise a display configured to display information from the device 1400 to the user.
- the user interface 1405 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1400 and further displaying information to the user of the device 1400.
- the user interface 1405 may be the user interface for communicating.
- the device 1400 comprises an input/output port 1409.
- the input/output port 1409 in some embodiments comprises a transceiver.
- the transceiver in such embodiments can be coupled to the processor 1407 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network.
- the transceiver or any suitable transceiver or transmitter and/or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
- the transceiver can communicate with further apparatus by any suitable known communications protocol.
- the transceiver can use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or can be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad-hoc networks (MANETs), cellular internet of things (IoT) RAN and Internet Protocol multimedia subsystems (IMS), any other suitable option and/or any combination thereof.
- LTE Advanced long term evolution advanced
- NR new radio
- 5G long term evolution advanced
- UMTS universal mobile telecommunications system
- UTRAN or E-UTRAN
- the transceiver input/output port 1409 may be configured to receive the signals.
- the device 1400 may be employed as at least part of the synthesis device.
- the input/output port 1409 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar and loudspeakers.
- headphones which may be a headtracked or a non-tracked headphones
- loudspeakers similar and loudspeakers.
- the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
- some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
- the software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
- the memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
- the data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
- Embodiments of the inventions may be practiced in various components such as integrated circuit modules.
- the design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules.
- the resultant design in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
- a standardized electronic format e.g., Opus, GDSII, or the like
- circuitry may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
- hardware-only circuit implementations such as implementations in only analog and/or digital circuitry
- software such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (i
- circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
- non-transitory is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Multimedia (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Mathematical Physics (AREA)
- Quality & Reliability (AREA)
- Stereophonic System (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2302631.3A GB2627482A (en) | 2023-02-23 | 2023-02-23 | Diffuse-preserving merging of MASA and ISM metadata |
| PCT/EP2024/052451 WO2024175321A1 (en) | 2023-02-23 | 2024-02-01 | Diffuse-preserving merging of masa and ism metadata |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4670156A1 true EP4670156A1 (en) | 2025-12-31 |
Family
ID=85794218
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24703704.7A Pending EP4670156A1 (en) | 2023-02-23 | 2024-02-01 | DIFFUSE PRESERVATION OF THE MERGING OF MASA AND ISM METADATA |
Country Status (8)
| Country | Link |
|---|---|
| EP (1) | EP4670156A1 (en) |
| KR (1) | KR20250150089A (en) |
| CN (1) | CN120752699A (en) |
| AU (1) | AU2024226319A1 (en) |
| CO (1) | CO2025012807A2 (en) |
| GB (1) | GB2627482A (en) |
| MX (1) | MX2025009724A (en) |
| WO (1) | WO2024175321A1 (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2639905A (en) * | 2024-03-27 | 2025-10-08 | Nokia Technologies Oy | Rendering of a spatial audio stream |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2217905A (en) | 1988-04-13 | 1989-11-01 | Ac Dc Holdings Limited | Discharge lamps |
| GB2574238A (en) | 2018-05-31 | 2019-12-04 | Nokia Technologies Oy | Spatial audio parameter merging |
| WO2020008112A1 (en) * | 2018-07-03 | 2020-01-09 | Nokia Technologies Oy | Energy-ratio signalling and synthesis |
| EP4462821A3 (en) * | 2018-11-13 | 2024-12-25 | Dolby Laboratories Licensing Corporation | Representing spatial audio by means of an audio signal and associated metadata |
| GB2598773A (en) * | 2020-09-14 | 2022-03-16 | Nokia Technologies Oy | Quantizing spatial audio parameters |
-
2023
- 2023-02-23 GB GB2302631.3A patent/GB2627482A/en not_active Withdrawn
-
2024
- 2024-02-01 AU AU2024226319A patent/AU2024226319A1/en active Pending
- 2024-02-01 WO PCT/EP2024/052451 patent/WO2024175321A1/en not_active Ceased
- 2024-02-01 CN CN202480014214.6A patent/CN120752699A/en active Pending
- 2024-02-01 EP EP24703704.7A patent/EP4670156A1/en active Pending
- 2024-02-01 KR KR1020257030751A patent/KR20250150089A/en active Pending
-
2025
- 2025-08-18 MX MX2025009724A patent/MX2025009724A/en unknown
- 2025-09-19 CO CONC2025/0012807A patent/CO2025012807A2/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| AU2024226319A1 (en) | 2025-10-09 |
| KR20250150089A (en) | 2025-10-17 |
| CO2025012807A2 (en) | 2025-09-29 |
| GB202302631D0 (en) | 2023-04-12 |
| CN120752699A (en) | 2025-10-03 |
| MX2025009724A (en) | 2025-09-02 |
| WO2024175321A1 (en) | 2024-08-29 |
| GB2627482A (en) | 2024-08-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12452619B2 (en) | Spatial audio representation and rendering | |
| US12451147B2 (en) | Spatial audio parameter encoding and associated decoding | |
| US20250157475A1 (en) | Parametric spatial audio rendering | |
| WO2024175321A1 (en) | Diffuse-preserving merging of masa and ism metadata | |
| WO2022223133A1 (en) | Spatial audio parameter encoding and associated decoding | |
| WO2024199801A1 (en) | Low coding rate parametric spatial audio encoding | |
| EP4690189A1 (en) | Spatial metadata direction harmonization | |
| US12400667B2 (en) | Spatial audio parameter encoding and associated decoding | |
| WO2023084145A1 (en) | Spatial audio parameter decoding | |
| WO2025078226A1 (en) | Parametric spatial audio decoding with pass-through mode | |
| WO2024175320A1 (en) | Priority values for parametric spatial audio encoding | |
| AU2023405231A1 (en) | Parametric spatial audio encoding | |
| EP4690188A1 (en) | Coding of frame-level out-of-sync metadata | |
| WO2025223950A1 (en) | Signalling of pass-through mode in spatial audio coding | |
| EP4662656A1 (en) | Audio rendering of spatial audio |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250923 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| REG | Reference to a national code |
Ref country code: HK Ref legal event code: DE Ref document number: 40129269 Country of ref document: HK |