EP4662656A1 - Audio rendering of spatial audio - Google Patents
Audio rendering of spatial audioInfo
- Publication number
- EP4662656A1 EP4662656A1 EP24700972.3A EP24700972A EP4662656A1 EP 4662656 A1 EP4662656 A1 EP 4662656A1 EP 24700972 A EP24700972 A EP 24700972A EP 4662656 A1 EP4662656 A1 EP 4662656A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio signal
- input
- audio
- parameter
- spatial
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/01—Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/03—Application of parametric coding in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
- H04S7/304—For headphones
Definitions
- the present application relates to apparatus and methods for audio rendering of spatial audio and the application of regularization in the rendering, but not exclusively for configuring a mixing solution regularization for the rendering.
- Background There are many ways to capture spatial audio.
- One option is to capture the spatial audio using a microphone array, e.g., as part of a mobile device. Using the microphone signals, spatial analysis of the sound scene can be performed to determine spatial metadata in frequency bands. Moreover, transport audio signals can be determined using the microphone signals. The spatial metadata and the transport audio signals can be combined to form a spatial audio stream. Metadata-assisted spatial audio (MASA) is one example of a spatial audio stream.
- MSA Metadata-assisted spatial audio
- the MASA stream can, e.g., be obtained by capturing spatial audio with microphones of, e.g., a mobile device, where the set of spatial metadata is estimated based on the microphone signals.
- the MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 mix) or other content by means of a suitable format conversion. It is also possible to use MASA tools inside a codec for the encoding of multichannel channel signals by converting the multichannel signals to a MASA stream and encoding that stream.
- specific spatial audio microphones such as Ambisonics
- studio mixes e.g., 5.1 mix
- MASA tools inside a codec for the encoding of multichannel channel signals by converting the multichannel signals to a MASA stream and encoding that stream.
- a method for generating an audio signal comprising: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input property parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input property parameter; determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- the input audio signal may be a spatial audio signal, the spatial audio signal further comprising at least one spatial parameter associated with the at least two audio channels.
- Determining processing parameters based at least on the at least two audio channels may further comprise determining the processing parameters further based at least on the at least one spatial parameter.
- the at least one input property parameter may comprise at least one of: a bitrate associated with the at least one input audio signal; a codec format indicating an origin of the at least one input audio signal; and a configuration of the at least one input audio signal.
- the configuration of the at least one input audio signal may comprise at least one of: a source format indicating an original or input format from which the input audio signal was created; transport channel description; number of channels; channel distance; and channel angle.
- the codec format indicating the origin of the at least one input audio signal may comprise at least one of: a metadata assisted spatial audio stream origin codec format value indicating the at least one input audio signal represents a metadata assisted spatial audio signal; a multichannel audio stream origin codec format value indicating the at least one input audio signal represents a multichannel audio signal; an audio object codec format value indicating the input audio signal represents audio objects; and an Ambisonic format value indicating the input audio signal represents an Ambisonic audio signal.
- the at least one spatial parameter may comprise at least one of: information that describes an organization of sound in space with respect to the at least one input audio signal, the information comprising: a direction parameter configured to indicate from where the sound arrives; and a ratio parameter configured to indicate a portion of the sound that arrives from that direction; information that describes properties of an original multi-channel or multi-object sound scene, the information comprising at least one of: channel levels; object levels; inter-channel correlations; inter-object correlations; and object directions; processing coefficients related to obtaining spatial audio format signals based at least on the at least one input audio signal.
- the codec format indicating the origin of the at least one input audio signal may comprise an indication that the origin is unknown or undefined.
- Determining the at least one control parameter based at least on the at least one input property parameter may comprise: determining a first set of control parameter values based on a first one of the at least one input property parameter; and selecting one control parameter value from the first set of control parameter values based on a second one of the at least one input property parameter.
- the first one of the at least one input property parameter may be the codec format and the second one of the at least one input property parameter may be the bitrate.
- Determining processing parameters based at least on the at least two audio channels wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter may comprise: generating processing matrices based at least on the at least two audio channels wherein the generating of the processing matrices may be controlled based at least partially on the at least one control parameter.
- Generating the processing matrices controlled based at least partially on the control parameter may comprise regularizing the generation of the processing matrices based at least on the control parameter.
- Generating the audio signal based at least on the at least two audio channels and the processing parameters may comprise processing the at least two audio channels using the regularized processing matrices to generate the audio signal.
- Generating the processing matrices controlled based at least partially on the at least one control parameter may comprise generating entries of a diagonal matrix based at least on at least two control parameter values.
- Obtaining at least one input property parameter associated with the input audio signal may comprise deducing from the input audio signal the at least one input property parameter.
- Obtaining at least one input property parameter associated with the input audio signal may comprise receiving the at least one input property parameter as part of the input audio signal or as configuration information.
- Receiving the at least one input property parameter as part of the input audio signal may further comprise: receiving an encoded input property parameter as part of the input audio signal; and decoding the encoded input property parameter to obtain the input property parameter.
- Generating the audio signal based at least on the at least two audio channels and the processing parameters may comprise generating at least one of: a binaural audio signal; and a multichannel audio signal.
- an apparatus for generating an audio signal comprising means configured to: obtain an input audio signal comprising at least two audio channels; obtain at least one input property parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input property parameter; determine processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- the input audio signal may be a spatial audio signal, the spatial audio signal further comprising at least one spatial parameter associated with the at least two audio channels.
- the means configured to determine processing parameters based at least on the at least two audio channels further comprises may further be configured to determine the processing parameters further based at least on the at least one spatial parameter.
- the at least one input property parameter may comprise at least one of: a bitrate associated with the at least one input audio signal; a codec format indicating an origin of the at least one input audio signal; and a configuration of the at least one input audio signal.
- the configuration of the at least one input audio signal may comprise at least one of: a source format indicating an original or input format from which the input audio signal was created; transport channel description; number of channels; channel distance; and channel angle.
- the codec format indicating the origin of the at least one input audio signal may comprise at least one of: a metadata assisted spatial audio stream origin codec format value indicating the at least one input audio signal represents a metadata assisted spatial audio signal; a multichannel audio stream origin codec format value indicating the at least one input audio signal represents a multichannel audio signal; an audio object codec format value indicating the input audio signal represents audio objects; and an Ambisonic format value indicating the input audio signal represents an Ambisonic audio signal.
- the at least one spatial parameter may comprise at least one of: information that describes an organization of sound in space with respect to the at least one input audio signal, the information comprising: a direction parameter configured to indicate from where the sound arrives; and a ratio parameter configured to indicate a portion of the sound that arrives from that direction; information that describes properties of an original multi-channel or multi-object sound scene, the information comprising at least one of: channel levels; object levels; inter-channel correlations; inter-object correlations; and object directions; processing coefficients related to obtaining spatial audio format signals based at least on the at least one input audio signal.
- the codec format indicating the origin of the at least one input audio signal may comprise an indication that the origin is unknown or undefined.
- the means configured to determine the at least one control parameter based at least on the at least one input property parameter may be further configured to: determine a first set of control parameter values based on a first one of the at least one input property parameter; and select one control parameter value from the first set of control parameter values based on a second one of the at least one input property parameter.
- the first one of the at least one input property parameter may be the codec format and the second one of the at least one input property parameter may be the bitrate.
- the means configured to determine processing parameters based at least on the at least two audio channels wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter may be further configured to: generate processing matrices based at least on the at least two audio channels wherein the generating of the processing matrices may be controlled based at least partially on the at least one control parameter.
- the means configured to generate the processing matrices controlled based at least partially on the control parameter may be further configured to regularize the generation of the processing matrices based at least on the control parameter.
- the means configured to generate the audio signal based at least on the at least two audio channels and the processing parameters may be further configured to process the at least two audio channels using the regularized processing matrices to generate the audio signal.
- the means configured to generate the processing matrices controlled based at least partially on the at least one control parameter may be further configured to generate entries of a diagonal matrix based at least on at least two control parameter values.
- the means configured to obtain at least one input property parameter associated with the input audio signal may be further configured to deduce from the input audio signal the at least one input property parameter.
- the means configured to obtain at least one input property parameter associated with the input audio signal may be further configured to receive the at least one input property parameter as part of the input audio signal or as configuration information.
- the means configured to receive the at least one input property parameter as part of the input audio signal may further be configured to: receive an encoded input property parameter as part of the input audio signal; and decode the encoded input property parameter to obtain the input property parameter.
- the means configured to generate the audio signal based at least on the at least two audio channels and the processing parameters may be further configured to generate at least one of: a binaural audio signal; and a multichannel audio signal.
- an apparatus for generating an audio signal comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input property parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input property parameter; determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- the input audio signal may be a spatial audio signal, the spatial audio signal further comprising at least one spatial parameter associated with the at least two audio channels.
- the apparatus caused to perform determining processing parameters based at least on the at least two audio channels may further be caused to perform determining the processing parameters further based at least on the at least one spatial parameter.
- the at least one input property parameter may comprise at least one of: a bitrate associated with the at least one input audio signal; a codec format indicating an origin of the at least one input audio signal; and a configuration of the at least one input audio signal.
- the configuration of the at least one input audio signal may comprise at least one of: a source format indicating an original or input format from which the input audio signal was created; transport channel description; number of channels; channel distance; and channel angle.
- the codec format indicating the origin of the at least one input audio signal may comprise at least one of: a metadata assisted spatial audio stream origin codec format value indicating the at least one input audio signal represents a metadata assisted spatial audio signal; a multichannel audio stream origin codec format value indicating the at least one input audio signal represents a multichannel audio signal; an audio object codec format value indicating the input audio signal represents audio objects; and an Ambisonic format value indicating the input audio signal represents an Ambisonic audio signal.
- the at least one spatial parameter may comprise at least one of: information that describes an organization of sound in space with respect to the at least one input audio signal, the information comprising: a direction parameter configured to indicate from where the sound arrives; and a ratio parameter configured to indicate a portion of the sound that arrives from that direction; information that describes properties of an original multi-channel or multi-object sound scene, the information comprising at least one of: channel levels; object levels; inter-channel correlations; inter-object correlations; and object directions; processing coefficients related to obtaining spatial audio format signals based at least on the at least one input audio signal.
- the codec format indicating the origin of the at least one input audio signal may comprise an indication that the origin is unknown or undefined.
- the apparatus caused to perform determining the at least one control parameter based at least on the at least one input property parameter may be further caused to perform: determining a first set of control parameter values based on a first one of the at least one input property parameter; and selecting one control parameter value from the first set of control parameter values based on a second one of the at least one input property parameter.
- the first one of the at least one input property parameter may be the codec format and the second one of the at least one input property parameter may be the bitrate.
- the apparatus caused to perform determining processing parameters based at least on the at least two audio channels wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter may be further caused to perform: generating processing matrices based at least on the at least two audio channels wherein the generating of the processing matrices may be controlled based at least partially on the at least one control parameter.
- the apparatus caused to perform generating the processing matrices controlled based at least partially on the control parameter may be further caused to perform regularizing the generation of the processing matrices based at least on the control parameter.
- the apparatus caused to perform generating the audio signal based at least on the at least two audio channels and the processing parameters may be futher caused to perform processing the at least two audio channels using the regularized processing matrices to generate the audio signal.
- the apparatus caused to perform generating the processing matrices controlled based at least partially on the at least one control parameter may be further caused to perform generating entries of a diagonal matrix based at least on at least two control parameter values.
- the apparatus caused to perform obtaining at least one input property parameter associated with the input audio signal may be further caused to perform deducing from the input audio signal the at least one input property parameter.
- the apparatus caused to perform obtaining at least one input property parameter associated with the input audio signal may be further caused to perform receiving the at least one input property parameter as part of the input audio signal or as configuration information.
- the apparatus caused to perform receiving the at least one input property parameter as part of the input audio signal may further be caused to perform: receiving an encoded input property parameter as part of the input audio signal; and decoding the encoded input property parameter to obtain the input property parameter.
- the apparatus caused to perform generating the audio signal based at least on the at least two audio channels and the processing parameters may be further caused to perform generating at least one of: a binaural audio signal; and a multichannel audio signal.
- an apparatus for generating an audio signal comprising: means for obtaining an input audio signal comprising at least two audio channels; means for obtaining at least one input property parameter associated with the input audio signal; means for determining at least one control parameter based at least on the at least one input property parameter; means for determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- an apparatus for generating an audio signal comprising: obtaining circuitry configured to obtain an input audio signal comprising at least two audio channels; obtaining circuitry configured to obtain at least one input property parameter associated with the input audio signal; determining circuitry configured to determine at least one control parameter based at least on the at least one input property parameter; determining circuitry configured to determine processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for generating an audio signal to perform at least the following: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input property parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input property parameter; determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- a non-transitory computer readable medium comprising program instructions for causing an apparatus for generating an audio signal to perform at least the following: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input property parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input property parameter; determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- a computer readable medium comprising program instructions for causing an apparatus for generating an output audio signal to perform at least the following: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input property parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input property parameter; determining processing parameters based at least on the at least two audio channels, wherein the determining of the processing parameters is controlled based at least partially on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameters.
- An apparatus comprising means for performing the actions of the method as described above.
- An apparatus configured to perform the actions of the method as described above.
- a computer program comprising program instructions for causing a computer to perform the method as described above.
- a computer program product stored on a medium may cause an apparatus to perform the method as described herein.
- An electronic device may comprise apparatus as described herein.
- a chipset may comprise apparatus as described herein.
- Figures 1 to 3 shows schematically an example system of capture or otherwise obtaining spatial audio signals in the form of transport audio signals and spatial metadata
- Figure 4 shows schematically example system of encoding spatial audio signals in the form of transport audio signals and spatial metadata and playback of spatial audio signals suitable for implementing some embodiments
- Figure 5 shows schematically an example system of multiple operating mode based encoding spatial audio signals in the form of transport audio signals and spatial metadata and playback of spatial audio signals suitable for implementing some embodiments
- Figure 6 shows schematically an example playback apparatus suitable for implementing some embodiments
- Figure 7 shows schematically an example decoder as shown in Figure 5 according to some embodiments
- Figure 8 shows a flow diagram of the operation of the example decoder apparatus shown in Figure 7 according to some embodiments
- Figure 9 shows schematically an example spatial synthesizer as shown in Figure 5 according to some embodiments
- Figure 10 shows a flow diagram of the operation of the example spatial synthesizer shown
- Metadata-Assisted Spatial Audio is an example of a parametric spatial audio format and representation suitable as an input format for IVAS. It can be considered an audio representation consisting of ‘N channels + spatial metadata’. It is a scene-based audio format particularly suited for spatial audio capture on practical devices, such as smartphones. The idea is to describe the sound scene in terms of time- and frequency-varying sound directions and, e.g., energy ratios. Sound energy that is not defined (described) by the directions, is described as diffuse (coming from all directions).
- spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total ratio, spread coherence, distance, etc.) per time-frequency tile.
- the spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene.
- a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency element (also known as a time-frequency portion or time- frequency tile).
- the metadata comprises other parameters such as direct-to-total ratios, spread coherence, distance values.
- parametric spatial metadata representation can use multiple concurrent spatial directions. With MASA, the proposed maximum number of concurrent directions is two. For each concurrent direction, there may be associated parameters such as: Direction index; Direct-to-total ratio; Spread coherence; and Distance. In some embodiments other parameters such as Diffuse- to-total energy ratio; Surround coherence; and Remainder-to-total energy ratio are defined.
- the parametric spatial metadata values are available for each time- frequency tile (the MASA format defines that there are 24 frequency bands and 4 temporal sub-frames in each frame). The frame size in IVAS is 20 ms.
- Example metadata parameters can be: Format descriptor, this parameter defines the MASA format for IVAS. This may be stored as a 64 bit value and may be eight 8-bit ASCII characters: 01001001, 01010110, 01000001, 01010011,01001101, 01000001, 01010011, 01000001 where the values are stored as 8 consecutive 8-bit unsigned integers; Channel audio format, this parameter defines a combined following fields stored in two bytes where the value is stored as a single 16-bit unsigned integer; Number of directions, this parameter defines the number of directions which are described by the spatial metadata. Each direction can be associated with a set of direction dependent spatial metadata as described afterwards.
- Direction index this parameter defines a direction of arrival of the sound at a time-frequency parameter interval. Typically this direction of arrival is a spherical representation at about 1-degree accuracy;
- Direct-to-total energy ratio this parameter defines an energy ratio for the direction index (in other words defining an energy ratio associated with the direction for the time-frequency subframe);
- Spread coherence this parameter defines a spread of energy for the direction index (in other words defining a measure of the ‘width’ of the direction for the time-frequency subframe);
- Transport definition this describes the configuration of the two transport channels; Channel angle, this describes symmetric angle positions for transport signals with directivity patterns; Channel distance, this describes distance between the two transport channels; and Channel layout, when source format is multichannel, this describes the channel layout of the multichannel source format.
- the rendering method is based on multi-channel mixing.
- the method processes the given audio signals in frequency bands so that a desired covariance matrix is obtained for the output signal in frequency bands.
- the covariance matrix contains the channel energies of all channels and inter-channel relationships between all channel pairs, namely the cross-correlation and the inter-channel phase differences.
- This target covariance matrix is determined based on the energy of the input audio signals and the spatial metadata that the renderer receives.
- this rendering method first attempts to achieve the target covariance matrix by mixing the input audio signals.
- a regularization is applied in the processing. For example, if the input audio signals are highly correlated (in other words have highly similar signals), but the target covariance matrix is set to have low correlation, this would lead to processing gains that would amplify the incoherent signal portion excessively. If there was no regularization applied, this would lead to significant amplification of noises and other unwanted sounds in the input audio signals, causing bad audio quality.
- the target covariance matrix is not always achieved by mixing alone.
- some signal portions are not amplified as much as needed to reach the target, and therefore some of the expected or desired signal energy is missing.
- these sounds would effectively be attenuated in the reproduction, and the desired incoherence between the output channels may not be reached.
- a proposed solution to achieving the desired incoherence is to decorrelate the input signals to obtain incoherent signals and process these signals with mixing gains to obtain the covariance matrix of the missing signal portion.
- the target covariance matrix is obtained for the output signals also when the regularization has limited the processing. There are positives and negatives to both means, mixing and decorrelating.
- the decorrelation processing has a drawback that it affects the sound quality, especially when there is significant amount of decorrelated signal in the output.
- a sound When a sound is decorrelated its phase spectrum is modified at least to a degree, typically time-invariantly to preserve the tones. It is known that the perceived quality of certain sounds such as speech or applause will degrade in this process.
- mixing with too large gains causes problems such as excessive amplification of small signal components that could be in some situations mostly noise.
- the amount of regularization can be used to control how much decorrelated energy is mixed to the output. If only mild regularization is applied (i.e., the maximum allowed mixing gains are large), the output contains mostly a mixture of input signals without much decorrelated signals.
- the IVAS codec supports multiple input formats.
- the MASA format often originates from mobile devices which typically have inexpensive microphones, and thus there is typically perceivable microphone noise in the MASA input signals.
- significant amount of regularization is needed in the renderer to avoid amplification of the microphone noises.
- IVAS supports also multichannel inputs (such as 5.1), which are typically recorded and produced in a studio with professional microphones having less noise.
- a fixed setting for regularization is not optimum in this respect either, as different formats would require different values to achieve a range of good experiences.
- Known rendering methods may have fixed regularization with a value set for situations where the signal SNR characteristics are large.
- a fixed regularization value based on a high SNR situation may not produce acceptable results for low- SNR situations. For example a spatial audio codec operating at low bit rates, such as 32 kbps or below, may not benefit with a regularization value set for a high bit rate low-noise input signal.
- a fixed regularization value produces a rendering which is optimal only for certain type of input signals and/or for certain bit rates where the transported audio signals are encoded.
- Too much decorrelation produces a sound which can be perceived as reverberant and non-engaging, and amplification of noises produces a reproduced sound with perceived artefacts.
- the following embodiments and the concept as discussed in the application herein is one of rendering spatial audio from a parametric spatial audio stream (audio signal(s) and associated spatial metadata) or more generally at least one audio signal comprising two or more audio channels.
- a rendering method (and thus a renderer apparatus) is proposed that enables obtaining improved audio quality by controlling the regularization in the determination of rendering gains based on an input property (or input parameter) associated with the at least one audio signal (such as bitrate and/or codec format).
- the controlled regularization is aimed at minimizing artefacts due to decorrelation (added reverberance and loss of engagement) and excessive amplification of noises (perception of loud musical and other noises), thus improving the perceived audio quality by mitigating the perception of those artefacts.
- this can be achieved by obtaining the (parametric spatial) audio stream; obtaining the input property; determining a regularization value based on the input property; determining regularized rendering gains based on the regularization value, the audio signal(s), and the associated spatial metadata; rendering spatial audio signals (such as binaural audio signals) from the audio signal(s) using the regulated rendering gains.
- the input property can mean various things in different embodiments.
- the input property can be: the bitrate that was used for coding the parametric spatial audio stream (in case it was encoded/decoded before the rendering); the codec format, stating the origin of the stream, e.g., if it is a MASA stream (i.e., typically from a mobile device) or if it is created from multichannel signals (such as 5.1); the configuration of the audio signal(s) the renderer receives (e.g., if they are from cardioid or omnidirectional microphones).
- the embodiments discussed in further detail hereafter present systems utilizing a spatial sound renderer and controlling the regularization of the rendering gains based on at least one input property.
- the renderer is implemented within a decoder and two example input properties are considered.
- the input properties discussed herein are Codec format and Bitrate.
- the renderer can be implemented elsewhere than in a decoder, and there can be other input properties.
- examples the “Codec format” input property is described in further detail, and how signals of different codec formats can be obtained.
- the “Bitrate” input property is described in further detail.
- a codec format input property can indicate the type of a parametric spatial audio signal being transmitted.
- a parametric spatial audio signal is a signal that comprises one or more audio signals and associated spatial metadata.
- the spatial metadata has information indicating how the sound is spatially organized, and it is typically provided at a substantially lower sampling rate than the audio signal.
- Examples of spatial metadata in different “codec formats” include: information that describes the organization of the sound in space. For example, in frequency bands, a direction parameter may indicate from where the sound arrives, and a ratio parameter may indicate the portion of sound that arrives from that direction; information that describes the properties of an original multi-channel or multi-object sound scene. For example, channel or object levels and inter- channel or inter-object correlations, or object directions; processing coefficients related to obtaining certain spatial audio format signals (such as Ambisonic audio signals) based on the transmitted audio signals.
- the examples shown above can be extended with other types of spatial metadata also used.
- the codec format input property can thus indicate the type of the spatial metadata being conveyed.
- the codec format input property can also indicate (or alternatively indicate) the type or origin of the transport audio signals.
- the codec format input property can in some embodiments indicate that the transport audio signals are a downmix of multichannel signals, or that they have been captured with a microphone array, or that they are a decomposition of Ambisonic audio signals.
- the codec input property can have a defined value which could indicate that the origin of the transport audio signals is not known or undefined.
- two different codec formats may have mutually same type of spatial metadata or same type of transport audio signals.
- any of the abovementioned audio formats may be rendered to a multitude of output spatial audio formats, such as a multichannel loudspeaker output (e.g., 5.1), head-tracked or non-head-tracked binaural output, or cross-talk-cancel stereo.
- a multichannel loudspeaker output e.g., 5.1
- head-tracked or non-head-tracked binaural output e.g., 5.1
- cross-talk-cancel stereo e.g., stereo
- the following examples describe the synthesis of binaural audio signals.
- the examples described herein can be straightforwardly extended and are applicable to reproduction of other output types.
- the following embodiments are described with respect to the application of regularization in a renderer where the input is a parametric spatial audio, it would be understood that some embodiments can be applied to any suitable renderer and rendering operation where regularization is applied to the input audio to be rendered and the regularization value is controlled based on input parameters such as the bitrate and codec format
- FIG. 1 an example of apparatus configured to determine a parametric spatial audio signal from microphone array signals.
- This apparatus can, for example be implemented on a mobile phone comprising a microphone array (integrated within the mobile phone).
- the microphone array signals 100 can be forwarded to a microphone array frontend 101.
- the microphone array frontend can, for example, be implemented using methods such as presented in US10873814, and is configured to output transport audio signals 102 and spatial metadata 104.
- the transport audio signals 102 and the spatial metadata 104 can, for example be in the form of a MASA stream.
- the microphone array frontend 101 is configured to measure inter-microphone correlations in different delays between the microphones in frequency bands, and find the delay that maximizes the correlation, and determine the direction of the arriving sound based on that delay.
- the microphone array frontend can in some embodiments also determine the direct-to- total energy ratio parameter in frequency bands based on the correlation value.
- the microphone array frontend can also provide the transport audio signals 102. The process of determining the transport audio signals 102 depends on what type or format microphone array signals are input. For example, if the microphone array signals 100 are from a mobile device, the microphone array frontend 101 can be configured to select a microphone signal from the left side of the device as the left transport signal and another one from the right side of the device as the right transport signal.
- the microphone array frontend 101 can furthermore in some embodiments also apply any suitable pre-processing steps, such as equalization, microphone noise suppression, wind noise suppression, automatic gain control, beamforming and other spatial filtering, ambient noise suppression, and limiter.
- a microphone array providing microphone array signals could be a dedicated microphone array, for example a microphone array that provides first- order Ambisonic (FOA) signals at its output.
- the apparatus comprises an ambisonic decomposer 201 configured to receive the ambisonics signals 200 and generate transport audio signals 102 and spatial metadata 104.
- the ambisonic decomposer 201 is configured to determine spatial metadata.
- the spatial metadata could be determined, for example, using methods similar to Directional Audio Coding (DirAC) such as described in Pulkki, V. (2007). Spatial sound reproduction with directional audio coding. Journal of the Audio Engineering Society, 55(6), 503-516, and the transport audio signals 104 could be cardioid patterns towards left and right directions generated from the ambisonics (FOA) signals 200.
- DIAC Directional Audio Coding
- the transport audio signals 104 could be cardioid patterns towards left and right directions generated from the ambisonics (FOA) signals 200.
- FOA ambisonics
- the ambisonic decomposer 201 can be configured to receive an Ambisonics signals 200, for example FOA or higher-order Ambisonics (HOA), and determine spatial metadata 104 such that it enables reconstructing the Ambisonic signals in the decoder based on the transport audio signals.
- an Ambisonics signals 200 for example FOA or higher-order Ambisonics (HOA)
- the Ambisonic signal is conveyed as one or more channel transport audio signals 102, where the first channel is the omnidirectional W component, and the remaining channels are residual signals.
- the ambisonic decomposer 201 determines these signals by determining prediction coefficients that enable predicting the remaining channels from the W component, and the residual signals are then the prediction error.
- the downmixer and metadata determiner 301 is configured to receive the channel-based audio signals 300 and generate transport audio signals 102 and spatial metadata 104.
- Example channel-based audio signals 300 are surround 5.1 sound or audio object sounds.
- the downmixer and metadata determiner 301 is configured to generate the spatial metadata 104 by converting the audio object signals and/or audio channels to a FOA format, and then determining the spatial metadata using DirAC or similar means.
- the downmixer and metadata determiner 301 is configured to generate the transport audio signals 104 using amplitude panning so that the channels/objects beyond ⁇ 30 degrees (and the corresponding cone of confusion) are panned to the left and right channels fully, and the channels/objects in between ⁇ 30 degrees are panned to both channels depending on their direction.
- the codec format input parameter can thus indicate various aspects described in the foregoing. Depending on the use case, it can indicate an entirely different signal format (e.g., stereo transports vs main-residual signal definition), and/or different characteristics of the same signal format (e.g., stereo transport signals from downmix vs stereo transport signals from microphones), and/or it can indicate the metadata format.
- the codec format input parameter can be an index in a predefined list of options defining both the kind of transport audio signals and the spatial metadata being used.
- a bitrate input parameter can be configured to indicate the number of bits used to transmit information within a specific time frame.
- IVAS is expected to operate between bitrates 13.2 kbps and 512 kbps.
- Bitrate may also be constant or varying through time.
- each transmission frame may have equal number of bits or variable number of bits. In IVAS, the transmission frame is expected to be 20 ms and the bitrate is expected to be constant in steady state situations.
- bitrates would translate to 264 bits/frame and 10240 bits/frame. It should be noted that this is usually the total bitrate, which is then further distributed for signaling, audio channel coding, and, if present, metadata coding. This further division may again be constant or variable.
- bitrate is indicative of the overall quality of the transmitted audio.
- the overall quality of transmitted audio increases in lossy audio codecs, such as IVAS, when bitrate is increased. Depending on the codec format, this increase in quality may come from increased bits used for audio channel coding which directly decrease presence of coding artifacts, or it may come from increased bits in spatial metadata coding which increase the accuracy of the reproduced spatial scene.
- Figure 4 shows an example apparatus which comprises an encoder 401 configured to receive the transport audio signals 102 and spatial metadata 104 and encode this to form a bitstream 402. Furthermore the apparatus comprises a decoder 402 which is configured to receive the bitstream 402 and output the spatial audio output 404.
- the system can be considered to operate in a first mode where processing software external to the encoder 401 determines and provides the transport audio signals 102 and spatial metadata 104.
- processing software could be a microphone array frontend software optimized to process the microphone signals of a specific device.
- Figure 5 shows further example apparatus which comprises the encoder 401 configured to receive the transport audio signals 102 and spatial metadata 104 and encode this to form a bitstream 402. Furthermore the apparatus comprises a decoder 402 which is configured to receive the bitstream 402 and output the spatial audio output 404. Furthermore is shown that prior to the encoder 401 is an encoder preprocessor 501. The encoder preprocessor 501 is configured to receive the audio signals 500 and generate the transport audio signals 102 and the spatial metadata 104. In other words the system can be considered to operate in a second mode where the encoding system 511 that performs the encoding also performs also the encoder preprocessing to obtain the transport audio signals 102 and spatial metadata 104.
- an IVAS encoder (which comprises both an encoder preprocessor 501 and encoder 401) could be configured to accept 5.1 or Ambisonics input to generate the transport audio signals 102 and spatial metadata 104 that are then encoded.
- the encoder 401 thus forms a bitstream 402, which can be stored or transmitted.
- the bitstream contains the transport audio signals 102 and the spatial metadata 104 in an encoded form.
- the audio signals can, e.g., be encoded using an IVAS core codec, EVS, or AAC encoder (or any other suitable encoder), and the metadata can, e.g., be encoded using the methods presented in US20210295855, US20220343928, US20220036906, EP4091166 (and/or any other suitable methods).
- the encoder 401 can furthermore in some embodiments multiplex the encoded audio and encoded spatial metadata to form the bitstream 402.
- the encoder 401 is configured to write or include within the bitstream 402 the input parameters such as codec format input parameter being used. For example, this could be signalled directly with signalling bits which define the codec format input parameter.
- the codec format can be signalled based on two bits such that a value “00” is a multi-channel format, “01” is MASA format, and “10” is Ambisonics format (these are merely example values).
- the encoder 401 may also write to the bit stream 402 other input parameters such as the bitrate input parameter.
- the bitrate input parameter may be constant or vary across time.
- the bitrate input parameter may be inferred in the decoder from the size of the received bitstream frame.
- the input parameter information such as the codec format input parameter and bitrate input parameter information is provided through a different communication channel other than the bitstream 402 that conveys the transport audio signal and spatial metadata.
- the bitstream 402 can be forwarded to a decoder 403, which can, for example, be an IVAS decoder (or any other suitable decoder).
- the decoder 403 decodes the audio signals and the metadata, and renders the spatial audio output 404, which can, e.g., be binaural audio signals.
- the decoder 403 may be at a different or same device than the encoder 401.
- FIG. 6 With respect to Figure 6 is shown an example (decoder) apparatus for implementing some embodiments.
- a mobile phone 601 coupled via a wired or wireless connection 613 with headphones 619 worn by the user of the mobile phone 601.
- the wired or wireless connection 613 may enable audio signals 615 to be passed to the headphones 619 and a mono audio signal (or more than one audio signal) 617 from the headphones 619 (which can in some embodiments further include head orientation/position information metadata from the headphones).
- the example device or apparatus is a mobile phone as shown in Figure 6.
- the example apparatus or device could also be any other suitable device, such as a tablet, a laptop, computer, or any teleconference device.
- the apparatus or device could furthermore be the headphones itself so that the operations of the exemplified mobile phone 601 are performed by the headphones.
- the mobile phone 601 comprises a processor 603.
- the processor 603 can be configured to execute various program codes such as the methods such as described herein.
- the processor 603 is configured to communicate with the headphones 619 using the wired or wireless headphone connection 615.
- the wired or wireless headphone connection 615 is a Bluetooth 5.3 or Bluetooth LE Audio connection.
- the connection 615 provides from a processor 603 a (two-channel) audio signal to be reproduced to the user with the headphones 619.
- the headphones 619 could be over-ear headphones as shown in Figure 6, or any other suitable type such as in-ear, or bone-conducting headphones, or any other type of headphones.
- the headphones 619 have a head orientation sensor providing head orientation information to the processor 603 via the connection 613.
- a head-orientation sensor is separate from the headphones 619 and the data is provided to the processor 603 separately.
- the head orientation is tracked by other means, such as using the device 601 camera and a machine-learning based face orientation analysis.
- the processor 603 is coupled with a memory 605 having program code 607 providing processing instructions according to the following embodiments.
- the program code 607 has instructions to process the transport audio signals and spatial metadata received by the transceiver 611 or retrieved from the storage 609 to a rendered form suitable for effective output to the headphones.
- the transceiver 611 can communicate with further apparatus by any suitable known communications protocol.
- the transceiver can use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or can be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad-hoc networks (MANETs), cellular internet of things (IoT) RAN and Internet Protocol multimedia subsystems (IMS), any other suitable option and/or any combination thereof.
- LTE Advanced long term evolution advanced
- NR new radio
- 5G long term evolution advanced
- UMTS universal mobile telecommunications system
- UTRAN or E-UTRAN
- the apparatus of Figure 6 may in some embodiments be also configured to implement the functions of the microphone array frontend 101, ambisonic decomposer 201, downmixer and metadata determiner 301, encoder 401 and encoder pre-processor 501 as well, for example, when the use case is two-way spatial audio communication with a remote apparatus.
- Figure 7 is shown a schematic view of an example decoder 403 as shown in Figure 5.
- the input bitstream 402 is forwarded to a demultiplexer and decoder 701.
- the demultiplexer and decoder 701 is configured to demultiplex and decode from the bitstream 402 the transport audio signals 706 and spatial metadata 708 from it.
- the decoding corresponds to the encoding applied in the encoder 401.
- the transport audio signals 706 and the spatial metadata 708 typically are not identical to the ones presented earlier as they have been encoded and decoded. Nevertheless, they are referred to using the same term (but with a different reference number) for simplicity.
- the demultiplexer and decoder 701 is configured to determine the bitrate 702 and the codec format 704 parameters based on the information within the bitstream 402. In one example, it could be that these are indicated by dedicated bits in the bitstream 402. In other examples, one or both can be detected from the bitstream, or they can be communicated through other channels, for example, via RTP or session control.
- the spatial metadata 708, transport audio signals 706, codec format 704, and bitrate 702 are then provided to the spatial synthesizer 703.
- the spatial synthesiser 703 is configured to synthesize the spatial audio output 404 in the desired format based on the spatial metadata 708, transport audio signals 706, codec format 704, and bitrate 702. In some examples only one of the codec format 704 and bitrate 702 parameters is used, and in other examples both. This information is used for balancing between using decorrelation and signal mixing in performing the spatial synthesis.
- the output provided by the spatial synthesiser 703 may, for example, be binaural audio signals.
- an example flow diagram shows the operations of the example decoder shown in Figure 7 according to some embodiments.
- the first operation can comprise as shown by 801, obtaining the (encoded spatial audio) bitstream.
- the (encoded spatial audio) bitstream is demultiplexed and decoded to generate transport audio signals, spatial metadata and input parameters such as codec format and bitrate.
- the spatial audio signals are synthesised from the transport audio signals based on the spatial metadata, and input parameters such as the codec format and bitrate.
- the spatial audio signals are output (for example binaural audio signals are output to the headphones).
- the spatial synthesiser 703 of Figure 7 is shown in further detail.
- the spatial synthesiser 703 is configured to receive the transport audio signals 706, the spatial metadata 708 and the input parameters as shown by the bitrate 702 and codec format 704.
- the spatial synthesiser 703 comprises a forward-filter bank 901.
- the transport audio signals 706 are provided to the forward-filter bank 901, which transforms the signals to a time-frequency representation.
- Any filter bank suitable for audio processing may be utilized, such as the complex-modulated quadrature mirror filter (QMF) bank, or a low-delay variant thereof, or the short-time Fourier transform (STFT).
- the forward- filter bank 901 can be implemented by any suitable time-frequency transformer.
- the filter bank is configured to have 60 frequency bins, and sufficient stop-band attenuation to avoid significant aliasing to occur when the frequency bin signals are processed.
- all frequency bins can be processed independently from each other, except that some frequency bins share the same spatial metadata.
- the spatial metadata may comprise spatial parameters in a limited number of frequency bands, for example 5, 12, or 24 bands, and each of these bands correspond to a set of one or more frequency bins provided by the forward filter bank 901.
- the output of the forward filter-bank 901 are time-frequency transport signals 906, which are provided to a decorrelator and mixer 907, a processing matrices determiner 905, and an input and target covariance matrix determiner 909.
- the transport audio signals have exactly two channels, but could be more than two in other examples.
- the spatial synthesiser 703 comprises an input and target covariance matrix determiner 909.
- the input and target covariance matrix determiner 909 is configured to receive the spatial metadata 708 and the time- frequency transport signals 906 and is configured to determine covariance matrices 906.
- the covariance matrices 906 comprise an input covariance matrix representing the time-frequency transport signals 906 and a target covariance matrix representing desired time-frequency spatial audio signals (that are to be rendered).
- the input covariance matrix can be determined or measured from the time-frequency transport signals 906, denoted as a column vector ⁇ ( ⁇ , ⁇ ), where the row indicates the transport signal channel.
- the superscript H indicates a conjugate transpose and ⁇ ⁇ ( ⁇ ) and ⁇ ⁇ ( ⁇ ) are the first and last time-frequency signal temporal indices corresponding to frame ⁇ (or sub-frame ⁇ in some embodiments). In this example, there are four time indices ⁇ at each frame ⁇ .
- the covariance matrix is determined for each bin. In other embodiments, the covariance matrix could be also averaged (or summed) over multiple frequency bins, in a resolution that approximates human hearing resolutions, or in the resolution of the determined spatial metadata parameters, or any suitable resolution.
- the target covariance matrix can be determined based on the spatial metadata and the overall signal energy.
- the overall signal energy ⁇ ⁇ ( ⁇ , ⁇ ) can be obtained as the mean of the diagonal values of ⁇ ⁇ ( ⁇ , ⁇ ) .
- the spatial metadata comprises the direction parameters azimuth ⁇ ( ⁇ , ⁇ ) and elevation ⁇ ( ⁇ , ⁇ ) and a direct-to-total ratio parameter ⁇ ( ⁇ , ⁇ ).
- the band index ⁇ is the one where the bin ⁇ resides.
- the target covariance matrix is where ⁇ , ⁇ ( ⁇ , ⁇ ) , ⁇ ( ⁇ , ⁇ ) ⁇ is a head-related transfer function column vector for bin ⁇ , azimuth ⁇ ( ⁇ , ⁇ ) and elevation ⁇ ( ⁇ , ⁇ ) , and it is a column vector of length two with complex values, where the values correspond to the HRTF amplitude and phase for left and right ears.
- the HRTF values may be also real because phase differences are not needed for perceptual reasons at high frequencies. Obtaining HRTFs for a given direction and frequency is known and any suitable method applied to obtain them.
- ⁇ ⁇ ( ⁇ ) is the diffuse field binaural covariance matrix, which can be determined for example in an offline stage by taking a spatially uniform set of HRTFs, formulating their covariance matrices independently, and averaging the result.
- the input covariance matrix ⁇ ⁇ ( ⁇ , ⁇ ) and the target covariance matrix are then output as covariance matrices 906 to the processing matrices determiner 905.
- the above example considered only directions and ratios.
- the procedure generating a target covariance matrix has been detailed also more broadly in WO2019086757A1 where additionally to the directions and ratios, using also spatial coherence parameters was also described, and furthermore, other output types than binaural output were also covered.
- the spatial synthesiser 703 comprises a regularization factor determiner 903.
- the regularization factor determiner 903 is configured to receive the input parameters, such as the codec format 704 and the bitrate 702 and determine a regularization factor ⁇ ( ⁇ ) 904.
- the regularization factor determiner 903 is configured to output the regularization factor ⁇ ( ⁇ ) 904 to the processing matrix determiner 905.
- the spatial synthesiser 703 comprises a processing matrix determiner 905.
- the processing matrix determiner 905 is configured to receive the covariance matrices 906 and the regularization factor ⁇ ( ⁇ ) 904 and determines processing matrices 908 ⁇ ( ⁇ , ⁇ ) and ⁇ ⁇ ( ⁇ , ⁇ ) .
- the method further uses a prototype matrix which is a matrix that informs the optimization procedure which kind of signals generally are meant for each output (with a constraint that the output must attain the target covariance matrix).
- the prototype matrix ⁇ can, for example, simply be or ⁇ 1 processing is head-tracked binaural, then when the user is 0.05 facing rear directions the transport audio signals may be processed (for example right after or before the forward filter-bank 901) so that they replace mutually each when the user is facing rear directions.
- the embodiments herein implements regularization based on ⁇ ( ⁇ ) in determining the processing matrices ⁇ ( ⁇ , ⁇ ) and ⁇ ⁇ ( ⁇ , ⁇ ) based on ⁇ ⁇ ( ⁇ , ⁇ ) , ⁇ ⁇ ( ⁇ , ⁇ ) and ⁇ ( ⁇ ).
- the derivation of the equations to determine the processing matrices is thoroughly explained in Vilkamo, J., Biffström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time–frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. Note that this implementation is an example implementation, and some operations can be performed in other ways to achieve a same or similar result.
- the example implementation provided in the appendix of Vilkamo, J., Biffström, T., & Kuntz, A. (2013) Optimized covariance domain framework for time–frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403- 411, can form the basis of an implementation for determining processing matrices with the difference being that the regularization factor ⁇ ( ⁇ ) is employed in place of the fixed number 0.2 in the program code line 21 (of the example implementation).
- the use of the regularization in the determination of the mixing matrices is here explained for completeness. Note that the notation for the matrices below is not exactly the same as in Vilkamo, J., Biffström, T., & Kuntz, A. (2013).
- ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ where the square root is an entry-wise operation and notation x,y indicates either of the input or target covariance matrix.
- ⁇ approaches 1 then ⁇ ⁇ , ⁇ , ⁇ approaches a matrix with all diagonal values being the same as the largest diagonal value of ⁇ ⁇ , ⁇ . This is the maximum regularization.
- There may be a non-zero missing covariance matrix ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ which may be called a residual covariance matrix.
- the same method as presented in the reference cited previously may be used to generate another processing matrix ⁇ ⁇ that is used to process the decorrelated version of the transport audio signals to attain properties of that missing portion, in other words, to attain the covariance matrix ⁇ ⁇ .
- This may be implemented by setting ⁇ ⁇ as the target covariance matrix, removing non-diagonal elements (because the sound is decorrelated) of ⁇ ⁇ and using that as the input covariance matrix, and otherwise performing the same operations as described to obtain ⁇ ⁇ .
- the function of the regularization value is not of significant importance at this stage, because the input is incoherent. Thus, for example, a fixed value of 0.2 can be used.
- the processing matrices determiner 905 can then be configured to output the processing matrices ⁇ ( ⁇ , ⁇ ) and ⁇ ⁇ ( ⁇ , ⁇ ) 908, which have been regularized based on the regularization factor ⁇ ( ⁇ ) 904.
- the spatial synthesiser 703 comprises a decorrelator and mixer 907.
- the decorrelator and mixer 907 is configured to receive the time- frequency transport signals ⁇ ( ⁇ , ⁇ ) 906 and the processing matrices ⁇ ( ⁇ , ⁇ ) and ⁇ ⁇ ( ⁇ , ⁇ ) 908.
- the decorrelator and mixer 907 is first configured to process the time- frequency transport signals 906 with decorrelators to generate decorrelated signals ⁇ ⁇ ( ⁇ , ⁇ ) . Then the decorrelator and mixer 907 is configured to apply the following mixing procedure to generate the time-frequency spatial audio signals ⁇ ( ⁇ , ⁇ ) 910 which then can be output by the decorrelator and mixer 907.
- the processing matrices may be linearly interpolated between frames ⁇ such that at each temporal index of the time-frequency signal the matrices take a step from ⁇ ( ⁇ , ⁇ ⁇ 1) towards ⁇ ( ⁇ , ⁇ ).
- the interpolation rate may be adjusted if an onset is detected (fast interpolation) or not (normal interpolation).
- the spatial synthesiser 703 comprises an inverse filter-bank 911.
- the inverse filter-bank 911 applies an inverse transform corresponding to that used by the forward filter-bank 901 to convert the time- frequency spatial audio signals 910 to spatial audio output 404, which is the output of the system.
- the first operation can comprise as shown by 1001, obtaining transport audio signals, spatial metadata and the input parameters such as codec format and bitrate.
- the transport audio signals are time-frequency transformed to generate time-frequency transport audio signals.
- the time-frequency transport audio signals can then be used to generate the input covariance matrix. Having determined the input covariance matrix then this can be used to determine an overall energy.
- the target covariance matrix can be generated.
- the generation of the input and the target covariance matrices is shown by 1007.
- the regularization factor is determined based on the input parameters such as the codec format and the bitrate as shown by 1005.
- the processing matrices are determined based on: the input covariance matrix; target covariance matrix; and regularization factor as shown by 1009.
- the time-frequency transport signal is decorrelated and mixed, based on the processing matrices to generate a time-frequency spatial audio signals as shown by 1011.
- the time-frequency spatial audio signals are inverse time-frequency transformed to generate spatial audio signals (for example using the inverse filter bank) as shown by 1013.
- the spatial audio signals can be output as the spatial audio output as shown by 1015.
- FIG 11 operations of the example regularization factor determiner 903 as shown in Figure 9.
- the regularization factor determiner is presented in the context of the above example spatial audio codec system.
- similar process can be applied in any context where a regularization factor or similar mixing or amplification limiting factor is used to control automatic mixing of input channels to output channels, and a constant value does not produce optimal quality.
- the initial operations are those of obtaining the input parameters affecting the regularization factor.
- this can be obtaining the codec format, as shown by 1101, which can describe the spatial audio format of the codec that is passed to the renderer. For example, this could be MASA format, premixed multi-channel format, Ambisonics format, etc. Additionally in some embodiments this can also comprise, as shown by 1105, obtaining a bitrate input parameter which describes the bitrate that was used for encoding the bitstream.
- the regularization factor itself can, in some embodiments, be defined as a decimal value in range [0,1] where value 1 restricts the rendering gains the most and value 0 does not restrict them at all.
- the codec format parameter is used, as shown by 1103, to select an initial set of available values for the regularization factor to be used in following operations steps.
- this could be implemented in a form of a number of tables where each table corresponds to a specific codec format.
- the table for MASA format could be [1.0, 0.8, 0.5, 0.2] and a similar table for Ambisonics format could be [1.0, 0.7, 0.4, 0.2].
- a similar example table can be [0.8, 0.5, 0.3, 0.2] as this content is usually professionally produced content and compression of the audio is usually less prone to compression artifacts.
- the use of tables is only one example implementation option and other equivalent or even more suitable ways to implement the selection can be created.
- the set of values are taken, and based on the obtained codec bitrate, a selection from the initial set of values an initial value for the regularization factor is selected. For example taking the aforementioned MASA format table, specific bitrates are associated to each table value.
- an example rule set can be: - If bitrate ⁇ 64 kbps, then use value 1.0 - If bitrate > 64 kbps and bitrate ⁇ 96 kbps, then use value 0.8 - If bitrate > 96 kbps and bitrate ⁇ 160 kbps, then use value 0.5 - If bitrate > 160 kbps, then use value 0.2
- an alternative way of implementing this determination would be to have a defined value in the set of values for each possible bitrate. Any suitable method can be employed which allows obtaining an initial regularization factor for a given bitrate and codec format combination.
- the final used values are usually obtained using a set of objective measures and/or subjective listening tests.
- the given example values are suitable, but they may not produce best quality in all situations and it is very much expected that various different sets of values can be used within the context of these embodiments.
- a single regularization factor is defined.
- the regularization is implemented by controlling the entries of a diagonal matrix based on this regularization factor. In some other embodiments, more fine-grained tuning of the regularization could be applied, for example, so that different regularization factors or different regularization schemes are used for different entries of that diagonal matrix.
- the regularization information may be not a single value but a more elaborate set of information.
- the above description shows the invention in the context of binaural rendering.
- Journal of the Audio Engineering Society, 61(6), 403-411 provides a general mixing framework and can be applied to any input-output mixing cases.
- the target output could be loudspeaker, Ambisonics, or any other audio channels in addition to binaural target output.
- the example embodiments describing adapting the regularization factor can be suitably adapted for any other input-output mixing cases.
- the output format may also be used to determine the selection of the regularization factor. For example, some artifacts may be more noticeable when using loudspeaker rendering than they are when using binaural rendering. This is turn suggests different regularization factor values for them.
- additional adjustments steps may be done for the regularization factor after an initial value has been selected based on the input parameters such as codec format and bitrate. These adjustments may change the regularization factor by absolute (e.g., +0.1 or -0.1) or relative (e.g., multiply with 0.9) steps.
- the source information for these adjustment steps can have many options.
- the source information can be one or more of the following: Descriptive metadata of MASA format or any similar data.
- the source format can distinguish between channel-based audio and microphone captures which may use different tables as described above.
- transport channel configuration may be used to perform adjustments as, for example, cardioid pattern transports may be more sensitive to artifacts than omnidirectional type transports.
- the order of selecting the regularization factor in the presented method could be reordered with trivial effort as the resulting regularization factor is simply a combination of multiple factors.
- the codec format can be signalled as a combination of bits indicating the kind of spatial metadata and/or transport audio signal used instead of directly signalling a format. It should be noted that although the above description implies regularization to be a control between mixing original signals and adding decorrelated versions of them to achieve the target covariance matrix, it is not mandatory to have decorrelated signals present at all.
- the presented method can be also used in a situation where only the mixing of original signals solution is used. In this case, target covariance matrix is not always achieved.
- the covariance-matrix-based rendering as presented in the reference above was used as an example.
- the presented methods can also be used with other kind of renderers, as long as there is some kind of regularization for the gains.
- the “direct” and the “ambient” sound may be rendered separately. In that case, it may be possible that prototype signals are created separately for each of them (using, e.g., decorrelation for the ambient part), and the target energies are also created separately for the direct and the ambient parts.
- gains are computed for the direct and the ambient part rendering. These gains typically need regularization to avoid excessive gaining of noises, as presented above for the main embodiment. These gains can be limited using the present invention. It should be noted that in this embodiment the regularization does not balance between mixing and decorrelation, but it balances between having artefacts from excessive gaining of noises and attenuation of some signals.
- Figure 12 is shown a processing example according to some embodiments. The processing is with a system such as described herein, where an original 5.1 sound is conveyed over two transport audio channels, and at a decoder the audio is rendered to a binaural sound using the transmitted spatial metadata.
- the spatial metadata and the audio signals are encoded with two different bitrates, 48 kbps and 256 kbps, represented by the two columns 48 kbps 1200 and 256 kbps 1202.
- the top row 1201 shows the binaural processed sound, where the middle 1203 and bottom 1205 rows show only the decorrelated path sound, in other words, where after formulating ⁇ ( ⁇ , ⁇ ) and ⁇ ⁇ ( ⁇ , ⁇ ) the first is set to zero prior to applying them to the signals.
- the middle row 1203 shows a prior art solution where the regularization factor is not adapted based on the bitrate, but instead, it is set to a default value 1.0, which is safer in terms of that it minimally amplifies the codec artefacts at the rendering.
- the bottom row 1205 shows the processing according to some embodiments where the regularization is adapted based on the bitrate.
- the regularization factor drops to 0.2, which causes the system to generate the required incoherence more with the mixing approach, and thus uses less decorrelation.
- the codec artefacts are of lower level and their amplification is not audible.
- better quality is provided, since less decorrelation is applied.
- the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
- aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
- the embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware.
- any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions.
- the software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
- the memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
- the data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
- Embodiments of the inventions may be practiced in various components such as integrated circuit modules.
- the design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc.
- the resultant design in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
- a standardized electronic format e.g., Opus, GDSII, or the like
- circuitry may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
- hardware-only circuit implementations such as implementations in only analog and/or digital circuitry
- software such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (
- circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
- non-transitory is a limitation of the medium itself (i.e., tangible, not a signal ) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Stereophonic System (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2301773.4A GB2626953A (en) | 2023-02-08 | 2023-02-08 | Audio rendering of spatial audio |
| PCT/EP2024/050876 WO2024165271A1 (en) | 2023-02-08 | 2024-01-16 | Audio rendering of spatial audio |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4662656A1 true EP4662656A1 (en) | 2025-12-17 |
Family
ID=89661084
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24700972.3A Pending EP4662656A1 (en) | 2023-02-08 | 2024-01-16 | Audio rendering of spatial audio |
Country Status (5)
| Country | Link |
|---|---|
| EP (1) | EP4662656A1 (en) |
| JP (1) | JP2026509128A (en) |
| CN (1) | CN120677524A (en) |
| GB (1) | GB2626953A (en) |
| WO (1) | WO2024165271A1 (en) |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP2838086A1 (en) * | 2013-07-22 | 2015-02-18 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | In an reduction of comb filter artifacts in multi-channel downmix with adaptive phase alignment |
| GB2556093A (en) | 2016-11-18 | 2018-05-23 | Nokia Technologies Oy | Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices |
| GB201718341D0 (en) | 2017-11-06 | 2017-12-20 | Nokia Technologies Oy | Determination of targeted spatial audio parameters and associated spatial audio playback |
| GB2575305A (en) | 2018-07-05 | 2020-01-08 | Nokia Technologies Oy | Determination of spatial audio parameter encoding and associated decoding |
| GB2577698A (en) | 2018-10-02 | 2020-04-08 | Nokia Technologies Oy | Selection of quantisation schemes for spatial audio parameter encoding |
| GB2587196A (en) | 2019-09-13 | 2021-03-24 | Nokia Technologies Oy | Determination of spatial audio parameter encoding and associated decoding |
| GB2592896A (en) | 2020-01-13 | 2021-09-15 | Nokia Technologies Oy | Spatial audio parameter encoding and associated decoding |
| GB2595475A (en) * | 2020-05-27 | 2021-12-01 | Nokia Technologies Oy | Spatial audio representation and rendering |
| GB2605190A (en) * | 2021-03-26 | 2022-09-28 | Nokia Technologies Oy | Interactive audio rendering of a spatial stream |
-
2023
- 2023-02-08 GB GB2301773.4A patent/GB2626953A/en not_active Withdrawn
-
2024
- 2024-01-16 JP JP2025546288A patent/JP2026509128A/en active Pending
- 2024-01-16 EP EP24700972.3A patent/EP4662656A1/en active Pending
- 2024-01-16 CN CN202480011718.2A patent/CN120677524A/en active Pending
- 2024-01-16 WO PCT/EP2024/050876 patent/WO2024165271A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| GB2626953A (en) | 2024-08-14 |
| WO2024165271A1 (en) | 2024-08-15 |
| JP2026509128A (en) | 2026-03-17 |
| CN120677524A (en) | 2025-09-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN114846542B (en) | Combination of spatial audio parameters | |
| CN114846541B (en) | Merging of spatial audio parameters | |
| US20260012742A1 (en) | Spatial Audio Representation and Rendering | |
| CN113597776A (en) | Wind noise reduction in parametric audio | |
| US20210250717A1 (en) | Spatial audio Capture, Transmission and Reproduction | |
| EP4627809A1 (en) | Binaural audio rendering of spatial audio | |
| US20250157475A1 (en) | Parametric spatial audio rendering | |
| WO2026057289A1 (en) | Spatial audio representation and rendering | |
| GB2582748A (en) | Sound field related rendering | |
| US20240137723A1 (en) | Generating Parametric Spatial Audio Representations | |
| WO2024175321A1 (en) | Diffuse-preserving merging of masa and ism metadata | |
| EP4662656A1 (en) | Audio rendering of spatial audio | |
| EP4312439A1 (en) | Pair direction selection based on dominant audio direction | |
| US20240236611A9 (en) | Generating Parametric Spatial Audio Representations | |
| WO2025201817A1 (en) | Rendering of a spatial audio stream | |
| US20240274137A1 (en) | Parametric spatial audio rendering | |
| GB2620593A (en) | Transporting audio signals inside spatial audio signal |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250908 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |