RENDERING OF A SPATIAL AUDIO STREAM Field The present application relates to apparatus and methods for rendering of a spatial audio stream, but not exclusively for gain factors in a combined format parametric spatial stream interactive audio rendering. Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the 3GPP Immersive Voice and Audio Services (IVAS) codec which is designed to be suitable for use over a communications network such as a 4G/5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR). This audio codec handles the encoding, decoding and rendering of speech, music and generic audio. It supports a variety of input formats, such as channel-based audio, object-based audio, and scene- based audio inputs including spatial information about the sound field and sound sources, as well as MASA (metadata-assisted spatial audio) inputs. IVAS operates with low latency to enable conversational services as well as supports high error robustness under various transmission conditions. The IVAS codec operates on a wide range bit rates from very low (13.4 kb/s) to relatively high bit rates (512 kb/s). Input signals can be presented to the IVAS encoder in one of a number of supported formats (and in some allowed combinations of the formats). For example, a mono audio signal (without metadata) may be encoded using an Enhanced Voice Service (EVS) encoder. Other input formats utilize new IVAS encoding tools. One input format proposed for IVAS is the Metadata-assisted spatial audio (MASA) format, where the encoder may utilize, e.g., a combination of mono and stereo encoding tools and metadata encoding tools for efficient transmission of the format. The use of “Audio objects”, or Independent streams with metadata (ISM), is another example of an input format proposed for IVAS. In this input format the scene is defined by a number (1 to N) of audio objects (where N is, e.g., 4). Each of the objects have an individual audio signal and some metadata describing its (spatial) features. The metadata may be a parametric representation of audio object
and may include such parameters as the direction of the audio object (e.g., azimuth and elevation angles). Furthermore, IVAS supports inputs in combined formats, e.g., the combined format input of audio object and MASA streams. This combination is referred to as OMASA. An example of such an input would be that the MASA format stream is obtained from the conferencing microphone in a room, capturing everything in the space, and the object format streams may be obtained using lapel microphones of the talkers. The combined encoding of the two input streams allows improvements in the quality of the coding. Being able to interact with an object or objects in the decoder side when employing a IVAS codec may be a desirable feature. For example, listener A may want to change a gain related to an audio object, whereas listener B may want to change the gain related to the same audio object. Furthermore listener B may want to change the main gain differently to the audio object or the gain of either the object relative to some other audio object. Furthemore in a combined format example, the listener may want to adjust the gain of the MASA component which may represent the signal background (and which may be relative to the object gain). Thus, rendering systems implementing codecs such as the above should be able to perform an interaction within the decoder/renderer so that each listener can have an individual experience. Summary There is provided according to a first aspect an apparatus for rendering a spatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the apparatus comprising means configured to: obtain the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtain gain control information; determine gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and render the spatial audio
signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. The gain control information may comprise at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion. The at least one gain for the at least one audio object signal may be based on a distance parameter associated with an at least one audio object of the at least one audio object portion. The means configured to determine the gain processing information based on the gain control information and the at least one audio object portion energy proportion may be configured to determine at least one first gain value based on the at least one audio signal and the at least one audio object portion energy proportion. The means may be further configured to apply the at least one first gain value to one of: the at least one audio object portion; and the at least one other spatial audio portion. The means configured to determine gain processing information may be further configured to determine gain processing information based on the at least one audio object portion position. The at least one other spatial audio portion may comprise at least two transport audio signals. The means configured to obtain the spatial audio stream may be configured to perform at least one of: receive information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receive at least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The means may be further configured to process the metadata associated with the at least one audio signal based on the gain control information. The means configured to process the metadata associated with the at least one audio signal based on the gain control information may be configured to at least one of: process metadata associated with the at least one audio object portion
of the at least one audio signal; and process metadata associated with the at least other spatial audio portion of the at least one audio signal. The at least one audio object portion and at least one other spatial audio portion may be a mixture within the at least one audio signal. According to a second aspect there is provided a method for an apparatus for rendering a spatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the method comprising: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information; determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. The gain control informationn may comprise at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion. The at least one gain for the at least one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion. Determining the gain processing information based on the gain control information and the at least one audio object portion energy proportion may comprise determining at least one first gain value based on the at least one audio signal and the at least one audio object portion energy proportion. The method may further comprise applying the at least one first gain value to one of: the at least one audio object portion; and the at least one other spatial audio portion. Determining gain processing information may further comprise determining gain processing information based on the at least one audio object portion position.
The at least one other spatial audio portion may comprise at least two transport audio signals. Obtaining the spatial audio stream may comprise at least one of: receiving information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receiving at least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The method may further comprise proessing the metadata associated with the at least one audio signal based on the gain control information. Processing the metadata associated with the at least one audio signal based on the gain control information may comprise at least one of: processing metadata associated with the at least one audio object portion of the at least one audio signal; and processing metadata associated with the at least other spatial audio portion of the at least one audio signal. The at least one audio object portion and at least one other spatial audio portion may be a mixture within the at least one audio signal. According to a third aspect there is provided an apparatus for rendering a spatial audio signal based on a spatial audio stream, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information; determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal.
The gain control information may comprise at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion. The at least one gain for the at least one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion. The apparatus caused to perform determining the gain processing information based on the gain control information and the at least one audio object portion energy proportion may be caused to perform determining at least one first gain value based on the at least one audio signal and the at least one audio object portion energy proportion. The apparatus may further be caused to perform applying the at least one first gain value to one of: the at least one audio object portion; and the at least one other spatial audio portion. The apparatus caused to perform determining gain processing information may further be caused to perform determining gain processing information based on the at least one audio object portion position. The at least one other spatial audio portion may comprise at least two transport audio signals. The apparatus caused to perform obtaining the spatial audio stream may be caused to perform at least one of: receiving information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receiving at least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The apparatus may be caused to further perform proessing the metadata associated with the at least one audio signal based on the gain control information. The apparatus caused to perform processing the metadata associated with the at least one audio signal based on the gain control information may be caused to perform at least one of: processing metadata associated with the at least one audio object portion of the at least one audio signal; and processing metadata associated with the at least other spatial audio portion of the at least one audio signal.
The at least one audio object portion and at least one other spatial audio portion may be a mixture within the at least one audio signal. According to a fourth aspect there is provided an apparatus for rendering a spatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the apparatus comprising: means for obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; means for obtaining gain control information; means for determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and means for rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for rendering a spatial audio signal based on a spatial audio stream to perform at least the following: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information; determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for rendering a spatial audio signal based on a spatial audio stream to perform at least
the following: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information; determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to a seventh aspect there is provided an apparatus for rendering a spatial audio signal based on a spatial audio stream, the apparatus comprising: obtaining circuitry configured to obtain the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining circuitry configured to obtain gain control information; determining circuitry configured to determine gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering circuitry configured to render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to an eighth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for rendering a spatial audio signal based on a spatial audio stream to perform at least the following: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information; determining gain processing information based on: the gain control information; and the at least one
audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically a system of apparatus suitable for implementing some embodiments; Figure 2 shows a flow diagram of the operation of the apparatus shown in Figure 1 according to some embodiments; Figure 3 shows schematically an example of the encoder as shown in Figure 1 according to some embodiments; Figure 4 shows a flow diagram of the operations of the example encoder shown in Figure 3 according to some embodiments; Figure 5 shows schematically an example of the decoder as shown in Figure 1 according to some embodiments; Figure 6 shows a flow diagram of the operations of the example decoder shown in Figure 5 according to some embodiments; Figure 7 shows schematically an example of the spatial synthesizer as shown in Figure 5 according to some embodiments;
Figure 8 shows a flow diagram of the operations of the example spatial synthesizer shown in Figure 7 according to some embodiments; and Figure 9 shows schematically an example device suitable for implementing the apparatus shown herein. Embodiments of the
The concept as discussed herein in further detail with respect to the following embodiments is related to parametric spatial audio rendering. In the following examples an IVAS codec is used to show practical implementations or examples of the concept. However, it would be appreciated that the embodiments presented herein may be extended to other codecs without inventive input. The concept as discussed herein in further detail in the following embodiments is one of providing an ability to modify or edit the gains of different audio objects and of a MASA part within a suitable decoder/renderer for a combined spatial audio format such as combined object and MASA formats (OMASA). As discussed earlier the metadata-assisted spatial audio (MASA) is one of the input formats supported by IVAS. It uses audio signal(s) together with corresponding spatial metadata (containing, e.g., directions and direct-to-total energy ratios in frequency bands). The MASA stream can, e.g., be obtained by capturing spatial audio with microphones of, e.g., a mobile device, where the set of spatial metadata is estimated based on the microphone signals. The MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 multichannel mix) or other content by means of a suitable format conversion. MASA spatial metadata values are available for each time-frequency tile (TF-tile) (there can, for example, be 24 frequency bands and 4 temporal sub-frames in each frame). The frame size in IVAS is 20 ms (and thus the temporal sub-frame is 5 ms). In addition, MASA supports 1 or 2 directions for each time-frequency tile (i.e., there are 1 or 2 direction index, direct-to-total energy ratio, and spread coherence parameters for each time-frequency tile). Additionally IVAS supports also audio objects (Independent streams with metadata, ISM) as an input. The audio objects contain for each object an audio signal and associated metadata (e.g., the direction of the object).
Furthermore, IVAS supports a combined format input of audio object and MASA streams. This combination is referred to as OMASA. The combined encoding of the two input streams allows improvements in the quality of the coding. For example, at certain bitrates with a certain number of objects (e.g., 3-4 objects at 64 kbps), the audio signals of some of the objects and the audio signals of the MASA stream are combined to common stereo transport audio signals, which are then encoded. As fewer audio signals are encoded than there were in the input, significant bitrate savings can be made. As a result, the effective bitrate per encoded audio channel is higher than in case all the audio signals would have been separately encoded leading to improved audio quality There have been discussed methods, such as UK patent application GB2104309.6 where a position of an object in a renderer or decoder can be changed for OMASA audio format signals in an IVAS codec. However, such an approach involved rotation only and did not propose methods for enabling changing the loudness (due to, e.g., distance or level) of an object. Spatial Audio Object Coding (SAOC) as described in Herre, J. et al. (2012). “MPEG spatial audio object coding—the ISO/MPEG standard for efficient coding of interactive audio scenes”, Journal of the Audio Engineering Society, 60(9), 655-673 describes encoding objects as a downmix and spatial metadata and then decoding them while allowing rendering-time spatial adjustments. Audio objects are downmixed for example to a stereo track and a set of spatial metadata is extracted in time-frequency regions. This metadata comprises Object level differences (OLD), inter-object cross coherences (IOC), downmix channel level differences (DCLD), downmix gains (DMG) and object energies (NRG). Such a set of metadata provides the information to manipulate or render the multi-object mixture to a spatialized output. However, the metadata involved, and therefore the techniques related do not provide means to account for mixtures that prominently have non- object content. Similarly, MPEG-H methods such as described in Herre, J., Hilpert, J., Kuntz, A., & Plogsties, J. (2015), ”MPEG-H 3D audio—The new standard for coding of immersive spatial audio”, IEEE Journal of selected topics in signal processing, 9(5), 770-779 applies the methods as described in context of SAOC, however in an extended form that is referred to as SAOC-3D (such as described in
Murtaza, A., Herre, J., Paulus, J., Terentiv, L., Fuchs, H., & Disch, S. (2015, October), “ISO/MPEG-H 3D audio: SAOC 3D decoding and rendering”, Audio Engineering Society Convention 139. Audio Engineering Society. These describe a system having audio objects and discrete surround audio channels at the same downmix stream, and effective decoding of them. However, this extension applies only to object-only coding as the original audio channels are conceptually close to static audio objects. In other words, SAOC-3D does not provide means to account for mixtures that prominently have non-object content, where “non-object content” is understood more broadly than loudspeaker channels, i.e., spatially static audio objects. The concept as discussed in the embodiments herein expands on these implementations in that the transport audio signals do not only include audio objects, but also audio signals (and associated spatial metadata) from other sources. The other sources may, for example, be spaced microphone captured audio signals (which could be captured using a mobile device), downmixed 5.1 audio signals, or any other suitable audio input format. The embodiments herein thus improve on the methods disclosed by SAOC in being configured to provide means to effectively handle gain based modifications (for example, object distance modification) of such signals. In particular in certain coding modes for the object and MASA combined input (OMASA) in IVAS, at least some of the object audio signals and the MASA audio signals are mixed together to a common stereo downmix signal. This enables obtaining higher bitrate per encoded audio channel, and thus improved audio quality. However, as a result of mixing the audio signals together, the object audio signals, nor the MASA audio signals, are no longer individually available in the decoder. This is not a problem, if the objects and the MASA stream are to be rendered according to the original input metadata without any modifications. However, the embodiments discussed herein enable the listener to modify the sound scene in some way before the rendering. This can be applied to OMASA (objects + MASA) but can also be applied to other combinations such as objects + Ambisonics and objects + multi-channel signals (e.g., 5.1 or 7.1+4). For example, the listener often wants to change the loudness of certain objects. For example, the objects may contain speech of different participants in a
teleconference. In such a case, the listener may want to make a selected participant louder, and some other participant softer. The embodiments as described herein enable such a scenario to be implemented. Moreover, in some circumstances the listener may wants to change the loudness of the MASA part. For example, a recording of a call at a live concert can result in the MASA part may contain some background sounds (e.g., a live concert) and the object part may contain speech of the caller. In this case, the listener at the other end may want to make the MASA part louder to hear the concert better, or alternatively may want to make the MASA part softer, to hear the speech of the caller better. The embodiments as discussed in detail hereafter enable this to be implemented. The examples presented herein thus enable changing the loudness of the different objects or the different audio streams despite not having independent audio signals for the objects and the MASA parts (for example, in certain coding modes of the OMASA input in IVAS where there are 3-4 objects encoded at 64 kbps). In some embodiments the changing of the loudness of the objects and/or the MASA part may be performed manually or automatically. For example, a listener or user may adjust the loudness of the different participants using a suitable user interface (UI) on, for example, a mobile phone app. This may be performed, in some embodiments, during a call. Alternatively, in some embodiments the loudness adjustment may be performed automatically, or semi-automatically, for example, based on the preferences of the user in the past manual adjustments. Hence in summary there is presented an interactive rendering of a parametric spatial audio stream, which contains at least one audio signal and associated spatial metadata, and where the audio signal is a mixture of audio objects and other audio content (e.g., MASA). The apparatus and methods as discussed herein enable the modification of the loudness of individual audio objects and the other audio content at the rendering and enabling, for example, different listeners having an individual balance between the different objects and the other spatial audio content. This can be achieved by determining information of the proportion (or energetic information or values) of the audio objects and the other content (for
example, in a time-frequency domain representation) as well as their spatial properties. Additionally there can be a determination of target proportions (or target energetic information or values) for the different audio objects and the other audio content based on the at least one audio signal, the spatial metadata and the values relared to the desired level modification. Furthermore there can be a modification of the at least one audio signal based on the determined energetic values and the determined target energetic values. Then there can be a modification of the spatial metadata based on the determined energetic values and the determined target energetic values. After this there can be a rendering of the spatial audio (for example, binaural audio signals) using the modified at least one audio signal and the modified spatial metadata. In such embodiments the values related to the desired level modification may be based on, for example, the desired levels, desired level changes, and/or the desired distances of the audio objects. The energetic values can furthermore be related, for example, to the energies or the levels of the objects and the other audio content. The values can be absolute, or they can be relative. The energetic values may also be called level values. The other spatial audio content can in some embodiments be in any suitable format, for example, in MASA, Ambisonics, or multi-channel (e.g., 5.1 or 7.1+4). The mixture audio signal may, for example, have been created in an encoder, where the audio signals from the objects and the other spatial audio content have been at least partially mixed together. Before discussing the concept in further detail we will initially describe in further detail some aspects of spatial encoding, decoding and reproduction which may be implemented in some embodiments. For example, with respect to Figure 1 is shown an example system suitable for implementing embodiments as described herein. The system comprises an encoder 101 which is configured to receive a number M of spatial audio signal streams. In Figure 1 is shown the spatial audio
stream 1104, spatial audio stream 2106, and spatial audio stream M 108 which is input to the encoder 101. The encoder 101 can in some embodiments comprise an IVAS encoder, though in other embodiments other suitable encoders can be employed. The spatial audio streams 104, 106, 108 can in some embodiments be different kind of streams. For example, the streams can be MASA streams, multichannel loudspeaker signal streams, and/or object streams. The encoder 101 is configured to generate an encoded bitstream 110. The encoded bitstream 110 in Figure 1 is shown being passed to a separate decoder 111. However in some embodiments the bitstream may be stored in a suitable storage medium for later retrieval. The system further comprises a decoder 111. The decoder 111 is configured to receive or retrieve the encoded bitstream 110. Additionally the decoder 111 is configured to receive gain control information 112. The decoder, which can be an IVAS decoder or any suitable format decoder (which matches the encoder) is configured to decode the bitstream 110 and render spatial audio signals 114 based on the gain control information 112. The gain control information 112 can, for example, comprise desired gains for the objects and/or the other streams (for example, the user of the decoder 111 apparatus may have set them). The spatial audio signals 114, in some embodiments, can be binaural audio signals. Figure 2 shows, for example, a flow diagram of the operation of the example system as shown in Figure 1. Thus the spatial audio streams are initially obtained as shown in step 201. Then the audio streams are encoded to generate the bitstream as shown in Figure 2 by step 203. The encoded bitstream is then transmitted to the decoder/received from the encoder (or stored/retrieved) as shown in Figure 2 by step 205. Additionally the gain control information is obtained as shown in Figure 2 by step 206. The encoded audio streams in the form of the encoded bitstream is then decoded and spatial audio signals are rendered based on the gain control information as shown in Figure 2 by step 207. Finally the spatial audio signals is output as shown in Figure 2 by step 209.
Figure 3 shows an example encoder 101 as shown in Figure 1 according to some embodiments. In this example, there are shown two input streams. The first input stream is a MASA stream, which comprises a MASA transport audio signals 302 and MASA metadata 300. The second input stream shown in Figure 3 is an object audio stream 320 (containing a number of, for example, N objects). It should be noted that in some embodiments there can be any suitable number of input streams and these two input streams are an example only. The encoder 101 comprises an object analyser 301. The object analyser 301 has an input which receives the object audio stream 320 and is configured to analyse the object audio stream 320 to determine which object to separate from the other objects. This can, for example, be performed in a manner as presented in WO2022/214730. The result of this analysis can be the separated object audio signal 316 and separated object metadata 314 which are outputted from the object analyser to an object encoder 309. The separated object metadata 314 can, for example, contain the object direction and the index of the object that was separated. In some embodiments the object encoder 309 is configured to receive the separated object audio signal 316 and separated object metadata 314 and encode these based on any suitable encoding mechanism. The resulting encoded separated object audio signal 326 and encoded separated object metadata 324 are output to a multiplexer or Mux 307. Additionally the remaining objects (from the object audio stream 320) are analyzed, to determine object transport audio signals 312 and object metadata 310. This remaining object analysis can be performed using any suitable method (for example, in a manner similar to that described in WO2022/200666). Furthermore, the metadata can contain any suitable metadata. As an example, the object transport audio signals 312 can be a stereo downmix using amplitude panning based on the object directions, and the object metadata 310 may contain the object directions and time-frequency domain object- to-total energy ratios (or, ISM ratios, in other words), which are obtained by analyzing the energies of the different objects in frequency bands and comparing them to the total energy of all objects in the corresponding bands and produce object transport audio signals 312 and object metadata 310.
The object transport audio signals 312 and the object metadata 310 can in some embodiments be passed to a metadata encoder 303 and the object transport audio signals 312 and MASA transport audio signals 302 passed to a transport audio signal combiner and encoder 305. In some embodiments the encoder 101 comprises a transport audio signal combiner and encoder 305. The transport audio signal combiner and encoder 305 is configured to obtain the MASA transport audio signals 302 and object transport audio signals 312 and combine and encode these inputs to generate encoded transport audio signals 306. The combination in some embodiments may be by summing them. In some embodiments, the transport audio signal combiner and encoder 305 is configured to perform other processing on the obtained transport audio signals or the combination of the transport signals. For example, in some embodiments the transport audio signal combiner and encoder 305 is configured to adaptively equalize the resulting signals in order to have the same energy in the time-frequency domain for the combined signals as the sum of the energies of the MASA and object transport audio signals. The encoding of the combined transport audio signals can employ any suitable codec. For example, in some embodiments the transport audio signal combiner and encoder 305 is configured to encode the combined transport audio signals using a EVS or AAC codec. In some implementations the transport audio signal combiner and encoder 305 is configured to encode the combined transport audio signals using a IVAS core coder. The encoded transport audio signals 306 can then be output to a multiplexer or Mux 307. In some embodiments the encoder 101 comprises a metadata encoder 303. The metadata encoder 303 is configured to receive the MASA metadata 300 and the object metadata 310 (in some embodiments the metadata encoder 303 is further configuered to receive the MASA transport audio signals 302 and the object transport audio signals 312). The metadata encoder 303 is configured to apply a suitable encoding to the metadata. The implementation of the metadata encoding may be any suitable encoding method, a few examples of which are described hereafter.
As a first example of metadata encoding, MASA-to-total energy ratios are determined using the MASA transport audio signals 302 and the object transport audio signals 312, for example, by computing the energies of the MASA and the object transport audio signals in time-frequency tiles, and then determining the MASA-to-total energy ratios by
where ^^^^^(^, ^) is the energy of the MASA transport audio signals 302 for the frequency band ^ and temporal subframe ^, and ^^^^(^, ^) the energy of the object transport audio signals 312. Then, the MASA metadata, the object directions, the ISM ratios, and the MASA-to-total energy ratios are encoded using any suitable methods (for example, using the methods presented in WO2020/089510, PCT/FI2019/050675, GB1811071.8, WO2020/193865, GB1913274.5, WO2022/200666, GB2217884.2, GB2217905.5, GB2217928.7, GB2217884.2). The resulting encoded metadata 304 is outputted from the block. The encoded transport audio signals 306, encoded metadata 304, encoded separated object audio signal 326, and encoded separated object metadata 324 are forwarded to the multiplexer or Mux 307, which multiplexes them to a bitstream 110, which is output from the encoder 101. Figure 4 shows a flow diagram of the operation of the example encoder as shown in Figure 3. As such in some embodiments the object audio streams are obtained as shown in Figure 4 by step 401. The object audio streams are then analysed to generate the object transport audio signals, object metadata, separated object audio signal and separated object metadata as shown in Figure 4 by step 403. Additionally the MASA transport audio signals are obtained as shown in Figure 4 by step 402. The MASA metadata is furthermore obtained as shown in Figure 4 by step 404.
Having obtained the MASA transport audio signals and the object transport audio signals these are combined and encoded to generate encoded combined transport audio signals as shown in Figure 4 by step 405. Having obtained the MASA metadata and the object metadata these are combined and encoded to generate encoded combined metadata as shown in Figure 4 by step 406. Furthermore the separated object audio signal and separated object metadata are encoded as shown in Figure 4 by step 407. Having generated the encoded separated object audio signal, the encoded separated object metadata, the encoded combined metadata and the encoded combined transport audio signals then these can be multiplexed as shown in Figure 4 by step 408. Then the bitstream (the multiplexed encoded signals) are output as shown in Figure 4 by step 409. Figure 5 shows an example decoder 111 as shown in Figure 1 according to some embodiments. In this example, there is shown the bitstream which is obtained by the decoder 111. The decoder 111 can in some embodiments comprise a demultiplexer or demux 501 which is configured to obtain the bitstream 110 and demultiplex the bitstream to generate: encoded metadata 502, which is passed to a metadata decoder and processor 503; encoded transport audio signals 512, which is passed to the transport audio signal decoder 513; and encoded separated object audio signal 522 and encoded separated object metadata 532 which are passed to the object decoder 523. Furthermore the decoder 111 can in some embodiments comprise a transport audio signal decoder 513 configured to receive the encoded transport audio signals 512. The transport audio signal decoder 513 can then be configured to decode the encoded transport audio signals 512 and generate decoded transport audio signals 514 which can be passed to a spatial synthesizer 505. The decoder 111 furthermore, in some embodiments, comprises a metadata decoder and processor 503 configured to receive the encoded metadata 502. The
metadata decoder and processor 503 furthermore is configured to decode and process the encoded metadata 502 and generate rendering metadata 504. As mentioned above, there various ways to encode the metadata, and also different possible sets of metadata transmitted. Hence, the decoding and processing implemented in some embodiments can vary. In some embodiments the encoded metadata 502 is decoded to generate decoded metadata. The decoded metadata can comprise the following parameters: decoded MASA metadata; MASA-to-total energy ratios; ISM ratios; and object directions. In some embodiments the decoded metadata parameters can be processed or converted to a form that is more suitable for rendering to generate the rendering metadata 504. For example firstly, the direct-to-total energy ratios ^^^^^(^, ^) in the MASA metadata are modified by multiplying them with the MASA-to-total energy ratio
The rest of the MASA metadata (directions ^^^^^^^(^, ^) , spread coherences ^^^^^(^, ^) , and surround coherences ^^^^^(^, ^) ) can be used without modifications. The ISM ratios ^^^^(^, ^, ^) can, in some embodiments, be modified by ^^^^,^^^^(^, ^, ^) = (1 − ^(^, ^)) ^^^^(^, ^, ^) where ^ is the object index. The object directions can be used without modifications (directions ^^^^^^(^, ^)). In some embodiments the processing can comprise a selection of one or more of the original decoded metadata parameters. The resulting rendering metadata 504 ( ^^^^^^^(^, ^) , ^^^^^,^^^^(^, ^) , ^^^^^(^, ^), ^^^^^(^, ^), ^^^^^^(^, ^), ^^^^,^^^^(^, ^, ^)) can then be provided to the spatial synthesizer 505 as an output of the metadata decoder and processor 503.
It should be noted that the rendering metadata 504 in some embodiments does not necessarily directly correspond to the original MASA metadata and object metadata (that were input to the metadata encoder as shown in the example encoder), as the original metadata was related to the separate transport audio signals, whereas the decoded metadata is related to the combined transport audio signals. Furthermore the generation of the rendering metadata may be implemented in the encoder 101 (as was mentioned above), or it may be performed elsewhere. The rendering metadata 504 can be passed to the spatial synthesizer 505. The encoded separated object audio signal 522 and encoded separated object metadata 532 are forwarded to the object decoder 523. The object decoder 523 is configured to decode the encoded separated object audio signal 522 and encoded separated object metadata 532 to generate decoded separated object audio signal 524 and decoded separated object metadata 534. The decoded separated object audio signal can comprise, for example, direction ^^^^^^(^) and the object index ^^^^(^). The decoded separated object audio signal 524 and decoded separated object metadata 534 are then passed to the spatial synthesizer 505. The decoder 111 in some embodiments comprises a spatial synthesizer 505. The spatial synthesizer 505 is configured to receive the rendering metadata 504, the decoded transport audio signals 514, the decoded separated object audio signal 524 and decoded separated object metadata 534 and the gain control information 112. The spatial synthesizer 505 can then be configured to generate the spatial audio signals 114 based on the rendering metadata 504, the decoded transport audio signals 514, the decoded separated object audio signal 524 and the decoded separated object metadata 534 and the gain control information 112. The gain control information 112, in some embodiments, comprises a gain ^^^^(^, ^) for each object and ^^^^^(^) for the MASA audio. The spatial audio signals 114 can then be output. With respect to Figure 6 is shown a flow diagram of the operations of the example decoder as shown in Figure 5. The bitstream is obtained as shown in Figure 6 by step 601.
The bitstream is then demultiplexed to generate the encoded metadata, the encoded transport audio signals, the encoded separated object audio signal and the encoded separated object metadata as shown in Figure 6 by step 603. The encoded transport audio signals are then decoded to generate the decoded transport audio signals as shown in Figure 6 by step 605. The encoded metadata furthermore is decoded and processed to generate the rendering metadata as shown in Figure 6 by step 606. The encoded separated object audio signal and encoded separated object metadata is decoded to generate decoded separated object audio signal and decoded separated object metadata as shown in Figure 6 by step 607. The gain control information is obtained as shown in Figure 6 by step 602. The spatial audio signals are generated from the decoded transport audio signals, decoded rendering metadata, decoded separated object audio signal, decoded separated object metadata and gain control information as shown in Figure 6 by step 608. Then the spatial audio signals are output as shown in Figure 6 by step 609. Figure 7 shows in further detail a schematic view of an example spatial synthesizer 505 as shown in Figure 5 according to some embodiments. The spatial synthesizer 505 is configured to receive the decoded transport audio signals 514, the gain control information 112, the rendering metadata 504, the decoded separated object audio signal 524 and decoded separated object metadata 534. In some embodments the spatial synthesizer 505 comprises a forward filter bank 701 or analysis filter bank. The forward filter bank 701 is configured to receive the decoded transport audio signals 514 and convert the signals to the time- frequency domain and generate time-frequency transport audio signals 702 to be passed to the transport signal and metadata processor 705. The time-frequency transport audio signals 702 ^(^, ^, ^) can be denoted as ^
either in vector or scalar form, where ^ is the frequency bin index, ^ is the time-frequency signal temporal index (or temporal slot index), and ^ is the channel
index. In this specific example, there are exactly two channels. Other number of channels can be employed in other examples. For example, the forward filter bank 701 comprises a short-time Fourier transform (STFT), the complex low-delay filter bank (CLDFB) or complex- modulated quadrature mirror filter (QMF) bank. In some embodiments, the filter bank is configured to have 60 frequency bins, and sufficient stop-band attenuation to avoid significant aliasing to occur when the frequency bin signals are processed. In this configuration, all frequency bins can be processed independently from each other, except that some frequency bins may share the same spatial metadata. For example, the spatial metadata may consist of spatial parameters in a limited number of frequency bands, for example, 5 bands, and each of these bands correspond to a set of one or more frequency bins provided by the Forward filter bank 701. In the following, however, it is assumed that if the spatial metadata is of lower frequency resolution than, for example, 60 bins, it is mapped (and by repetition when needed) to the appropriate, for example, 60 bins before the processing as described in the following. In some embodiments the spatial synthesiser 505 comprises a transport signal and metadata processor 705 configued to process the time-frequency transport audio signals 702 based on the gain control information 112 and rendering metadata 504 to provide processed time-frequency transport signals 706, so that they attain level properties for the objects and/or other sounds as determined in the gain control information 112. The processed time-frequency transport signals 706 can be denoted as ^
The processed time-frequency transport signals 706 can be then provided to a decorrelator/mixer 707 and a mix matrix determiner 709. As described in detail further below, the transport signal and metadata processor 705 can also be configued to also modify the rendering metadata 504 based on the gain control information 112 to obtain processed rendering metadata 716 which is passed to the mix matrix determiner 709. The spatial synthesizer 505 in some embodiments comprises a mix matrix determiner 709 configured to receive the processed time-frequency transport
signals 706 ^(^, ^) , the processed rendering metadata 716 and the decoded separated object metadata 534. The mix matrix determiner 709 is configured to determine a mixing matrix that, when applied to the processed time-frequency transport signals 706, enables a spatialized (e.g., binaural) output to be generated. In some embodiments the mix matrix determiner 709 is configured to initially determine first the processed transport signal covariance matrix
where the superscript H indicates a conjugate transpose and ^^(^) and ^^(^) are the first and last time-frequency signal temporal indices corresponding to subframe ^. In this example, there are four time indices ^ at each subframe ^. As said, the covariance matrix is determined for each bin. In other embodiments, it could be also averaged (or summed) over multiple frequency bins, in a resolution that approximates human hearing resolutions, or in the resolution of the determined spatial metadata parameters, or any suitable resolution. The mix matrix determiner 709 can then be configured to determine an overall energy value ^^(^, ^) as the sum of the diagonal values of ^^(^, ^). The mix matrix determiner 709 can furthermore be configured to determine a target covariance matrix, which consists of the levels and correlations for the spatial audio signals, which in this example is a binaural signal. To determine a target covariance matrix in a binaural form, the mix matrix determiner 709 is configured to be able to determine (for example, employing lookup from a database) the head related transfer functions (HRTFs) for any direction of arrival (DOA). The HRTF is denoted ^(^^^, ^) which is a 2x1 column vector having complex-valued gains for left and right ears for bin ^ and direction ^^^ . The corresponding HRTF covariance matrix is ^(^^^, ^) = ^(^^^, ^)^^(^^^, ^). The mix matrix determiner 709 can in some embodiments be configured to determine information of a diffuse-field covariance matrix ^^^^^(^), which may be formulated, for example, by selecting a spatially equally spaced set of directions ^^^^ where
. The target covariance matrix can therefore in some embodiments be determined by
The parameters ^′^^^^,^^^^(^, ^) and ^′^^^,^^^^(^, ^, ^) used in this equation are described in further detail later, and are a part of the processed rendering metadata. In this example implementation, there is one simultaneous MASA direction ^^^^^^^(^, ^) and ^^ object directions ^^^^ . In other embodiments, there may be more than one MASA direction, and those directions can be straightforwardly added to the equation above. Similarly, if the metadata indicates, the target covariance matrix could be generated by taking into account various other features such as coherent or incoherent spatial spreads, spatial coherences, or any other spatial features known in the art. The rendering based on spread coherences ^^^^^(^, ^) , surround coherences ^^^^^(^, ^) is described in GB2572650. The mix matrix determiner 709 can then be configured to employ any suitable method to generate a mixing matrix ^(^, ^) based on the matrices ^^(^, ^) and ^^(^, ^). For example a suitable method has been described in Vilkamo, J., Bäckström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time–frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. The formula provided in the appendix of the above publication can be used to formulate a mixing matrix ^(^, ^). The same notation for matrices has been employed as in the publication for assist the implementation of the publication approach to determine the mixing matrix In some embodiments there is also determined a prototype matrix ^ = ^ 1 0.05 0.05 1 ^ that guides the generation of the mixing matrix. The rationale of these matrices and the formula to obtain a mixing matrix ^(^, ^) based on them has been thoroughly explained in the above cited publication. In summary, the method is such that provides a mixing matrix ^(^, ^) that when applied to a signal with a covariance matrix ^^(^, ^) produces a signal with covariance matrix ^^(^, ^), in a
least-squares optimized way. In these examples the prototype matrix ^ is simply the identity matrix with slight leakage for regularization purposes. Having an identity prototype matrix means that the processing aims to produce an output that is as similar as possible to the input (i.e., with respect to the prototype signals) while obtaining the target covariance matrix
. For simplicity, the above example does not account for head orientation or changes in head orientation. However when head tracking is enabled, this could present the situation where the user or listener is facing a ‘rear’ direction. In such examples the channels of the processed time-frequency transport audio signals 706 can be mutually flipped before the above processing. Furthermore, in some embodiments the processed rendering metadata can be rotated based on any received head orientation data. Other processing can also be implemented, such as signal-dependently cross-mixing the transport audio signals before the presented rendering operations, if the user is facing side directions. Such and other procedures to account for head-tracked rendering has been described thoroughly in the cited documents and are not repeated here. The mixing matrix determiner 709 in some configurations also determines a residual processing matrix ^^(^, ^) . In some situations, it is possible that the processed transport signals do not have suitable inter-channel incoherence enabling rendering of incoherent outputs (for example, in situations where there are ambience or spread sounds). The determination of the residual processing matrix was also described in the earlier cited publication. In summary the residual processing matrix can be determined, after implementing matrix regularizations, to determine how the processing of the transport signals with ^(^, ^) falls short in obtaining the target covariance matrix
. The residual processing matrix can then be determined such that it is able to process a decorrelated version of the processed transport signals ^(^, ^) to obtain that missing portion of the target covariance matrix. In other words, the residual processing matrix achieves to produce a signal with a covariance matrix
. The mixing matrix determiner 709 furthemore can be configured to also determines HRTF gains for the separated object for each frequency bin as ^^^^^^^^, ^^.
The mixing matrix determiner 709 can then be configured to provide the mixing matrix ^(^, ^), the residual mixing matrix ^^(^, ^), and the separated object processing gains ^^^^^^^^, ^^ as processing matrices 710 to the decorrelator/mixer 707. Additionally the spatial synthesizer 505 comprises a forward filter bank 721 (or analysis filter bank) configured to receive the decoded separated object audio signal 524 and generate time-frequency object audio signal 722. The forward filter bank 721 in some embodiments can be similar to or implemented with the forward filter bank 701. The spatial synthesizer 505 furthermore comprises a separate object signal processor 703. The separate object signal processor 703 is configured to receive the time-frequency object audio signals 722 and the gain control information 112 and generate a processed time-frequency object audio signals 732 which can be passed to the decorrelator/mixer 707. The spatial synthesizer 505 furthemore comprises a decorrelator/mixer 707. The decorrelator/mixer 707 is configured to receive the processed time-frequency transport audio signals 706 ^(^, ^), the processed time-frequency object audio signal 732 ^^^^(^, ^) and the processing matrices 710 ^(^, ^) , ^^(^, ^) and ^^^^^^^^, ^^ . The decorrelator/mixer 707 is configured to first process the processed time-frequency transport audio signals 706 with decorrelators to generate decorrelated signals ^^(^, ^) . It then can be configued to apply the following mixing procedure to generate the time-frequency spatial audio signals 708 output by the decorrelator/mixer 707.
In the above processing, although not explicitly written in the equation, the processing matrices may be linearly interpolated between subframes ^ such that at each temporal index ^ of the time-frequency signal the matrices take a step from ^(^, ^ − 1) towards ^(^, ^), and similarly for the other processing coefficients ^^(^, ^) and ^^^^^^^^, ^^. The interpolation rate may be adjusted if an onset is detected (fast interpolation) or not (normal interpolation). The time-frequency spatial audio signals 708 ^(^, ^) can then be output to a inverse filter bank 711.
In some embodiments the spatial synthesizer 505 comprises an inverse filter bank 711 which is configured to apply an inverse transform corresponding to that used by the forward filter bank 701/721 to convert the time-frequency spatial audio signals 708 to spatial audio signals 114, which is the output of the spatial synthesizer 505 as shown in Figures 5 and 7 and decoder 111 of Figure 1 and 5, and also of Figure 3. It is understood that the mix matrix determiner 709 and decorrelator/mixer 707 represent only one way to synthesize a spatial output signal based on transport signals (in our example, the processed time-frequency transport signals) and spatial metadata (in our example, the processed rendering metadata), and other methods of generating the processing matrices 710 and the time-frequency spatial audio signals 708 (and the spatial audio signals 114) based on determined processed T-F object audio signals 732 and processed T-F transport audio signals 706 are known in the literature. With respect to Figure 8 is shown a flow diagram showing the operations of the spatial synthesiser 505 shown in Figure 7. Thus the decoded transport audio streams are obtained as shown in Figure 8 by step 801. A forward filter bank is configured to time-frequency domain transform the decoded transport audio streams to generate time-frequency transport audio signals as shown in Figure 8 by step 807. Furthermore the rendering metadata is obtained as shown in Figure 8 by step 802. Additionally the decoded separated object metadata is obtained as shown in Figure 8 by step 804. Furthermore the gain control information is obtained as shown in Figure 8 by step 803.Furthemore the separated object audio signal is obtained as shown in Figure 8 by step 805. A forward filter bank is configured to time-frequency domain transform the separated object audio signal to generate time-frequency separated object audio signal as shown in Figure 8 by step 806.
A processed time-frequency transport signals and processed rendering metadata based on the time-frequency transport signals, rendering metadata and gain control information is determined as shown in Figure 8 by step 809. Furthermore a processed time-frequency object signal based on a time- frequency object signal and gain control information is obtained as shown in Figure 8 by step 808. Then there is a determination of a mix matrix (or more generally processing matrices) based on processed time-frequency transport signals, processed rendering metadata and decoded separated object metadata as shown in Figure 8 by step 810. Then decorrelation can be applied and processing matrices also applied to the processed T-F transport signals and processed T-F object signal to generate T-F spatial audio signals as shown in Figure 8 by step 811. Then an inverse filter bank (or synthesis filter bank) is applied to the time- frequency domain spatial audio signals as shown in Figure 8 by step 813 to generate the spatial audio signals. The spatial audio signals can then be output as shown in Figure 8 by step 815. The transport signal and metadata processor 705 and separate object signal processor 703 are further described hereafter. As described above the transport signal and metadata processor 705 is configured to receive the time-frequency transport signals 702, gain control information 112 and the rendering metadata 504. The aim of the transport signal and metadata processor 705 is to formulate target gain coefficients for every time- frequency instance based on the parametric representation of the objects and MASA in the rendering metadata 504 and target gain coefficients in gain control information 112. In other words, the determined or formulated target gain coefficients should amplify or attenuate only the intended energetic parts of the time-frequency transport signals 702. The transport signal and metadata processor 705 thus in some embodiments attempts to generate processed time-frequency transport signals 706 which have modified time-frequency energetic content based on the presented gaining procedure below. Furthermore, the transport signal and metadata
processor 705 is configured to produce processed rendering metadata 716 which comprises modified metadata values according to the performed gaining operations. The transport signal and metadata processor 705 receives the time- frequency transport signals 702 in frequency bins ^ and temporal indices ^, and further determines total channel energetic values and total energetic value per frame ^. The total channel energetic values can be determined by
And the total energy value can be determined by
The transport signal and metadata processor 705 can also receive the rendering metadata which comprises the following parameters (as also referenced above): MASA directions ^^^^^^^(^, ^); MASA direct-to-total energy ratio ^^^^^,^^^^(^, ^); spread coherences
; surround coherences ^^^^^(^, ^); Object directions ^^^^^^(^, ^); ISM
. As stated above, in some embodiments, where the spatial metadata is of lower frequency resolution than the time-frequency audio signals (for example, fewer than the example 60 bins), the transport signal and metadata processor 705 is configured to map metadata (by repetition when needed) to the appropriate bins before the processing. In addition, the transport signal and metadata processor 705 can receive gain control information, which comprises gain coefficients ^^^^(^, ^) and , for each object ^ and MASA. In some example embodiments, the gain coefficient is a linear value, in other words, if the target amplification of the modified object is +6 dB, the related gain coefficient is 2.
The transport signal and metadata processor 705 in some embodiments is configured to have (or have access to) the information of a panning function how the object signals have been mixed into the transport audio signals. The panning function provides panning gains ^(^^^, ^) for each channel ^ for any ^^^. For example, the panning function could be the tangent panning law for loudspeakers at ±30 degrees, such that any angle beyond this interval is hard panned to the nearest loudspeaker (except for the rear ±30 arc which could also use the same panning rule). Another optional embodiment is that the panning follows a cardioid pattern shape towards left or right directions. Any further panning rule can be employed in some optional embodiments, as long as the decoder knows which panning rule was applied by the encoder. This may be known (e.g., fixed), or signalled among the spatial metadata, for example, as an index value to a table containing a set of pre-determined panning rules. Regardless of the panning rule, in the following example, the panning gains are assumed to be limited between 0 and 1, and that the square sum of the panning gains is always 1. The transport signal and metadata processor 705 in some embodiments applies the following steps for each frequency bin ^ and subframe ^ in every channel ^. First, for each object ^, the following steps are performed: - Determining total original object energetic value ^^^^(^, ^, ^) based on the object ratio ^^^^,^^^^(^, ^, ^) and total energetic value ^(^, ^): ^^^^(^, ^, ^) = ^^^^,^^^^(^, ^, ^)^(^, ^), - Determining object energetic pan values as ^(^, ^, ^) = (^(^^^^^^(^, ^), ^))^ - Formulating target channel energetic value of object ^ ^ ^^^ (^, ^, ^, ^) = ^ ^ ^^^ (^, ^) ^(^, ^, ^)^^^^(^, ^, ^) Then, for modifying the level of the MASA part, the following steps are performed:
- Channel-specific MASA energetic value is determined as the remainder of the total original channel energy subtracted by the total original channel energy value of each object
- Target channel energetic value for MASA is determined as ^ ^ ^^^^ (^, ^, ^) = ^ ^ ^^^^ (^)^^^^^(^, ^, ^) The gain value ^^(^, ^, ^) is determined based on the square root ratio of total target channel energy and total original channel energy
The obtained total target and total original energy values may be temporally smoothed, for example, using an infinite impulse response (IIR) or finite impulse response (FIR) filter, before the gains are computed. To obtain processed time-frequency transport audio signals 706 ^(^, ^), the determined gain values ^^(^, ^, 1) and ^^(^, ^, 2) are applied to the time-frequency transport audio signals 702 ^(^, ^) ^
It should be noted that, as the applied gain coefficients are determined within a temporal accuracy of a subframe ^, the same target gain coefficient is applied for every temporal index ^ within the subframe ^. The gains may also be interpolated for different slots ^ between the values for the adjacent subframes so that the values change more smoothly. In some embodiments the energetic gain value ^^^(^, ^, ^) is determined based on the ratio of total target channel energy and total original channel energy
The gain value ^^(^, ^, ^) is determined based on the energetic gain value, for example, with the square root
In some embodiments the square root is some other root, some other function, or may be omitted completely. The obtained total target and total original energy values may, in a similar manner to the method above, be temporally smoothed, for example, using an infinite impulse response (IIR) or finite impulse response (FIR) filter, before the gains are computed. For generating the processed rendering metadata 716, ISM ratios ^^^^,^^^^(^, ^, ^) and MASA ratio ^^^^^,^^^^(^, ^) are modified according to the gain control information 112. For example, firstly, the new total ratio ^′^^^^^ is determined:
where the diffuse-to-total energy ratio ^^^^^ is defined as:
The new ISM ratios and the MASA ratio are then obtained from the ratios between the original ratios and the new total ratio. Thus, the modified ISM ratios are:
And, the modified MASA ratio is:
In addition to the obtained new ISM ratios ^^ ^^^,^^^^(^, ^, ^) and new MASA direct-to-total energy ratio ^^ ^^^^,^^^^(^, ^), the processed rendering metadata 716 comprises unmodified metadata parameters of rendering metadata 504: Object directions ^^^^^^(^, ^); MASA directions ^^^^^^^(^, ^);
spread coherences ; surround coherences ^^^^^(^, ^). In some embodiments the presented processing steps described above can include the option that if a certain object and/or MASA is not modified, the corresponding ^^^^(^, ^) and/or ^^^^^(^) is 1. In some embodiments the separate object signal processor 703 receives time-frequency object audio signal 722 ^^^^(^, ^) and gain control information 112. Based on the gain control information 112, the separate object signal processor 703 can apply target amplification or attenuation to the time-frequency object audio signal 722. Gain coefficients for separated objects can be obtained from the gain control information 112 as ^^^^(^) =
As the time-frequency object audio signal 722 comprises only the time- frequency signal of the separated object, the gain coefficient determined in the gain control information 112 can be applied directly:
As previously, the gains may also be temporally interpolated for different temporal indices ^. The separate object signal processor 703 can then be configured to produce processed time-frequency object audio signals 732 ^^^^(^, ^). In some alternative embodiments, no objects are separated from the mix. Thus, processing related to the separated object audio signal 524 and separated object metadata 534 can be omitted in these embodiments. In some alternative embodiments, there may be two directions in the MASA datastream. In some alternative embodiments, the metadata can comprise other parameters additionally or instead of the parameters presented above. In some alternative embodiments, the gain editing can be implemented jointly with the position edition according to GB2104309.6. In this case, the processing matrices to be applied on the audio signals can be combined, smoothed together, and applied only once to the audio signals. In a further alternative embodiment, the operations presented herein are combined with the operations of GB2104309.6 such that when determining the
target channel energetic value, the object energetic pan values are determined based on panning gains ^(^^^, ^) when the direction of the object is not modified or using panning gains of the modified direction as described in GB2104309.6 for the objects whose direction is modified. In a further alternative embodiment, the operations presented herein are combined with the operations of GB2104309.6 such that target energetic values are defined as presented above, and direction modifications and equalization as described in GB2104309.6 are done based on the determined target energetic values. In such alternative embodiment, before determing gain values ^^(^, ^, ^) and applying them to the input signal, the processing steps described in GB2104309.6 are done based on the determined total target energies ^^^^^^^(^, ^, ^) as described above. Gaining is then applied after the required processing steps described in GB2104309.6. In some alternative embodiments, the gain editing can be implemented to a mono transport signal. In this case, the target energetic values are determined based only on the rendering metadata and determined total energetic values. In some alternative embodiments, a new total ratio for obtaining the processed rendering metadata may comprise different weighting of applied ratio components for calculations. In some examples, the gain control information may be in part or fully based on a set distance parameter, so that higher distance entails lower gain values. Furthermore in some embodiments, the gain control information can be frequency-dependent, for example, when it is based on a set distance parameter, where a larger distance may cause more attenuation at high frequencies than at low frequencies. In some embodiments, the operations are performed on bands comprising one or more bins ^. In such embodiments the signal energy values are summed over the bins in the band, and the determined gain control information is expanded to all bins in the band before applying on the signal. In some further embodiments, the input may be a combination of the audio objects and some other spatial audio format, instead of MASA. For example, an input fomat can comprise: objects + Ambisonics or objects + multi-channel audio signals (e.g., 5.1 or 7.1+4) or any suitable combination. In such embodiments, the
Ambisonics or multi-channel audio signals may, for example, be first converted to a MASA stream, and then the processing can be applied as presented herein. Alternatively, the Ambisonics or multi-channel audio signals may be converted to some other suitable parametric spatial audio format, and the subsequent processing parts can be adjusted to be compatible with that format. With respect to Figure 9 an example electronic device which may be used as the computer, encoder processor, decoder processor or any of the functional blocks described herein is shown. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, a laptop, or a teleconferencing system. In some embodiments the device 1600 comprises at least one processor or central processing unit (CPU or processor) 1607. The processor 1607 can be configured to execute various program codes such as the methods such as described herein. The device 1600 furthermore comprises a transceiver 1609 which is configured to receive the bitstream and provide it to the processor 1607. Typically, the connection is wirelessly received data from a remote device or a server, however, in some embodiments the bitstream is received via a wired connection or read from a local memory of the device. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The device may furthermore comprise a user interface (UI) 1605 which may display to the user an interface changing the loudness of the audio objects and the MASA stream, for example, using sliders. This loudness information is the gain control information 1615 provided to the processor, CPU, 1607. The device 1600 may further comprise memory (MEM) 1611 which is coupled to the processor 1607. In some embodiments the memory 1611 comprises the program code 1621 which is executed by the processor 1607. The program
code may involve instructions to perform the operations of the spatial synthesizer described above. The processor 1607 can then be configured to output the spatial audio signals, which in this example was a binaural output, to a digital to analogue converter (DAC)/Bluetooth 1601 converter. The combination of the processor, CPU, 1607 and memory, MEM, can implement the IVAS decoder 1631 functionality described above. The DAC/Bluetooth 1601 is configured to convert the spatial audio signals to an analogue form if the headphones are conventional wired (analogue) headphones. For wireless connections, the DAC/Bluetooth 1601 may be a Bluetooth transceiver. The DAC/Bluetooth 1601 block provides (either wired or wirelessly) the spatial audio to be played back with the headphones 1603 to the user. In some embodiments, the headphones 1603 may have a head tracker which may provide orientation and/or position information of the user’s head to the processor 1607 of the rendering apparatus, so that user’s head orientation is accounted for at the spatial synthesizer. In some embodiments the remote device (not shown in Figure 11) may generate the bitstream in various ways. In one situation, the remote device consists of multiple devices, for example, a device with a microphone array at a room with multiple participants, and multiple other devices with near-microphones (e.g., headset microphones) of remote participants. The microphone array may generate the MASA stream, and the remote participants may generate single-channel audio streams treated as object signals. Depending on the bit rates, these streams may be combined by a server, and conveyed to the device of Figure 11. In another example, the MASA stream is a captured spatial stream, for example, an audio recording at a sports event, and the object stream would originate from a commentator. For the present invention, the bitstream may originate from any kind of a setting. The device of Figure 11 may also capture the audio locally, and transmit it to a remote device, where the remote device may perform the rendering similarly to the device of Figure 11. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media, and optical media. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. The foregoing description has provided by way of exemplary and non- limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.