EP4649690A1 - A method and apparatus for complexity reduction in 6dof rendering - Google Patents

A method and apparatus for complexity reduction in 6dof rendering

Info

Publication number
EP4649690A1
EP4649690A1 EP23828166.1A EP23828166A EP4649690A1 EP 4649690 A1 EP4649690 A1 EP 4649690A1 EP 23828166 A EP23828166 A EP 23828166A EP 4649690 A1 EP4649690 A1 EP 4649690A1
Authority
EP
European Patent Office
Prior art keywords
higher order
determined
audio
order ambisonics
spatial metadata
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23828166.1A
Other languages
German (de)
French (fr)
Inventor
Lauros PAJUNEN
Jussi Artturi LEPPÄNEN
Sujeet Shyamsundar Mate
Mikko-Ville Laitinen
Archontis Politis
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Technologies Oy
Original Assignee
Nokia Technologies Oy
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Technologies Oy filed Critical Nokia Technologies Oy
Publication of EP4649690A1 publication Critical patent/EP4649690A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • H04S7/304For headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/15Aspects of sound capture and related signal processing for recording or reproduction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2420/00Techniques used stereophonic systems covered by H04S but not provided for in its groups
    • H04S2420/11Application of ambisonics in stereophonic audio systems

Definitions

  • a METHOD AND APPARATUS FOR COMPLEXITY REDUCTION IN 6DOF RENDERING Field The present application relates to apparatus and methods for reduction in spatial metadata related calculations in audio rendering by avoiding perceptually insignificant computations.
  • Background Spatial audio capture approaches attempt to capture an audio environment or audio scene such that the audio environment or audio scene can be perceptually recreated to a listener in an effective manner and furthermore may permit a listener to move and/or rotate within the recreated audio environment.
  • a high-end microphone array is needed for spatial audio capture and recording spatial sound linearly at one position at the recording space.
  • One such microphone is the spherical 32- microphone Eigenmike.
  • HOA Ambisonics
  • the spatial audio can be rendered so that sounds arriving from different directions are satisfactorily separated in a reasonable auditory bandwidth.
  • multiple microphone locations enable a multi-point HOA (MPHOA) capture system where there are multiple HOA audio signals at locations within an audio scene.
  • MPHOA multi-point HOA
  • basic microphone array for up to first order ambisonics may also be used for recording the the audio scene.
  • the audio scene may comprise two or more synthetic FOA or HOA sources. Audio rendering, where the captured audio signals are presented to a listener can be part of a virtual reality (VR) or augmented reality (AR) system.
  • VR virtual reality
  • AR augmented reality
  • the audio rendering furthermore can be performed as part of a VR or AR where the listener can freely move within the environment or audio scene and rotate their head, which is known as a 6 degrees of freedom (6DoF) configuration.
  • the audio rendering can be Multi-Point HOA (MPHOA) audio rendering where the audio scene comprises multiple HOA audio signals recordings which are rendered to a user in a 6DoF manner. That is, the user is able to listen to the recorded scene from positions that may be other than the positions of the recorded HOA sources.
  • MPHOA Multi-Point HOA
  • a method for generating a spatialized audio output comprising: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on
  • the method may further comprise performing a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • Performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may comprise processing the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position.
  • Determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may comprise analysing the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • the method may further comprise obtaining an indicator identifying the frequency limit.
  • the indicator may be within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources.
  • the indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator.
  • Determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may comprise: determining an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and selecting the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices.
  • the at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources.
  • the selection of frequencies upto the frequency limit may be all frequencies upto the limit, such that the average spatial metadata may be based on the determined spatial metadata for frequencies upto the frequency limit.
  • the at least one respective channel signals of the determined at least one active higher order ambisonics audio source may comprise more than one frequency bin of the at least one respective channel signals of the determined at least one active higher order ambisonics audio source, wherein the frequency limit may define at least one of the more than one frequency bin.
  • an apparatus for generating a spatialized audio output comprising means configured to: obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtain a listener position within the audio environment; determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determine interpolated spatial metadata for spatial metadata upto the frequency limit; determine average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; use the average interpolated spatial metadata as spatial metadata above the frequency limit; and generate the spatialized audio output based on the determined spatial
  • the means may be further configured to perform a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • the means configured to perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may be configured to process the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position.
  • the means configured to determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may be configured to analyse the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • the means may be further configured to obtain an indicator identifying the frequency limit.
  • the indicator may be within a received bitstream, the bitstream may further comprising the at least two higher order ambisonics audio sources.
  • the indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator.
  • the means configured to determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may be configured to: determine an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and select the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices.
  • the at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources.
  • an apparatus for generating a spatialized audio output comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto
  • the apparatus may further be caused to perform performing a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • the apparatus caused to perform performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may be caused to perform processing the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position.
  • the apparatus caused to perform determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may be caused to perform analysing the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source.
  • the apparatus may be caused to further perform obtaining an indicator identifying the frequency limit.
  • the indicator may be within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources.
  • the indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator.
  • the apparatus caused to perform determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may be caused to further perform: determining an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and selecting the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices.
  • the at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources.
  • the selection of frequencies upto the frequency limit may be all frequencies upto the limit, such that the average spatial metadata may be based on the determined spatial metadata for frequencies upto the frequency limit.
  • the at least one respective channel signals of the determined at least one active higher order ambisonics audio source may comprise more than one frequency bin of the at least one respective channel signals of the determined at least one active higher order ambisonics audio source, wherein the frequency limit may define at least one of the more than one frequency bin.
  • an apparatus for generating a spatialized audio output comprising: obtaining circuitry configured to obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining circuitry configured to obtain a listener position within the audio environment; determining circuitry configured to determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation circuitry configured to perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining circuitry configured to determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining circuitry configured to determine interpolated spatial metadata for spatial metadata upto the frequency limit; determining circuitry configured to determine average interpolated spatial metadata based on
  • a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for generating a spatialized audio output, the apparatus caused to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency
  • a non-transitory computer readable medium comprising program instructions for causing an apparatus, for generating a spatialized audio output, to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpol
  • an apparatus for generating a spatialized audio output, comprising: means for obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; means for obtaining a listener position within the audio environment; means for determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; mean for performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; means for determining interpolated spatial metadata for spatial metadata upto the frequency limit; means for determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; means for using the average interpolated spatial metadata as spatial metadata above the
  • a computer readable medium comprising instructions for causing an apparatus, for generating a spatialized audio output, to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above
  • An apparatus comprising means for performing the actions of the method as described above.
  • An apparatus configured to perform the actions of the method as described above.
  • a computer program comprising program instructions for causing a computer to perform the method as described above.
  • a computer program product stored on a medium may cause an apparatus to perform the method as described herein.
  • An electronic device may comprise apparatus as described herein.
  • a chipset may comprise apparatus as described herein.
  • Figure 1 shows schematically a system of apparatus showing the audio rendering or reproduction of an example audio scene and within which a user can move within the audio scene according to some embodiments
  • Figure 2 shows schematically an example audio scene comprising reproduction of an audio scene where a user moves within an area determined by higher order ambisonic audio signal sources
  • Figure 3 shows schematically a full channel processing example for ambisonic audio sources
  • Figure 4 shows schematically a sub-set channel processing example for ambisonic audio sources
  • Figure 5 shows schematically a high frequency processing example according to some embodiments
  • Figure 6 shows an example flow diagram of the operation of the example apparatus according to some embodiments
  • Figure 7 shows an example
  • Figure 8 shows schematically an example device suitable for implementing the apparatus shown.
  • Embodiments of the Application The concept as discussed herein in further detail with respect to the following embodiments is related to the rendering of audio scenes wherein the audio scene was captured based on a linear or parametric spatial audio methods with two or more microphone-arrays corresponding to different positions at the recording space (or in other words with audio signal sets which are captured at respective signal set positions in the recording space). Furthermore the concept is related to attempting to lower computational complexity of the spatial analysis required for MPHOA processing. Computational complexity of the MPHOA processing is relatively high. It is of importance to attempt to lower the computational complexity where possible (preferably without sacrificing audio quality). When the computational complexity is too high for the scene and rendering system, the listener may encounter glitches in audio playback.
  • the designated frequency limit can be modifiable during bitstream creation and is delivered as part of the bitstream delivered to the renderer to enable efficient 6DoF HOA rendering.
  • the designated frequency limit is defined as part of the renderer implementation.
  • the audio signal sets are generated by microphones (or microphone-arrays).
  • a microphone arrangement may comprise one or more microphones and generate for the audio signal set one or more audio signals.
  • the audio signal set comprises audio signals which are virtual or generated audio signals (for example a virtual speaker audio signal with an associated virtual speaker location).
  • the microphone-arrays are furthermore separate from or physically located away from any processing apparatus, however this does not preclude examples where the microphones are located on the processing apparatus or are physically connected to the processing apparatus.
  • an example apparatus which can be configured to implement MPHOA processing according to some embodiments.
  • the apparatus is part of a suitable MPEG-I Audio reference audio renderer.
  • the apparatus 101 comprises a pre-processor 103.
  • the pre-processor 103 is configured to receive the head related impulse responses (HRIRs) 100 and the Higher order ambisonic microphone positions ⁇ ⁇ 104.
  • HRIRs head related impulse responses
  • the audio scene can then be segmented into triangle sections ⁇ by performing Delauney triangulation.
  • An example is shown and has been described with respect to the example audio scene shown in Figure 2.
  • the triangulation can be used later in the processing to determine which HOA sources surround the listener and are used for generating the binaural signal at the listener position.
  • the audio scene 201 comprises microphones ⁇ ⁇ 203, ⁇ ⁇ 205, ⁇ ⁇ 207, ⁇ ⁇ 209, which are arranged such that there are two triangles defined by the locations of the microhpones.
  • the scene thus comprises a first triangle is formed defined by the ‘connection’ 204 between ⁇ ⁇ 203 and ⁇ ⁇ 205, the ‘connection’ 208 between ⁇ ⁇ 205 and ⁇ ⁇ 209 and the ‘connection’ 206 between 203 and ⁇ ⁇ 209.
  • the scene comprises a second triangle is formed defined by the ‘connection’ 210 between ⁇ ⁇ 207 and ⁇ ⁇ 205, the ‘connection’ 208 between ⁇ ⁇ 205 and ⁇ ⁇ 209 and the ‘connection’ 212 between ⁇ ⁇ 209 and ⁇ ⁇ 207.
  • the listener ⁇ ⁇ 211 is located within the second triangle.
  • the active sources are the microphones which form the vertices or corners of the second triangle ⁇ ⁇ 205, ⁇ ⁇ 207 and ⁇ ⁇ 209.
  • the pre-processor 103 can be configured to sample the HRIR filters 100 at a uniform grid of directions and convert these into frequency domain HRTFs 110. This is performed as the MPHOA processing is implemented in the frequency domain.
  • the pre-processor is configured to perform these operations during the initialization of the apparatus.
  • the pre- processor 103 is employed in some embodiments once for each audio scene.
  • the apparatus 101 comprises a position pre- processor 105 configured to receive the position of the listener ⁇ ⁇ 104, the position of the Higher order ambisonic microphone positions ⁇ ⁇ ( ⁇ ) 106 and ⁇ 108 and from these generate ⁇ ⁇ ( ⁇ ) 112 - the active triangle for frame j, ⁇ ⁇ ( ⁇ , ⁇ ) 128 – the chosen interpolation weights for subframe k of frame j, ⁇ ⁇ ( ⁇ ) 126 – the chosen HOA source for frame j.
  • the position pre-processor 105 can be configured, for every frame of audio, to determine the active triangle ⁇ ⁇ 112, that is used for processing at the spatial analysis block.
  • the active triangle ⁇ ⁇ 112 is the triangle from the available triangles sections ⁇ 108 which surrounds the listener (or in other words the triangle which the listener is located or positioned in).
  • the position pre-processor 105 in some embodiments is configured to determine or select a “chosen” HOA source for signal interpolation.
  • the “chosen” HOA source is the source that determined to be closest to the listener position.
  • the position pre-processor 105 furthermore in some embodiments is also configured to determine interpolation weights ⁇ ⁇ ( ⁇ , ⁇ ) 128, these are weights referring to the HOA sources in the active triangle. The closer to the user the HOA source is, the higher the weighting factor.
  • the ⁇ ⁇ , ⁇ value contains the coordinates of the HOA sources in the active triangle on the x-y plane.
  • the apparatus 101 comprises a spatial analyser 107 configured to receive ⁇ ⁇ ( ⁇ , ⁇ ) 102, the input time domain HOA signals in Equivalent Spatial Domain representation, from the inputs and ⁇ ⁇ ( ⁇ ) 112, the active triangle for frame j, from the position pre-processor 105.
  • the spatial analyser 107 can be based on the spatial analyser with respect to GB2007710.8 and EP21201766.9 as well as the MPEG-I Immersive Audio standard working draft (ISO/IEC 23090-4 WD), Section 6.6.18 From these the spatial analyser 107 is configured to generate metadata ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 116, the azimuth for HOA source i, frame j, subframe k and frequency bin b, ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 118 elevation for HOA source i, frame j, subframe k and frequency bin b ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 120 direct-to-total energy ratio for HOA source i, frame j, subframe k and frequency bin b ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 122, energy for HOA source i, frame j, subframe k and frequency bin b and a time-frequency domain signal ⁇ ( ⁇ , ⁇ , ⁇
  • the output HOA signals are then split into ⁇ ⁇ subframes of equal length: Time-frequency domain conversion is then applied for all active HOA sources ⁇ .
  • the conversion can be performed using a suitable function such as the afSTFT function which is found for example in https://github.com/jvilkamo/afSTFT.
  • ⁇ ( ⁇ , ⁇ , ⁇ ) is a ⁇ ⁇ ⁇ ⁇ ⁇ matrix containing the time-frequency domain signals of length ⁇ ⁇ for each HOA channel.
  • the afSTFT conversion can be run for each channel ⁇ h separately.
  • the spatial analysis block calculates spatial metadata comprising direction, diffuseness and energy information. These are then passed on to the spatial metadata interpolator 111.
  • the signal covariance matrix calculation ( ⁇ ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) , above, the full covariance matrix is calculated. However, only the diagonal values and the values on the first row are used in the further processing steps.
  • Spatial metadata is then calculated for each frequency bin of each active HOA source from the covariance matrix. This includes direction information, diffuseness information as well as energy: Where ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 116 is the azimuth, ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 118 is the elevation, ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 120, the direct-to-total energy ratio and ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 122 is the energy for HOA source ⁇ , for frame ⁇ (subframe ⁇ ) and frequency bin ⁇ . These can be obtained as follows.
  • all of the channels 303 are processed or analysed by the application of STFTs 305 to all of the channels 303 to generate a range of frequency bins 307 for all of the channels and from which metadata is generated from all of the frequency bins 307.
  • This is computationally the ‘heaviest’ approach requiring time-domain transforms to be applied to all of the channels and furthermore processing of all of the frequency bins generated.
  • a reduction in processing complexity can thus be as shown in the example 401 in Figure 4 where a subset 413 of all of the channels 303 are processed or analysed by the application of STFTs 405 to the subset of all of the channels to generate a range of frequency bins 407 for the subset of all of the channels and from which metadata is generated from the (subset) of the frequency bins 407.
  • This is computationally an easier or ‘lighter’ approach requiring fewer time-domain transforms to be applied and furthermore a reduction of processing of the generated frequency bins.
  • the approach followed is that as shown in Figure 5.
  • the example 501 in Figure 5 achieves a reduction in complexity by processing all of the channels 303 by the application of STFTs 305 to all of the channels to generate a full range of frequency bins 307 for all of the channels.
  • a frequency threshold is determined or selected and based on the frequency threshold metadata is generated from the subset of the frequency bins 517 from all of the frequency bins 307. This is also computationally an easier or ‘lighter’ approach producing a reduction of processing of the generated frequency bins.
  • the generated metadata from the analysis of the subset of frequency bins (below the frequency threshold) can be interpolated as described below and used to form interpolated metadata for frequency bins above the frequency threshold.
  • the spatial analyser 107 is configured to receive or otherwise obtain a frequency threshold value, for example ⁇ ⁇ 152.
  • the spatial analyser 107 can be configured to implement the spatial analysis operations as discussed above but rather than determining the spatial metadata values for each frequency bin [0, ⁇ ⁇ ], the energy values are calculated for each frequency bin, but the Direction-of-arrival (DOA) and Direct-to-total energy ratios (DTR) are calculated up to bin ⁇ ⁇ , which indicates the maximum frequency limit index.
  • DOA Direction-of-arrival
  • DTR Direct-to-total energy ratios
  • a bitstream information parameter lowestHighBandIndex is replaced by the bitstream parameter ⁇ ⁇ , and the parameters highestSingleBinBandsIndex and intermediateBandsERBWidth are not used or employed.
  • the parameter hoaGroupHasFreqBandConfig defines a value equal to 0 indicates that the 6DoF HOA rendering utilizes all the frequency bins specified.
  • the apparatus 101 comprises a spatial metadata interpolator 111 configured to receive ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 116 ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 118 ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 120 ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 122 from the spatial analyser 107, and ⁇ ⁇ ( ⁇ , ⁇ ) from the position pre- processor 105 and from these generate interpolated metadata ⁇ ⁇ ( ⁇ , ⁇ , ⁇ ) 134 ⁇ ( ⁇ , ⁇ , ⁇ ) 136 ⁇ ( ⁇ , ⁇ , ⁇ ) 138 ⁇ ( ⁇ , ⁇ , ⁇ ) 132.
  • the spatial metadata interpolator 111 takes the metadata 116, 118, 120, 122 related to the HOA sources of the active triangle (calculated in the spatial analyser 107) and creates interpolated metadata, that is, metadata at the listener position.
  • the aim of the spatial metadata interopolator 111 is to describe the sound field at the listener position (what it should sound like at the listener position, which frequencies are coming from which direction at which energy etc.).
  • the output of the spatial metadata interpolator 111 is interpolated metadata which is a weighted sum of the spatial metadata of the HOA sources of the active triangle.
  • the weights for the weighted interpolation can be the weights ⁇ ⁇ ( ⁇ , ⁇ ) 128 calculated in the position pre-processor 105.
  • the interpolation is applied to all frequency bins for energy, and to bins [0, ⁇ ⁇ ] for DOA and DTR.
  • average DOA and DTR values are computed from the previous interpolated bin values as: where ⁇ ⁇ ( ⁇ , ⁇ , ⁇ ) , ⁇ ( ⁇ , ⁇ , ⁇ ) and ⁇ ( ⁇ , ⁇ , ⁇ ) represent the interpolated azimuth, elevation and DTR values, respectively, for processing frame ⁇ , sub-frame ⁇ and frequency bin ⁇ .
  • ⁇ ⁇ is the low frequency limit index for the averaging.
  • the value of ⁇ ⁇ is obtained as follows.
  • the centre frequency ⁇ ( ⁇ ⁇ ) is rounded to the closest one-third- octave band center frequency.
  • contains the band centre frequencies used in the afSTFT processing, such as shown in ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022.
  • the rounded frequency indexes ⁇ h ⁇ and the one-third-octave frequencies are listed in the following table which shows the closest centre frequency bin (b) for each one-third-octave band frequency.
  • the ⁇ h ⁇ ( ⁇ ⁇ ) represents the closest rounded center frequency index.
  • the metadata is adjusted as explained in 6.6.18.4.5 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022.
  • the metadata adjustments can be applied to all metadata in frequency bins [0, ⁇ ⁇ ] and to the average metadata at bin ⁇ ⁇ + 1.
  • the spatial metadata values are copied to each sub-frame.
  • the spatial metadata DOA and DTR values are converted into vector form and rotated.
  • the conversion and rotation operations can be applied to each sub-frame, which, due to the copying operation before, contain the same data for each sub- frame.
  • the DOA and DTR spatial metadata is not copied to each sub-frame in the metadata calculation phase.
  • the conversions and rotations can be applied to only to the first sub-frame of data. After the rotated metadata vector for the first sub- frame is obtained, they can be copied to the other sub-frames.
  • the energy metadata can in some embodiments be copied to each sub-frame.
  • the signal interpolator 109 takes as input the chosen HOA source frequency domain signal and provides as output a prototype frequency domain signal.
  • the prototype signal creation involves applying an equalizer gain (based on the interpolated signal energy) to the signal, rotating it according to listener head orientation and then multiplying it with an HOA to binaural transformation matrix.
  • the apparatus 101 comprises a signal interpolator 109 configured to receive ⁇ ⁇ ( ⁇ ) 126 from the position pre-processor 105, ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 114 and ⁇ ( ⁇ , ⁇ , ⁇ , ⁇ ) 122 from the spatial analyser 107 and ⁇ ( ⁇ , ⁇ , ⁇ ) 132 from the spatial metadata interpolator 111 from these generate interpolated signal ⁇ ⁇ ( ⁇ , ⁇ , ⁇ ) 130.
  • the signal interpolator 109 is configured to select as an input audio signal an audio signal associated with a HOA source closest to the listener. For example in the situation shown in Figure 2, the chosen source is ⁇ ⁇ 205.
  • This selection can be implemented in the manner as described in section 6.6.18.3.2.4 “Determine HOA source for signal interpolation” in ISO/IEC 23090-4 WD.
  • the signal ⁇ ( ⁇ ⁇ , ⁇ , ⁇ ), where ⁇ ⁇ is the index of the chosen HOA source for signal interpolation is passed on to the signal interpolator 109, where a prototype binaural signal can be calculated from it.
  • the generation of the prototype signal involves applying an equalizer gain on the signal, rotating it according to listener or users head orientation and then multiplying it with an HOA to binaural transformation matrix.
  • the equalizer gain can be calculated as follows: where ⁇ ⁇ ( ⁇ ) is the index of the chose HOA source for frame j.
  • the interpolated signal is the calculated as follows: This, for example, can be implemented in a form similar to that described in Section 6.6.18.3.4.2 “Signal interpolation” in ISO/IEC 23090-4 WD.
  • the apparatus 101 comprises a mixer 113 configured to receive interpolated signal ⁇ ⁇ ( ⁇ , ⁇ , ⁇ ) 130 from the signal interpolator 109, and interpolated metadata ⁇ ⁇ ( ⁇ , ⁇ , ⁇ ) 134 ⁇ ( ⁇ , ⁇ , ⁇ ) 136 ⁇ ( ⁇ , ⁇ , ⁇ ) 138 ⁇ ( ⁇ , ⁇ , ⁇ ) 132 from the spatial metadata interpolator 111. From these the mixer generates output audio O( ⁇ , ⁇ , ⁇ ) 142.
  • the mixer 113 is configured to take as an input the interpolated spatial metadata 134, 136, 138, 140 as well as the prototype signal 130.
  • interpolated spatial metadata interpolated spatial metadata
  • a binaural signal that is an approximation of the output that is desired (signal of the closest HOA source to the listener which has been equalized based on the interpolated signal energy).
  • the mixer 113 creates a binaural signal 142 from the interpolated signal 130 such that it has the same characteristics as the interpolated metadata 134, 136, 138, 140.
  • an optimal mixing algorithm is used, such as for example Vilkamo, J., Biffström, T., & Kuntz, A.
  • a direct portion of the covariance matrix can be calculated as: ⁇ ) ⁇ ⁇ ( ⁇ , ⁇ ) And the diffuse portion of the covariance matrix can be calculated as: Where And the final target covariance matrix can be determined as: Mixing matrices are then obtained, for example by employing an optimal mixing algorithm such as Vilkamo, J., Bburgström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time--frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411.
  • the mixer is configured to control the generation of the calculations of optimal mixing matrices with frequency boundaries ⁇ and ⁇ .
  • the comparison of ⁇ ( ⁇ ) ⁇ ⁇ is replaced with ⁇ ⁇ ⁇ ⁇ in ⁇ _ ⁇ _ ⁇ ( ⁇ ) function.
  • the calculations of the covariance matrices ⁇ ⁇ and ⁇ ⁇ are explained in ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 20226.6.18.3.5.2 for the interior processing and in 6.6.18.4.6 for the exterior processing. In both processing cases, the full covariance matrices are computed for each frequency bin.
  • the optimal mixing matrix ⁇ is formed from the diagonal values of ⁇ ⁇ and ⁇ ⁇ .
  • the covariance matrix calculations are modified to only compute the diagonal values for frequency bins ⁇ > ⁇ ⁇ .
  • the computation of ⁇ ⁇ is shown above.
  • the diagonal values can be calculated independently as: where ⁇ ( ⁇ , ⁇ ) is the prototype signal for processing frame ⁇ and frequency bin ⁇ .
  • the computation of the direct portion of the target covariance matrix ⁇ ⁇ ⁇ is shown above.
  • the interpolated energy ⁇ ( ⁇ , ⁇ , ⁇ ) values are updated for each frequency bin ⁇ .
  • the interpolated DTR value ⁇ ( ⁇ , ⁇ , ⁇ ⁇ + 1 ) and the HRTF matrix multiplication ⁇ ( ⁇ ⁇ + 1, ⁇ ) ⁇ ( ⁇ ⁇ , ⁇ ) ⁇ ⁇ ( ⁇ ⁇ + 1, ⁇ ) are calculated only once and used for all bins ⁇ > ⁇ ⁇ .
  • ⁇ ( ⁇ , ⁇ , ⁇ ⁇ + 1) and ⁇ ( ⁇ ⁇ + 1, ⁇ ) contain the DTR and HRTF data for the frequency averaged spatial metadata.
  • For bins ⁇ > ⁇ ⁇ , only the diagonal values of ⁇ ⁇ ⁇ are calculated as: ⁇ ⁇ > ⁇ ⁇ , ⁇ ⁇ [ 1,2 ] where ⁇ ⁇ is the number of sub-frames.
  • the DTR value ⁇ ( ⁇ , ⁇ , ⁇ ⁇ + 1 ) is used for all bins ⁇ > ⁇ ⁇ , and only the diagonal values of the covariance matrix are calculated: ⁇ ⁇ > ⁇ ⁇ , ⁇ ⁇ [ 1,2 ]
  • the diffuse portion of the target covariance matrix ⁇ ⁇ ⁇ is computed as shown in 6.6.18.4.6 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022.
  • the weighted DTR ⁇ ⁇ , ⁇ is calculated from the average metadata values stored at bin ⁇ ⁇ + 1, resulting in ⁇ ⁇ , ⁇ ( ⁇ , ⁇ , ⁇ ⁇ + 1). Only the diagonal values are calculated for ⁇ ⁇ : The matrix ⁇ ⁇ already contains values only on the diagonal and it is not modified.
  • the apparatus 101 comprises an output processor 115 configured to receive output audio O( ⁇ , ⁇ , ⁇ ) 142, an output time-frequency domain signal (binaural), from the mixer 113 and generates output audio signals ⁇ ⁇ ( ⁇ ) 144, an output time domain audio signal (binaural).
  • the output buffer 115 can be configured to perform an inverse time- frequency domain transform (such as STFT) to the output frequency domain signal ⁇ ( ⁇ , ⁇ , ⁇ ) to provide the final time domain output signal.
  • an inverse time- frequency domain transform such as STFT
  • STFT inverse time- frequency domain transform
  • time-frequency domain transforms STFT
  • STFT time-frequency domain transforms
  • Determine interpolated metadata for metadata above high frequency limit
  • signal covariance matrix (only diagonal values) as shown by Figure 6 by 615.
  • FIG. 7 An example system employing some emodiments is described in Figure 7, above.
  • the figure illustrates an end to end system overview for an audio scene comprising multiple HOA sources, which is rendered according to the above examples.
  • the renderer receives the scene description and audio bitstreams and performs rendering accordingly.
  • the MPHOA processing described in Figure 1 and presented in this invention is performed in the MPEG-I Audio Renderer whenever the scene comprises multiple HOA sources.
  • the system can comprise a content creator 701 which can be implemented on any suitable computer or processing device.
  • the content creator 701 comprises an (MPEG-I) encoder 711 which is configured to receive the audio scene description 700 and the audio signals or data 702.
  • the audio scene description 700 can be provided in the MPEG-I Encoder Input Format (EIF) or in other suitable format.
  • EIF MPEG-I Encoder Input Format
  • the audio scene description contains an acoustically relevant description of the contents of the audio scene, and contains, for example, the scene geometry as a mesh or voxel, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not.
  • the MPEG-I encoder 711 is configured to output encoded data 712.
  • the content creator 701 furthermore in some embodiments comprises a bitstream encoder 713 which is configured to receive the output 712 of the MPEG- I encoder 711 and the encoded audio signals from the MPEG-H encoder 711 and generate the bitstream 714.
  • the bitstream 714 in some embodiments can be streamed to end-user devices or made available for download or stored. Additionally the system comprises a server configured to obtain the bitstream 714, and store it and supply it to the player 705. In some emobodiments this is implemented by a streaming server 721 which is configured to supply the audio data 722 and MPEG-I audio 6DoF metadata bitstream 724. The relevant bitstream 724 and audio data 722 is retrieved by the player 705. In some embodiments other implementation options are feasible such as broadcast, multicast.
  • the player 705 in some embodiments comprises a playback device 731 configured to obtain or receive the audio data 722 and MPEG-I audio 6DoF metadata bitstream 724, and furthermore can be configured to receive or otherwise obtain the 6 DoF tracking information (listener orientation or position information) 734 from a suitable listener user interface, for example from the head mounted device (HMD) 741. These can for example be generated by sensors within the HMD 741 or from sensors in the environment sensing the orientation or position of the listener.
  • HMD head mounted device
  • the playback device 731 comprises a bitstream parser 733 configured to obtain the encoded metadata bitstream 724 and decode these in an opposite or inverse operation to the bitstream encoder 713 and mpeg I encoder 711 to generate audio scene description information 732 which can be passed to a MPEG-I audio renderer 735.
  • the playback device 731 comprises the MPEG-I audio renderer 735 configured to implement the rendering operations as described above and generate audio output signals which can be output to the head mounted device 741.
  • the playback device 731 can be implemented in different form factors depending on the application.
  • the playback device is equipped with its own listener position tracking apparatus or receives the listener position information from an external apparatus.
  • the playback device can in some embodiments be also equipped with headphone connector to deliver output of the rendered binaural audio to the headphones.
  • the device may be any suitable electronics device or apparatus.
  • the device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
  • the device 1600 comprises at least one processor or central processing unit 1607.
  • the processor 1607 can be configured to execute various program codes such as the methods such as described herein.
  • the device 1600 comprises a memory 1611.
  • the at least one processor 1607 is coupled to the memory 1611.
  • the memory 1611 can be any suitable storage means.
  • the memory 1611 comprises a program code section for storing program codes implementable upon the processor 1607. Furthermore in some embodiments the memory 1611 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1607 whenever needed via the memory-processor coupling.
  • the device 1600 comprises a user interface 1605.
  • the user interface 1605 can be coupled in some embodiments to the processor 1607.
  • the processor 1607 can control the operation of the user interface 1605 and receive inputs from the user interface 1605.
  • the user interface 1605 can enable a user to input commands to the device 1600, for example via a keypad. In some embodiments the user interface 1605 can enable the user to obtain information from the device 1600.
  • the user interface 1605 may comprise a display configured to display information from the device 1600 to the user.
  • the user interface 1605 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1600 and further displaying information to the user of the device 1600.
  • the device 1600 comprises an input/output port 1609.
  • the input/output port 1609 in some embodiments comprises a transceiver.
  • the transceiver in such embodiments can be coupled to the processor 1607 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network.
  • the transceiver or any suitable transceiver or transmitter and/or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
  • the transceiver can communicate with further apparatus by any suitable known communications protocol.
  • the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).
  • UMTS universal mobile telecommunications system
  • WLAN wireless local area network
  • IRDA infrared data communication pathway
  • the transceiver input/output port 1609 may be configured to transmit/receive the audio signals, the bitstream and in some embodiments perform the operations and methods as described above by using the processor 1607 executing suitable code.
  • the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
  • some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
  • the software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media, and optical media.
  • the memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
  • the data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
  • Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Stereophonic System (AREA)

Abstract

A method for generating a spatialized audio output comprising: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation.

Description

A METHOD AND APPARATUS FOR COMPLEXITY REDUCTION IN 6DOF RENDERING Field The present application relates to apparatus and methods for reduction in spatial metadata related calculations in audio rendering by avoiding perceptually insignificant computations. Background Spatial audio capture approaches attempt to capture an audio environment or audio scene such that the audio environment or audio scene can be perceptually recreated to a listener in an effective manner and furthermore may permit a listener to move and/or rotate within the recreated audio environment. For spatial audio capture and recording spatial sound linearly at one position at the recording space, a high-end microphone array is needed. One such microphone is the spherical 32- microphone Eigenmike. From the high-end microphone array higher-order Ambisonics (HOA) signals can be obtained and used for rendering. With the HOA audio signals, the spatial audio can be rendered so that sounds arriving from different directions are satisfactorily separated in a reasonable auditory bandwidth. In some systems multiple microphone locations enable a multi-point HOA (MPHOA) capture system where there are multiple HOA audio signals at locations within an audio scene. In some embodiments even basic microphone array for up to first order ambisonics may also be used for recording the the audio scene. In some other embodiments, the audio scene may comprise two or more synthetic FOA or HOA sources. Audio rendering, where the captured audio signals are presented to a listener can be part of a virtual reality (VR) or augmented reality (AR) system. The audio rendering furthermore can be performed as part of a VR or AR where the listener can freely move within the environment or audio scene and rotate their head, which is known as a 6 degrees of freedom (6DoF) configuration. Furthmore the audio rendering can be Multi-Point HOA (MPHOA) audio rendering where the audio scene comprises multiple HOA audio signals recordings which are rendered to a user in a 6DoF manner. That is, the user is able to listen to the recorded scene from positions that may be other than the positions of the recorded HOA sources. Summary There is provided according to a first aspect a method for generating a spatialized audio output comprising: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. The method may further comprise performing a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source. Performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may comprise processing the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position. Determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may comprise analysing the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source. The method may further comprise obtaining an indicator identifying the frequency limit. The indicator may be within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources. The indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator. Determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may comprise: determining an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and selecting the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices. The at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources. The selection of frequencies upto the frequency limit may be all frequencies upto the limit, such that the average spatial metadata may be based on the determined spatial metadata for frequencies upto the frequency limit. The at least one respective channel signals of the determined at least one active higher order ambisonics audio source may comprise more than one frequency bin of the at least one respective channel signals of the determined at least one active higher order ambisonics audio source, wherein the frequency limit may define at least one of the more than one frequency bin. According to a second aspect there is provided an apparatus for generating a spatialized audio output, the apparatus comprising means configured to: obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtain a listener position within the audio environment; determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determine interpolated spatial metadata for spatial metadata upto the frequency limit; determine average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; use the average interpolated spatial metadata as spatial metadata above the frequency limit; and generate the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. The means may be further configured to perform a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source. The means configured to perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may be configured to process the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position. The means configured to determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may be configured to analyse the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source. The means may be further configured to obtain an indicator identifying the frequency limit. The indicator may be within a received bitstream, the bitstream may further comprising the at least two higher order ambisonics audio sources. The indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator. The means configured to determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may be configured to: determine an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and select the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices. The at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources. The selection of frequencies upto the frequency limit is all frequencies upto the limit, such that the average spatial metadata may be based on the determined spatial metadata for frequencies upto the frequency limit. According to a third aspect there is provided an apparatus for generating a spatialized audio output, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. The apparatus may further be caused to perform performing a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source. The apparatus caused to perform performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position may be caused to perform processing the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position. The apparatus caused to perform determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit may be caused to perform analysing the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source. The apparatus may be caused to further perform obtaining an indicator identifying the frequency limit. The indicator may be within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources. The indicator identifying the frequency limit may further comprise a pre- determined frequency limit indicator. The apparatus caused to perform determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position may be caused to further perform: determining an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and selecting the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices. The at least one of the determined at least two higher order ambisonics audio sources may be at least one of the determined at least one active higher order ambisonics audio sources. The selection of frequencies upto the frequency limit may be all frequencies upto the limit, such that the average spatial metadata may be based on the determined spatial metadata for frequencies upto the frequency limit. The at least one respective channel signals of the determined at least one active higher order ambisonics audio source may comprise more than one frequency bin of the at least one respective channel signals of the determined at least one active higher order ambisonics audio source, wherein the frequency limit may define at least one of the more than one frequency bin. According to a fourth aspect there is provided an apparatus for generating a spatialized audio output, the apparatus comprising: obtaining circuitry configured to obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining circuitry configured to obtain a listener position within the audio environment; determining circuitry configured to determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation circuitry configured to perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining circuitry configured to determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining circuitry configured to determine interpolated spatial metadata for spatial metadata upto the frequency limit; determining circuitry configured to determine average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating circuitry configurd to generate the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for generating a spatialized audio output, the apparatus caused to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus, for generating a spatialized audio output, to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. According to a seventh aspect there is provided an apparatus, for generating a spatialized audio output, comprising: means for obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; means for obtaining a listener position within the audio environment; means for determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; mean for performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; means for determining interpolated spatial metadata for spatial metadata upto the frequency limit; means for determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; means for using the average interpolated spatial metadata as spatial metadata above the frequency limit; and means for generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. According to an eighth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for generating a spatialized audio output, to perform at least the following: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically a system of apparatus showing the audio rendering or reproduction of an example audio scene and within which a user can move within the audio scene according to some embodiments; Figure 2 shows schematically an example audio scene comprising reproduction of an audio scene where a user moves within an area determined by higher order ambisonic audio signal sources; Figure 3 shows schematically a full channel processing example for ambisonic audio sources; Figure 4 shows schematically a sub-set channel processing example for ambisonic audio sources; Figure 5 shows schematically a high frequency processing example according to some embodiments; Figure 6 shows an example flow diagram of the operation of the example apparatus according to some embodiments; Figure 7 shows an example and Figure 8 shows schematically an example device suitable for implementing the apparatus shown. Embodiments of the Application The concept as discussed herein in further detail with respect to the following embodiments is related to the rendering of audio scenes wherein the audio scene was captured based on a linear or parametric spatial audio methods with two or more microphone-arrays corresponding to different positions at the recording space (or in other words with audio signal sets which are captured at respective signal set positions in the recording space). Furthermore the concept is related to attempting to lower computational complexity of the spatial analysis required for MPHOA processing. Computational complexity of the MPHOA processing is relatively high. It is of importance to attempt to lower the computational complexity where possible (preferably without sacrificing audio quality). When the computational complexity is too high for the scene and rendering system, the listener may encounter glitches in audio playback. Furthermore, higher computational complexity reduces the target addressable market of playback devices that can be used to consume audio scenes requiring MPHOA audio rendering. The concept as discussed in further detail in the embodiments herein relates to complexity reduction in 6DoF audio rendering of audio scenes comprising two or more HOA sources where there is provided a method for obtaining spatially interpolated spatial metadata for frequencies above a threshold frequency (upper threshold for perceptual significance) to achieve reduction in computational complexity for spatial metadata calculation. This is achieved in some embodiments by calculating spatially interpolated spatial metadata for frequency bins above the designated frequency limit by using an average of the spatially interpolated spatial metadata calculated up to a designated frequency limit. This results in some embodiments to avoiding calculation of spatial metadata as well as spatially interpolated spatial metadata for all the frequency bins above the designated frequency limit. This can be implemented in some embodiments by the following operations: calculating spatial metadata and performing spatial metadata interpolation for frequencies up to a predefined or designated frequency limit; obtaining average interpolated spatial metadata from a subset of frequency bins of the calculated spatially interpolated spatial metadata; and using the average spatially interpolated spatial metdata as metadata for frequency bins corresponding to frequencies above the predefined frequency limit. In some embodiments the designated frequency limit can be modifiable during bitstream creation and is delivered as part of the bitstream delivered to the renderer to enable efficient 6DoF HOA rendering. In some further embodiments, the designated frequency limit is defined as part of the renderer implementation. As discussed above 6DoF is presently commonplace in virtual reality, such as VR games, where movement at the audio scene is straightforward to render as all spatial information is readily available (i.e., the position of each sound source as well as the audio signal of each source separately). In the following examples the audio signal sets are generated by microphones (or microphone-arrays). For example a microphone arrangement may comprise one or more microphones and generate for the audio signal set one or more audio signals. In some embodiments the audio signal set comprises audio signals which are virtual or generated audio signals (for example a virtual speaker audio signal with an associated virtual speaker location). In some embodiments the microphone-arrays are furthermore separate from or physically located away from any processing apparatus, however this does not preclude examples where the microphones are located on the processing apparatus or are physically connected to the processing apparatus. With respect to Figure 1 an example apparatus is shown which can be configured to implement MPHOA processing according to some embodiments. In some embodiments the apparatus is part of a suitable MPEG-I Audio reference audio renderer. In some embodiments the apparatus 101 comprises a pre-processor 103. The pre-processor 103 is configured to receive the head related impulse responses (HRIRs) 100 and the Higher order ambisonic microphone positions ^^104. The audio scene can then be segmented into triangle sections ^ by performing Delauney triangulation. An example is shown and has been described with respect to the example audio scene shown in Figure 2. The triangulation can be used later in the processing to determine which HOA sources surround the listener and are used for generating the binaural signal at the listener position. An example scene 201 is shown in Figure 2. The audio scene 201 comprises microphones ^^ 203, ^^ 205, ^^ 207, ^^ 209, which are arranged such that there are two triangles defined by the locations of the microhpones. The scene thus comprises a first triangle is formed defined by the ‘connection’ 204 between ^^ 203 and ^^ 205, the ‘connection’ 208 between ^^ 205 and ^^ 209 and the ‘connection’ 206 between 203 and ^^ 209. Furthermore the scene comprises a second triangle is formed defined by the ‘connection’ 210 between ^^ 207 and ^^ 205, the ‘connection’ 208 between ^^ 205 and ^^ 209 and the ‘connection’ 212 between ^^ 209 and ^^ 207. In this example the listener ^^ 211 is located within the second triangle. As such the active sources are the microphones which form the vertices or corners of the second triangle ^^ 205, ^^ 207 and ^^ 209. Furthermore the pre-processor 103 can be configured to sample the HRIR filters 100 at a uniform grid of directions and convert these into frequency domain HRTFs 110. This is performed as the MPHOA processing is implemented in the frequency domain. The pre-processor is configured to perform these operations during the initialization of the apparatus. Thus in some embodiments the pre- processor 103 is employed in some embodiments once for each audio scene. In some embodiments the apparatus 101 comprises a position pre- processor 105 configured to receive the position of the listener ^^ 104, the position of the Higher order ambisonic microphone positions ^^(^) 106 and ^ 108 and from these generate ^^(^) 112 - the active triangle for frame j, ^^(^, ^) 128 – the chosen interpolation weights for subframe k of frame j, ^^ (^) 126 – the chosen HOA source for frame j. Thus in some embodiments the position pre-processor 105 can be configured, for every frame of audio, to determine the active triangle ^^ 112, that is used for processing at the spatial analysis block. The active triangle ^^ 112 is the triangle from the available triangles sections ^ 108 which surrounds the listener (or in other words the triangle which the listener is located or positioned in). Thus the position pre-processor 105 in some embodiments is configured to determine or select a “chosen” HOA source for signal interpolation. In some embodiments the “chosen” HOA source is the source that determined to be closest to the listener position. The position pre-processor 105 furthermore in some embodiments is also configured to determine interpolation weights ^^(^, ^) 128, these are weights referring to the HOA sources in the active triangle. The closer to the user the HOA source is, the higher the weighting factor. The weighting factors can in some embodiments be obtained by calculating the barycentric coordinates for the triangle by solving, ^^^,^^^ = ^^,^^ where ^^,^^ = [^^ ^^ 1] is the listener position on the x-y plane. The ^^^,^ value contains the coordinates of the HOA sources in the active triangle on the x-y plane. The barycentric coordinates can then in some embodiments be used as the weighting factors. In some embodiments the apparatus 101 comprises a spatial analyser 107 configured to receive ^^^^ (^, ^) 102, the input time domain HOA signals in Equivalent Spatial Domain representation, from the inputs and ^^(^) 112, the active triangle for frame j, from the position pre-processor 105. The spatial analyser 107 can be based on the spatial analyser with respect to GB2007710.8 and EP21201766.9 as well as the MPEG-I Immersive Audio standard working draft (ISO/IEC 23090-4 WD), Section 6.6.18 From these the spatial analyser 107 is configured to generate metadata ^(^, ^, ^, ^) 116, the azimuth for HOA source i, frame j, subframe k and frequency bin b, ^(^, ^, ^, ^) 118 elevation for HOA source i, frame j, subframe k and frequency bin b ^(^, ^, ^, ^) 120 direct-to-total energy ratio for HOA source i, frame j, subframe k and frequency bin b ^(^, ^, ^, ^) 122, energy for HOA source i, frame j, subframe k and frequency bin b and a time-frequency domain signal ^(^, ^, ^, ^) 114. In some embodiments this can be achieved by first converting the input signals into higher-order Ambisonics (HOA) signals as follows: ^^^^(^, ^) = ^^^^^^^^^^^^^(^, ^), where ^^^^^^^^^ is a ^^^ × ^^^ ESD to HOA conversion matrix, ^ is the HOA Source index and ^ is the frame index. The output HOA signals are then split into ^^^ subframes of equal length: Time-frequency domain conversion is then applied for all active HOA sources ^. The conversion can be performed using a suitable function such as the afSTFT function which is found for example in https://github.com/jvilkamo/afSTFT. where ^(^, ^, ^) is a ^^^ × ^^ matrix containing the time-frequency domain signals of length ^^ for each HOA channel. The afSTFT conversion can be run for each channel ^ℎ separately. Thus, the more channels there are to process, the more computationally heavy the processing is. For each frequency bin ^ of signal the spatial analysis block calculates spatial metadata comprising direction, diffuseness and energy information. These are then passed on to the spatial metadata interpolator 111. The spatial metadata furthermore can be calculated from a signal covariance matrix ^^^^, which is obtained from the signal as follows: ^^^^(^, ^, ^, ^) = ^(^, ^, ^, ^)^^(^, ^, ^, ^), where: where ^^,^^ (^, ^, ^) is the value in matrix ^(^, ^, ^) corresponding to channel ^ℎ and frequency bin ^. In some embodiments the signal covariance matrix calculation (^^^^ (^, ^, ^, ^), above, the full covariance matrix is calculated. However, only the diagonal values and the values on the first row are used in the further processing steps. Unnecessary calculations can be avoided by only explicitly calculating the diagonal and first row values of the signal covariance matrix. Spatial metadata is then calculated for each frequency bin of each active HOA source from the covariance matrix. This includes direction information, diffuseness information as well as energy: Where ^(^, ^, ^, ^) 116 is the azimuth, ^(^, ^, ^, ^) 118 is the elevation, ^(^, ^, ^, ^) 120, the direct-to-total energy ratio and ^(^, ^, ^, ^) 122 is the energy for HOA source ^, for frame ^ (subframe ^) and frequency bin ^. These can be obtained as follows. First an intensity vector is calculated from the covariance matrix: Then the energy: Then the average (over sub-frames ^^^) of the intensity vector and energy: And the rest of the spatial metadata: The signal ^(^^ , ^, ^), where ^^ is the index of the chosen HOA source for signal interpolation is passed on to the signal interpolation block, where a prototype binaural signal is calculated from it. One approach that has applied to reduce the complexity is shown in Figures 3 and 4. In some approaches such as shown by the example 301 in Figure 3 all of the channels 303 are processed or analysed by the application of STFTs 305 to all of the channels 303 to generate a range of frequency bins 307 for all of the channels and from which metadata is generated from all of the frequency bins 307. This is computationally the ‘heaviest’ approach requiring time-domain transforms to be applied to all of the channels and furthermore processing of all of the frequency bins generated. A reduction in processing complexity can thus be as shown in the example 401 in Figure 4 where a subset 413 of all of the channels 303 are processed or analysed by the application of STFTs 405 to the subset of all of the channels to generate a range of frequency bins 407 for the subset of all of the channels and from which metadata is generated from the (subset) of the frequency bins 407. This is computationally an easier or ‘lighter’ approach requiring fewer time-domain transforms to be applied and furthermore a reduction of processing of the generated frequency bins. In the following embodiments the approach followed is that as shown in Figure 5. The example 501 in Figure 5 achieves a reduction in complexity by processing all of the channels 303 by the application of STFTs 305 to all of the channels to generate a full range of frequency bins 307 for all of the channels. However in this example a frequency threshold is determined or selected and based on the frequency threshold metadata is generated from the subset of the frequency bins 517 from all of the frequency bins 307. This is also computationally an easier or ‘lighter’ approach producing a reduction of processing of the generated frequency bins. Additionally in some embodiments the generated metadata from the analysis of the subset of frequency bins (below the frequency threshold) can be interpolated as described below and used to form interpolated metadata for frequency bins above the frequency threshold. In some embodiments the spatial analyser 107 is configured to receive or otherwise obtain a frequency threshold value, for example ^^^^ 152. The spatial analyser 107 can be configured to implement the spatial analysis operations as discussed above but rather than determining the spatial metadata values for each frequency bin [0, ^^], the energy values are calculated for each frequency bin, but the Direction-of-arrival (DOA) and Direct-to-total energy ratios (DTR) are calculated up to bin ^^^^ , which indicates the maximum frequency limit index. In some embodiments a bitstream information parameter lowestHighBandIndex is replaced by the bitstream parameter ^^^^ , and the parameters highestSingleBinBandsIndex and intermediateBandsERBWidth are not used or employed. In some embodiments the semantics for the parameters are described after the syntax. hoaGroups(){ unsigned int(8) hoaGroupsCount; for(int i=0; i<hoaGroupsCount){ unsigned int(1) hoaGroupId; unsigned int(1) hoaGroupHasFreqBandConfig; unsigned int(1) hoaGroupHasRegion; if(hoaGroupHasRegion) { unsigned int(16) hoaGroupRegionId; } if(hoaGroupHasFreqBandConfig) { unsigned int(8) bmax; } } } In these embodiments the parameter hoaGroupHasFreqBandConfig defines a value equal to 0 indicates that the 6DoF HOA rendering utilizes all the frequency bins specified. A value equal to 1 indicates that the 6DoF HOA rendering utilizes the frequency bins specified in according to the embodiments. bmax indicates the lowest bin index above all the frequency bins are merged and treated as a single bin for the subsequent rendering for the particular HOA group. In some embodiments the apparatus 101 comprises a spatial metadata interpolator 111 configured to receive ^(^, ^, ^, ^) 116 ^(^, ^, ^, ^) 118 ^(^, ^, ^, ^) 120 ^(^, ^, ^, ^) 122 from the spatial analyser 107, and ^^(^, ^) from the position pre- processor 105 and from these generate interpolated metadata ^^(^, ^, ^) 134 ^^(^, ^, ^) 136 ^̂(^, ^, ^) 138 ^̂(^, ^, ^) 132. The spatial metadata interpolator 111 takes the metadata 116, 118, 120, 122 related to the HOA sources of the active triangle (calculated in the spatial analyser 107) and creates interpolated metadata, that is, metadata at the listener position. The aim of the spatial metadata interopolator 111 is to describe the sound field at the listener position (what it should sound like at the listener position, which frequencies are coming from which direction at which energy etc.). The output of the spatial metadata interpolator 111 is interpolated metadata which is a weighted sum of the spatial metadata of the HOA sources of the active triangle. The weights for the weighted interpolation can be the weights ^^(^, ^) 128 calculated in the position pre-processor 105. In the embodiments as discussed herein, the interpolation is applied to all frequency bins for energy, and to bins [0, ^^^^] for DOA and DTR. For frequency bin ^^^^ + 1, average DOA and DTR values are computed from the previous interpolated bin values as: where ^^(^, ^, ^), ^^(^, ^, ^) and ^̂(^, ^, ^) represent the interpolated azimuth, elevation and DTR values, respectively, for processing frame ^, sub-frame ^ and frequency bin ^. ^^^^ is the low frequency limit index for the averaging. In some embodiments the value of ^^^^ is obtained as follows. The centre frequency ^^^^^^^(^^^^) is rounded to the closest one-third- octave band center frequency. ^^^^^^^ contains the band centre frequencies used in the afSTFT processing, such as shown in ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022. The rounded frequency indexes ^^^^ℎ^^^^^^^^^^^^^^^^ and the one-third-octave frequencies are listed in the following table which shows the closest centre frequency bin (b) for each one-third-octave band frequency. Hz b Hz b Hz b Hz b 16.0 0 160.0 1 1600.0 13 16000.0 89 20.0 0 200.0 2 2000.0 15 20000.0 111 25.0 0 250.0 2 2500.0 17 31.5 0 315.0 3 3150.0 21 40.0 0 400.0 4 4000.0 25 50.0 0 500.0 5 5000.0 31 63.0 0 630.0 6 6300.0 38 80.0 1 800.0 8 8000.0 47 100.0 1 1000.0 9 10000.0 57 125.0 1 1250.0 11 12500.0 71 The ^^^^ℎ^^^^^^^^^^^^^^^^(^^^^^^^^) represents the closest rounded center frequency index. ^^^^ which can be then determined with: if (^^^^^^^^ > 0) { ^^^^ = ^^^^ℎ^^^^^^^^^^^^^^^^(^^^^^^^^ − 1) if (^^^^ == ^^^^ + 1) { ^^^^ = ^^^^ } } else if (^^^^ > 0) { ^^^^ = ^^^^ } else { ^^^^ = 0 } In the exterior processing, the metadata is adjusted as explained in 6.6.18.4.5 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022. The metadata adjustments can be applied to all metadata in frequency bins [0, ^^^^] and to the average metadata at bin ^^^^ + 1. As is described in In 6.6.18.3.3.10 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022, the spatial metadata values are copied to each sub-frame. Later in the metadata interpolation phase, the spatial metadata DOA and DTR values are converted into vector form and rotated. The conversion and rotation operations can be applied to each sub-frame, which, due to the copying operation before, contain the same data for each sub- frame. In some embodiments to avoid unnecessary conversion and rotation operations, the DOA and DTR spatial metadata is not copied to each sub-frame in the metadata calculation phase. The conversions and rotations can be applied to only to the first sub-frame of data. After the rotated metadata vector for the first sub- frame is obtained, they can be copied to the other sub-frames. The energy metadata, can in some embodiments be copied to each sub-frame. The signal interpolator 109 takes as input the chosen HOA source frequency domain signal and provides as output a prototype frequency domain signal. In summary, the prototype signal creation involves applying an equalizer gain (based on the interpolated signal energy) to the signal, rotating it according to listener head orientation and then multiplying it with an HOA to binaural transformation matrix. In some embodiments the apparatus 101 comprises a signal interpolator 109 configured to receive ^^ (^) 126 from the position pre-processor 105, ^(^, ^, ^, ^) 114 and ^(^, ^, ^, ^) 122 from the spatial analyser 107 and ^̂(^, ^, ^) 132 from the spatial metadata interpolator 111 from these generate interpolated signal ^^(^, ^, ^) 130. The signal interpolator 109 is configured to select as an input audio signal an audio signal associated with a HOA source closest to the listener. For example in the situation shown in Figure 2, the chosen source is ^^ 205. This selection can be implemented in the manner as described in section 6.6.18.3.2.4 “Determine HOA source for signal interpolation” in ISO/IEC 23090-4 WD. The signal ^(^^ , ^, ^), where ^^ is the index of the chosen HOA source for signal interpolation is passed on to the signal interpolator 109, where a prototype binaural signal can be calculated from it. In some embodiments, the generation of the prototype signal involves applying an equalizer gain on the signal, rotating it according to listener or users head orientation and then multiplying it with an HOA to binaural transformation matrix. The equalizer gain can be calculated as follows: where ^^ (^) is the index of the chose HOA source for frame j. The interpolated signal is the calculated as follows: This, for example, can be implemented in a form similar to that described in Section 6.6.18.3.4.2 “Signal interpolation” in ISO/IEC 23090-4 WD. In some embodiments the apparatus 101 comprises a mixer 113 configured to receive interpolated signal ^^(^, ^, ^) 130 from the signal interpolator 109, and interpolated metadata ^^(^, ^, ^) 134 ^^(^, ^, ^) 136 ^̂(^, ^, ^) 138 ^̂(^, ^, ^) 132 from the spatial metadata interpolator 111. From these the mixer generates output audio O(^, ^, ^) 142. The mixer 113 is configured to take as an input the interpolated spatial metadata 134, 136, 138, 140 as well as the prototype signal 130. Thus, there is a description of the sound field at the listener position (interpolated spatial metadata) and a binaural signal that is an approximation of the output that is desired (signal of the closest HOA source to the listener which has been equalized based on the interpolated signal energy). To get the final output the mixer 113 creates a binaural signal 142 from the interpolated signal 130 such that it has the same characteristics as the interpolated metadata 134, 136, 138, 140. For this, an optimal mixing algorithm is used, such as for example Vilkamo, J., Bäckström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time--frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. This can for example be implemented as generating a prototype binaural signal ^(^, ^) from the interpolated signal: where ^^^(^) is a spherical harmonics rotation matrix calculated according to the listeners head and source orientation and ^^^^^^^^(^) is the Ambisonics to binaural matrix for frequency bin b. From the prototype signal a covariance matrix is calculated: A target coavariance matrix can be calculated from the interpolated metadata and the HRTFs calculated in the pre-processing step. A direct portion of the covariance matrix can be calculated as: ^)^^(^, ^) And the diffuse portion of the covariance matrix can be calculated as: Where And the final target covariance matrix can be determined as: Mixing matrices are then obtained, for example by employing an optimal mixing algorithm such as Vilkamo, J., Bäckström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time--frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. The output of the optimal mixing algorithm can be mixing matrices (^^(^, ^, ^) and ^^ ^ (^, ^, ^) ) which, when applied to the binaural prototype signal is configured to produce an output binaural signal ^(^, ^, ^) 142 with a covariance matrix equal to ^^ : ^(^, ^, ^) = ^^(^, ^, ^) ∗ ^(^ − 1, ^, ^) + ^^ ^ (^, ^, ^) ∗ ^(^, ^) where ^(^, ^) is a decorrelated time-frequency domain signal obtained from a buffer of previous binaural signals B. In some embodiments the mixer is configured to control the generation of the calculations of optimal mixing matrices with frequency boundaries ^^^^^^^^^^^^^^^^^^ and ^^^^^^^^^^^^^^^^^^. With the new upper boundary index ^^^^ present, the comparison of ^^^^^^^(^) < ^^^^^^^^^^^^^^^^^^ is replaced with ^ ≤ ^^^^ in ^^^^^^^^^_^^^^^^_^^^^(^) function. Thus for example the calculations of the covariance matrices ^^ and ^^ are explained in ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 20226.6.18.3.5.2 for the interior processing and in 6.6.18.4.6 for the exterior processing. In both processing cases, the full covariance matrices are computed for each frequency bin. However, when the frequency bin ^ > ^^^^ , the optimal mixing matrix ^ is formed from the diagonal values of ^^ and ^^. In order to save computations, the covariance matrix calculations are modified to only compute the diagonal values for frequency bins ^ > ^^^^ . The computation of ^^ is shown above. Instead of the full prototype matrix multiplication for frequency bins ^ > ^^^^, the diagonal values can be calculated independently as: where ^(^, ^) is the prototype signal for processing frame ^ and frequency bin ^. The computation of the direct portion of the target covariance matrix ^^ ^^^^^^ is shown above. The interpolated energy ^̂(^, ^, ^) values are updated for each frequency bin ^. The interpolated DTR value ^̂(^, ^, ^^^^ + 1) and the HRTF matrix multiplication ^^(^^^^ + 1, ^) = ^(^^^^ , ^)^^(^^^^ + 1, ^) are calculated only once and used for all bins ^ > ^^^^ . ^̂(^, ^, ^^^^ + 1) and ^(^^^^ + 1, ^) contain the DTR and HRTF data for the frequency averaged spatial metadata. For bins ^ > ^^^^ , only the diagonal values of ^^ ^^^^^^ are calculated as: ^^ ^ > ^^^^ , ^ ∈ [1,2] where ^^^ is the number of sub-frames. When the listener is inside the triangulated capturing area as discussed in 6.6.18.3.1.1 and Figure 56 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022, the diffuse portion of the target covariance matrix ^ ^^^^^^^ ^ is computed as shown in Equation (223) of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022. For bins ^ > ^^^^ , the modifications to the computations are similar to the direct portion. The DTR value ^̂(^, ^, ^^^^ + 1) is used for all bins ^ > ^^^^ , and only the diagonal values of the covariance matrix are calculated: ^^ ^ > ^^^^ , ^ ∈ [1,2] When the listener is outside the triangulated capturing area (as discussed in 6.6.18.4 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022, the diffuse portion of the target covariance matrix ^ ^^^^^^^ ^ is computed as shown in 6.6.18.4.6 of ISO/IEC JTC1/SC29/WG6 N0168 "WD1 of ISO/IEC 23090-4 Immersive Audio", October 2022. For frequency bins ^ > ^^^^, the weighted DTR ^^^^,^^^^^^^^ is calculated from the average metadata values stored at bin ^^^^ + 1, resulting in ^^^^,^^^^^^^^(^, ^, ^^^^ + 1). Only the diagonal values are calculated for ^^^^^^^: The matrix ^^^^^^^ already contains values only on the diagonal and it is not modified. The diagonal values for the weighted covariance matrix are obtained with: ^^ ^ > ^^^^ , ^ ∈ [1,2] The diagonal values for the diffuse target covariance matrix in the exterior rendering is formed as: ^^ ^ > ^^^^ , ^ ∈ [1,2] In some embodiments the apparatus 101 comprises an output processor 115 configured to receive output audio O(^, ^, ^) 142, an output time-frequency domain signal (binaural), from the mixer 113 and generates output audio signals ^^^^(^) 144, an output time domain audio signal (binaural). The output buffer 115 can be configured to perform an inverse time- frequency domain transform (such as STFT) to the output frequency domain signal ^(^, ^, ^) to provide the final time domain output signal. With respect to Figure 6 is shown a flow diagram of the method steps for some embodiments: For example obtain listener position as shown in Figure 6 by 601. In other words receive from the listener position and orientation interface in terms of the audio scene coordinates. Then obtain bitstream parameters and active HOA sources based on the listener position as shown in Figure 6 by 603. This can be evaluated regularly, in other words for every scene state update. Thus all the HOA sources comprising the triangle the listener is in are classified as active. Furthermore then implement time-frequency domain transforms (STFT) for the HOA sources as shown by Figure 6 by 605. Then determine (at least some, for example direction and direct energy ratios) spatial metadata upto a high frequency limit as shown by Figure 6 by 607. Vectorize the metadata and apply rotations according to the orientations of the HOA sources and the head of the listener. Then copy the metadata vectors to subframes as shown by Figure 6 by 608. Determine interpolated metadata (for metadata above high frequency limit) as shown by Figure 6 by 609. (optionally) perform Mixing adjustment based on the max frequency band information as shown by Figure 6 by 613. Furthermore determine signal covariance matrix (only diagonal values) as shown by Figure 6 by 615. Then implement rendering based on determined values as shown by Figure 6 by 617. An example system employing some emodiments is described in Figure 7, above. The figure illustrates an end to end system overview for an audio scene comprising multiple HOA sources, which is rendered according to the above examples. The renderer receives the scene description and audio bitstreams and performs rendering accordingly. The MPHOA processing described in Figure 1 and presented in this invention is performed in the MPEG-I Audio Renderer whenever the scene comprises multiple HOA sources.The system can comprise a content creator 701 which can be implemented on any suitable computer or processing device. The content creator 701 comprises an (MPEG-I) encoder 711 which is configured to receive the audio scene description 700 and the audio signals or data 702. The audio scene description 700 can be provided in the MPEG-I Encoder Input Format (EIF) or in other suitable format. Generally, the audio scene description contains an acoustically relevant description of the contents of the audio scene, and contains, for example, the scene geometry as a mesh or voxel, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not. The MPEG-I encoder 711 is configured to output encoded data 712. The content creator 701 furthermore in some embodiments comprises a bitstream encoder 713 which is configured to receive the output 712 of the MPEG- I encoder 711 and the encoded audio signals from the MPEG-H encoder 711 and generate the bitstream 714. The bitstream 714 in some embodiments can be streamed to end-user devices or made available for download or stored. Additionally the system comprises a server configured to obtain the bitstream 714, and store it and supply it to the player 705. In some emobodiments this is implemented by a streaming server 721 which is configured to supply the audio data 722 and MPEG-I audio 6DoF metadata bitstream 724. The relevant bitstream 724 and audio data 722 is retrieved by the player 705. In some embodiments other implementation options are feasible such as broadcast, multicast. The player 705 in some embodiments comprises a playback device 731 configured to obtain or receive the audio data 722 and MPEG-I audio 6DoF metadata bitstream 724, and furthermore can be configured to receive or otherwise obtain the 6 DoF tracking information (listener orientation or position information) 734 from a suitable listener user interface, for example from the head mounted device (HMD) 741. These can for example be generated by sensors within the HMD 741 or from sensors in the environment sensing the orientation or position of the listener. In some embodiments the playback device 731 comprises a bitstream parser 733 configured to obtain the encoded metadata bitstream 724 and decode these in an opposite or inverse operation to the bitstream encoder 713 and mpeg I encoder 711 to generate audio scene description information 732 which can be passed to a MPEG-I audio renderer 735. In some embodiments the playback device 731 comprises the MPEG-I audio renderer 735 configured to implement the rendering operations as described above and generate audio output signals which can be output to the head mounted device 741. The playback device 731 can be implemented in different form factors depending on the application. In some embodiments the playback device is equipped with its own listener position tracking apparatus or receives the listener position information from an external apparatus. The playback device can in some embodiments be also equipped with headphone connector to deliver output of the rendered binaural audio to the headphones. With respect to Figure 8 an example electronic device which may be used as the computer, encoder processor, decoder processor or any of the functional blocks described herein is shown. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. In some embodiments the device 1600 comprises at least one processor or central processing unit 1607. The processor 1607 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 1600 comprises a memory 1611. In some embodiments the at least one processor 1607 is coupled to the memory 1611. The memory 1611 can be any suitable storage means. In some embodiments the memory 1611 comprises a program code section for storing program codes implementable upon the processor 1607. Furthermore in some embodiments the memory 1611 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1607 whenever needed via the memory-processor coupling. In some embodiments the device 1600 comprises a user interface 1605. The user interface 1605 can be coupled in some embodiments to the processor 1607. In some embodiments the processor 1607 can control the operation of the user interface 1605 and receive inputs from the user interface 1605. In some embodiments the user interface 1605 can enable a user to input commands to the device 1600, for example via a keypad. In some embodiments the user interface 1605 can enable the user to obtain information from the device 1600. For example the user interface 1605 may comprise a display configured to display information from the device 1600 to the user. The user interface 1605 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1600 and further displaying information to the user of the device 1600. In some embodiments the device 1600 comprises an input/output port 1609. The input/output port 1609 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1607 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and/or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The transceiver input/output port 1609 may be configured to transmit/receive the audio signals, the bitstream and in some embodiments perform the operations and methods as described above by using the processor 1607 executing suitable code. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media, and optical media. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. The foregoing description has provided by way of exemplary and non- limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

CLAIMS: 1. A method for generating a spatialized audio output comprising: obtaining at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtaining a listener position within the audio environment; determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determining interpolated spatial metadata for spatial metadata upto the frequency limit; determining average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; using the average interpolated spatial metadata as spatial metadata above the frequency limit; and generating the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation.
2. The method as claimed in claim 1, further comprising performing a Short- time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source.
3. The method as claimed in claim 2, wherein performing signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position comprises processing the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position.
4. The method as claimed in any of claims 2 or 3, wherein determining spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit comprises analysing the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source.
5. The method as claimed in any of claims 1 to 4, further comprising obtaining an indicator identifying the frequency limit.
6. The method as claimed in claim 5, wherein the indicator is within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources.
7. The method as claimed in any of claims 5 or 6, wherein the indicator identifying the frequency limit further comprises a pre-determined frequency limit indicator.
8. The method as claimed in any of claims 1 to 7, wherein determining at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position comprises: determining an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and selecting the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices.
9. The method as claimed in any of claims 1 to 8, wherein the at least one of the determined at least two higher order ambisonics audio sources is at least one of the determined at least one active higher order ambisonics audio sources.
10. The method as claimed in any of claims 1 to 9, wherein the selection of frequencies upto the frequency limit is all frequencies upto the limit, such that the average spatial metadata is based on the determined spatial metadata for frequencies upto the frequency limit.
11. The method as claimed in any of claims 1 to 10, wherein the at least one respective channel signals of the determined at least one active higher order ambisonics audio source comprises more than one frequency bin of the at least one respective channel signals of the determined at least one active higher order ambisonics audio source, wherein the frequency limit defines at least one of the more than one frequency bin.
12. An apparatus for generating a spatialized audio output, the apparatus comprising means configured to: obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtain a listener position within the audio environment; determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determine interpolated spatial metadata for spatial metadata upto the frequency limit; determine average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; use the average interpolated spatial metadata as spatial metadata above the frequency limit; and generate the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation.
13. The apparatus as claimed in claim 12, wherein the means is further configured to perform a Short-time Fourier transform on channel audio signals of the determined at least one active higher ambisonics audio source to generate time-frequency representations of the channel audio signals of the determined at least one active higher ambisonics audio source.
14. The apparatus as claimed in claim 13, wherein the means configured to perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position is configured to process the time-frequency representations of the channel audio signals of the determined the at least one higher ambisonics audio source of the determined at least two higher order ambisonics audio sources based on the listener position.
15. The apparatus as claimed in any of claims 13 or 14, wherein the means configured to determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto the frequency limit is configured to analyse the time-frequency representations upto the frequency limit of the channel audio signals of the determined at least one active higher ambisonics audio source.
16. The apparatus as claimed in any of claims 12 to 15, wherein the means is further configured to obtain an indicator identifying the frequency limit.
17. The apparatus as claimed in claim 16, wherein the indicator is within a received bitstream, the bitstream further comprising the at least two higher order ambisonics audio sources.
18. The apparatus as claimed in any of claims 16 or 17, wherein the indicator identifying the frequency limit further comprises a pre-determined frequency limit indicator.
19. The apparatus as claimed in any of claims 12 to 18, wherein the means configured to determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position is configured to: determine an area within which the listener position is located, the area defined by vertice positions of at least three higher order ambisonics audio sources; and select the at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources, the at least one active audio source being those whose positions define the the area vertices.
20. The apparatus as claimed in any of claims 13 to 19, wherein the at least one of the determined at least two higher order ambisonics audio sources is at least one of the determined at least one active higher order ambisonics audio sources.
21. An apparatus for generating a spatialized audio output comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain at least two higher order ambisonics audio sources, wherein the at least two higher order ambisonics audio sources are associated with respective audio source positions within an audio environment; obtain a listener position within the audio environment; determine at least one active higher order ambisonics audio source from the at least two higher order ambisonics audio sources based on the listener position; perform signal interpolation by processing, channel signals of at least one of the determined at least two higher order ambisonics audio sources based on the listener position; determine spatial metadata by processing at least one respective channel signals of the determined at least one active higher order ambisonics audio source upto a frequency limit; determine interpolated spatial metadata for spatial metadata upto the frequency limit; determine average interpolated spatial metadata based on the determined interpolated spatial metadata upto the frequency limit; use the average interpolated spatial metadata as spatial metadata above the frequency limit; and generate the spatialized audio output based on the determined spatial metadata upto the frequency limit, determined average interpolated spatial metadata above the frequency limit and performed signal interpolation.
EP23828166.1A 2023-01-09 2023-12-12 A method and apparatus for complexity reduction in 6dof rendering Pending EP4649690A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GB202300291 2023-01-09
PCT/EP2023/085358 WO2024149548A1 (en) 2023-01-09 2023-12-12 A method and apparatus for complexity reduction in 6dof rendering

Publications (1)

Publication Number Publication Date
EP4649690A1 true EP4649690A1 (en) 2025-11-19

Family

ID=89307991

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23828166.1A Pending EP4649690A1 (en) 2023-01-09 2023-12-12 A method and apparatus for complexity reduction in 6dof rendering

Country Status (3)

Country Link
EP (1) EP4649690A1 (en)
CN (1) CN120530655A (en)
WO (1) WO2024149548A1 (en)

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10924876B2 (en) * 2018-07-18 2021-02-16 Qualcomm Incorporated Interpolating audio streams

Also Published As

Publication number Publication date
WO2024149548A1 (en) 2024-07-18
CN120530655A (en) 2025-08-22

Similar Documents

Publication Publication Date Title
US20250260941A1 (en) Rendering Reverberation
JP7728775B2 (en) Audio rendering with spatial metadata interpolation
CN115955622B (en) 6DOF rendering of audio captured by a microphone array for locations outside the microphone array
GB2614537A (en) Conditional disabling of a reverberator
WO2024115045A1 (en) Binaural audio rendering of spatial audio
US20250157455A1 (en) Reverberation Level Compensation
EP4178231A1 (en) Spatial audio reproduction by positioning at least part of a sound field
US20250260940A1 (en) Adjustment of Reverberator Based on Source Directivity
WO2024149548A1 (en) A method and apparatus for complexity reduction in 6dof rendering
WO2024149557A1 (en) A method and apparatus for complexity reduction in 6dof audio rendering
WO2024149567A1 (en) 6dof rendering of microphone-array captured audio
US20250218446A1 (en) Spatial Rendering of Reverberation
GB2634316A (en) A method and apparatus for control in 6DoF rendering
WO2025011907A1 (en) Beamforming control for 6-degrees of freedom audio rendering
WO2025218311A1 (en) Acoustic scene playback method and apparatus
WO2025218310A1 (en) Acoustic scene playback method and apparatus
CN121312155A (en) Audio rendering method, apparatus and non-volatile computer-readable storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250811

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)