EP4627572A1 - Parametric spatial audio encoding - Google Patents
Parametric spatial audio encodingInfo
- Publication number
- EP4627572A1 EP4627572A1 EP23805488.6A EP23805488A EP4627572A1 EP 4627572 A1 EP4627572 A1 EP 4627572A1 EP 23805488 A EP23805488 A EP 23805488A EP 4627572 A1 EP4627572 A1 EP 4627572A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- vector
- audio
- ratio
- quantized
- generating
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/02—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
- G10L19/032—Quantisation or dequantisation of spectral components
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
- G10L19/20—Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L2019/0001—Codebooks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L2019/0001—Codebooks
- G10L2019/0004—Design or structure of the codebook
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L2019/0001—Codebooks
- G10L2019/0007—Codebook element generation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/03—Application of parametric coding in stereophonic audio systems
Definitions
- the directions and direct-to-total energy ratios in frequency bands are thus a parameterization that is particularly effective for spatial audio capture.
- Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency.
- An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec which is being designed to be suitable for use over a communications network such as a 3GPP 4G/5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR).
- IVAS Immersive Voice and Audio Services
- This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources.
- the codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.
- the stereo signal could be encoded, for example, with an AAC encoder and the mono signal could be encoded with an EVS encoder.
- a decoder can decode the audio signals into PCM signals and process the sound in frequency bands (using the spatial metadata) to obtain the spatial output, for example a binaural output.
- the aforementioned immersive audio codecs are particularly suitable for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, stand-alone microphone arrays).
- microphone arrays e.g., in mobile phones, VR cameras, stand-alone microphone arrays.
- an encoder can have other input types, for example, loudspeaker signals, audio object signals, Ambisonic signals.
- the means for generating the vector from the selection of all but one of the quantized ratio parameters may be for: generating a full vector from the quantized ratio parameters for the audio objects; and generating the vector from a selection of all but one of the quantized ratio parameters for the audio objects.
- the means for quantizing the ratio parameter with respect to the audio object using the first number of bits may be for scalar quantizing the ratio parameter with respect to the audio object using the first number of bits.
- the first number of bits may be three and wherein the integer value may be an integer value in base ten.
- the valid vector may be one in which one of: a sum of vector element values may be less than or equal to seven; or no element of the vector has a value which is greater than seven and the sum of vector element values may be less than or equal to seven.
- an apparatus for decoding ratio parameters for audio objects, the apparatus comprising means for: obtaining an integer value representing ratio parameters for the audio objects; converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- the means for converting the integer value to the vector representing the selection of quantized ratio parameters based on the indexing of the vector may be for: generating a single number from the integer value, by performing an iteration loop from a zeroth iteration up to and including the single number of iterations and sequentially associating index values to iteration loop iteration numbers which have valid vectors, wherein the integer value is the highest index value; and separating the single number into vector component values to generate the vector.
- the means for dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify the distribution of the specific object within the object part of the total audio environment may be for scalar dequantizing the ratio parameter with respect to the audio object using a first number of bits.
- the first number of bits may be three, the expected sum value may be seven and wherein the integer value may be an integer value in base ten.
- a method for an apparatus for encoding an audio object parameter; the method comprising: obtaining a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; quantizing the ratio parameters with respect to the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- Generating a single number value by appending elements from the vector may further comprise transforming the elements from the vector into a base representation based on the first number of bits.
- Transforming the elements from the vector into a base representation based on the first number of bits may comprise transforming the elements into one of: a base 10 representation when the first number of bits is three; base 16 representation when the first number of bits is four; or base 32 representation when the first number of bits is five.
- Generating the vector from the selection of the quantized ratio parameters may comprise generating the vector from the selection of all but one of the quantized ratio parameters.
- Generating the vector from the selection of all but one of the quantized ratio parameters may comprise: generating a full vector from the quantized ratio parameters for the audio objects; and generating the vector from a selection of all but one of the quantized ratio parameters for the audio objects.
- Quantizing the ratio parameter with respect to the audio object using the first number of bits may comprise scalar quantizing the ratio parameter with respect to the audio object using the first number of bits.
- the first number of bits may be three and wherein the integer value may be an integer value in base ten.
- the valid vector may be one in which one of: a sum of vector element values may be less than or equal to seven; or no element of the vector has a value which is greater than seven and the sum of vector element values may be less than or equal to seven.
- a method for an apparatus for decoding ratio parameters for audio objects comprising: obtaining an integer value representing ratio parameters for the audio objects; converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- Regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters may comprise generating at least one further quantized ratio parameter based on a value of summed elements of the vector subtracted from an expected sum value.
- Dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify the distribution of the specific object within the object part of the total audio environment may comprise scalar dequantizing the ratio parameter with respect to the audio object using a first number of bits.
- the first number of bits may be three, the expected sum value may be seven and wherein the integer value may be an integer value in base ten.
- an apparatus for encoding an audio object parameter
- the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; quantizing the ratio parameters with respect to the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- the apparatus caused to perform generating the integer value based on the indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects may be further caused to perform: generating a single number value by appending elements from the vector; and generating the index from the single number, by performing an iteration loop from a zeroth iteration up to and including the single number of iterations and sequentially associating index values to iteration loop iteration numbers which have valid vectors, wherein the integer value is the highest index value reached at the end of the iteration loop.
- the apparatus caused to perform generating a single number value by appending elements from the vector may be further be caused to perform transforming the elements from the vector into a base representation based on the first number of bits.
- the apparatus caused to perform transforming the elements from the vector into a base representation based on the first number of bits may be caused to perform transforming the elements into one of: a base 10 representation when the first number of bits is three; base 16 representation when the first number of bits is four; or base 32 representation when the first number of bits is five.
- the apparatus caused to perform generating the vector from the selection of the quantized ratio parameters may be further caused to perform generating the vector from the selection of all but one of the quantized ratio parameters.
- the apparatus caused to perform generating the vector from the selection of all but one of the quantized ratio parameters may be further caused to perform: generating a full vector from the quantized ratio parameters for the audio objects; and generating the vector from a selection of all but one of the quantized ratio parameters for the audio objects.
- the apparatus caused to perform quantizing the ratio parameter with respect to the audio object using the first number of bits may be caused to perform scalar quantizing the ratio parameter with respect to the audio object using the first number of bits.
- the valid vector may be one in which one of: a sum of vector element values may be less than or equal to seven; or no element of the vector has a value which is greater than seven and the sum of vector element values may be less than or equal to seven.
- an apparatus for decoding ratio parameters for audio objects, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining an integer value representing ratio parameters for the audio objects; converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- the apparatus caused to perform converting the integer value to the vector representing the selection of quantized ratio parameters based on the indexing of the vector may be caused to perform: generating a single number from the integer value, by performing an iteration loop from a zeroth iteration up to and including the single number of iterations and sequentially associating index values to iteration loop iteration numbers which have valid vectors, wherein the integer value is the highest index value; and separating the single number into vector component values to generate the vector.
- the apparatus caused to perform regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters may be further caused to perform generating at least one further quantized ratio parameter based on a value of summed elements of the vector subtracted from an expected sum value.
- the first number of bits may be three, the expected sum value may be seven and wherein the integer value may be an integer value in base ten.
- an apparatus for encoding an audio object parameter comprising: means for obtaining a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; means for quantizing the ratio parameters with respect to the audio objects using a first number of bits; means for generating a vector from a selection of the quantized ratio parameters; and means for generating an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- an apparatus for decoding ratio parameters for audio objects comprising: means for obtaining an integer value representing ratio parameters for the audio objects; means for converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; means for regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and means for dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- an apparatus for encoding an audio object parameter comprising: obtaining circuitry configured to obtain a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; quantizing circuitry configured to quantize the ratio parameters with respect to the audio objects using a first number of bits; vector generating circuitry for generating a vector from a selection of the quantized ratio parameters; and integer value generating circuitry configured to generate an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- an apparatus for decoding ratio parameters for audio objects comprising: obtaining circuitry configured to obtain an integer value representing ratio parameters for the audio objects; converting circuitry configured to convert the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating circuitry configured to regenerate at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing circuitry configured to dequantize the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus for encoding an audio object parameter, the apparatus caused to perform at least the following: obtaining a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; quantizing the ratio parameters with respect to the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus for decoding ratio parameters for audio objects, the apparatus caused to perform at least the following: obtaining an integer value representing ratio parameters for the audio objects; converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- a non-transitory computer readable medium comprising program instructions for causing an apparatus for encoding an audio object parameter, the apparatus caused to perform at least the following: obtaining a ratio parameter associated with a respective audio object within an audio environment, the audio environment comprising at least two audio objects and the ratio parameters configured to identify a distribution of the respective object within the object part of the total audio environment; quantizing the ratio parameters with respect to the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating an integer value based on an indexing from the vector, wherein the generated integer value represents the ratio parameters for the at least two audio objects.
- a fourteenth there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for decoding ratio parameters for audio objects, the apparatus caused to perform at least the following: obtaining an integer value representing ratio parameters for the audio objects; converting the integer value to a vector representing a selection of quantized ratio parameters based on an indexing of the vector; regenerating at least one further quantized ratio parameter from the vector selection of the quantized ratio parameters; and dequantizing the quantized ratio parameter to obtain ratio parameters for the audio objects, the ratio parameters configured to identify a distribution of a specific object within the object part of a total audio environment.
- An apparatus comprising means for performing the actions of the method as described above.
- a computer program comprising program instructions for causing a computer to perform the method as described above.
- a computer program product stored on a medium may cause an apparatus to perform the method as described herein.
- An electronic device may comprise apparatus as described herein.
- a chipset may comprise apparatus as described herein.
- Embodiments of the present application aim to address problems associated with the state of the art.
- Figure 1 shows schematically a system of apparatus suitable for implementing some embodiments
- Figure 2 shows schematically an example metadata extractor and metadata compressor and packer as shown in the system of apparatus as shown in Figure 1 according to some embodiments;
- FIG 4 shows schematically an example ISM vector index generator as shown in Figure 2 according to some embodiments
- Figure 5 shows a flow diagram of the operation of the example ISM vector index generator as shown in Figures 4 according to some embodiments;
- the 3GPP IVAS codec is configured to be receive a combined input format mode.
- the combined input format mode will enable simultaneous encoding of two different audio input formats.
- An example of two different audio input formats being currently considered is the combination of the MASA format with audio objects.
- the audio objects data can also be known as independent stream with metadata (ISM) and is interchangeably described herein.
- ISM ratio is used to describe the distribution of the ISM related audio content with respect to the objects. Specifically the ISM ratio identifies the distribution of a certain object within the object part of the total audio scene.
- MASA-to-total energy ratio identifies a portion of MASA stream within the total audio scene (containing the objects and the MASA).
- (1 - MASA-to-total energy ratio) identifies a portion of all the objects within the total audio scene
- the following concept as discussed in detail herein is the efficient encoding of these ISM radios.
- These ISM ratios could be indexed within a pyramidal truncation of a Zn lattice or to encoded by a suitable entropy encoder (such as a Golomb Rice coder) or a context arithmetic encoder.
- a suitable entropy encoder such as a Golomb Rice coder
- a context arithmetic encoder such as a Golomb Rice coder
- the encoding using pyramidal truncation of a Zn lattice is more efficient in terms of compression efficiency, but it needs memory to store index offsets and description of the vectors of the Zn lattice.
- the arithmetic encoding methods are generally less efficient because there is typically not sufficient data within an audio frame in order to determine the distribution of the index values.
- the embodiments as discussed herein attempt to provide an indexing method for the lattice Zn vectors which does not need to store the index offsets nor the information relative to layer values.
- the embodiments which employ such methods are efficient for lower dimensions of the lattice and can be used for the encoding the ISM ratio index vectors.
- Metadata-Assisted Spatial Audio is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.
- spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, spread coherence, distance, etc.) per time-frequency tile.
- the spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene.
- a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency subframe (and associated with each direction direct-to- total ratios, spread coherence, distance values etc) are determined.
- the quantization steps are relatively large.
- the quantization points are at 0, ⁇ 45, ⁇ 90, ⁇ 135, and 180 degrees of azimuth.
- the audio objects input format can comprise independent streams with metadata (ISM).
- the metadata may not be available (in which cases, e.g., some default values may be assumed).
- ISM ratio a parameter named ISM ratio has been defined and which identifies the distribution of a certain object within the object part of the total audio scene. The concept as discussed herein is the efficient encoding and decoding of these ISM ratio parameters.
- Figure 1 depicts an example apparatus 100 and system for implementing embodiments of the application.
- Figure 1 depicts an example apparatus and system for implementing embodiments of the application.
- the system is shown with an ‘analysis’ part.
- the ‘analysis’ part is the part from receiving the multi-channel signals up to an encoding of the metadata and downmix signal.
- the input to the system ‘analysis’ part is the multi-channel audio signals 102.
- a microphone channel signal input is described, however any suitable input (or synthetic multi-channel) format may be implemented in other embodiments.
- the spatial analyser and the spatial analysis may be implemented external to the encoder.
- the spatial (MASA) metadata associated with the audio signals may be provided to an encoder as a separate bit-stream.
- the spatial (MASA) metadata may be provided as a set of spatial (direction) index values.
- Figure 1 also depicts multiple audio objects 104 as a further input to the analysis part.
- these multiple audio objects (or audio object stream) 104 may represent various sound sources within a physical space.
- Each audio object may be characterized by an audio (object) signal and accompanying metadata comprising directional data (in the form of azimuth and elevation values) which indicate the position or direction of the audio object within a physical space on an audio frame basis.
- the multi-channel signals 102 are passed to an analyser and encoder 101 , and specifically a transport signal generator 105 and to a metadata generator 103.
- the metadata generator 103 is also configured to receive the multi-channel signals and analyse the signals to produce metadata 104 associated with the multi-channel signals and thus associated with the transport signals 106.
- the analysis processor 103 may be configured to generate the metadata which may comprise, for each time-frequency analysis interval, a direction parameter and an energy ratio parameter and a coherence parameter (and in some embodiments a diffuseness parameter).
- the direction, energy ratio and coherence parameters may in some embodiments be considered to be MASA spatial audio parameters (or MASA metadata).
- the spatial audio parameters comprise parameters which aim to characterize the sound-field created/captured by the multi-channel signals (or two or more audio signals in general).
- the parameters generated may differ from frequency band to frequency band.
- band X all of the parameters are generated and transmitted, whereas in band Y only one of the parameters is generated and transmitted, and furthermore in band Z no parameters are generated or transmitted.
- band Z no parameters are generated or transmitted.
- the transport signals 106 and the metadata 104 may be passed to a combined encoder core 109.
- the audio objects 104 may be passed to the audio object analyser 107 for processing.
- the audio object analyser 107 analyses the object audio input stream 104 in order to produce suitable audio object transport signals and audio object metadata.
- the audio object analyser may be configured to produce the audio object transport signals by downmixing the audio signals of the audio objects into a stereo channel using amplitude panning based on the associated audio object directions.
- the audio object analyser may also be configured to produce the audio object metadata associated with the audio object input stream 104.
- the audio object metadata may comprise direction values which are applicable for all sub-bands. So, if there are 4 objects, there are 4 directions.
- This metric may be used to drive the encoding of the audio object metadata 108 and the metadata 104. Furthermore, the metric as determined by the separation metadata determiner and encoder may also be used as an influencing factor in the process of encoding the transport audio signals 106 and audio object transport audio signal 128 performed by the combined encoder core 109.
- the output metric from the stream separation metadata determiner and encoder can furthermore be represented as encoded stream separation metadata and be combined into the encoded metadata stream from the combined encoder core 109.
- an associated decoder and renderer 119 which is configured to obtain the bitstream 118 comprising encoded metadata 116, encoded transport audio signals 138 and encoded audio object metadata 112 and from these generate suitable spatial audio output signals.
- the decoding and processing of such audio signals are known in principle and are not discussed in detail hereafter other than the decoding of the encoded ISM ratio metadata.
- the energies of the objects are computed in frequency bands where b k low is the lowest and b k high the highest bin of the frequency band k.
- the ISM ratios (k, n, i) can be computed as where I is the number of objects.
- the audio object metadata encoder 111 comprises an ISM ratio quantizer 203.
- the ISM ratio quantizer 203 is configured to receive the ISM ratio values 202 and quantize them.
- the quantization of each of the ratios returns a positive integer value in binary from 000 to 111 (or 0 to 7 in decimal or base 10 form).
- the quantization can be performed using any suitable number of bits.
- the following examples show a uniform scalar quantizer based on 3 bits for each value. It can also be a non-uniform scalar quantizer. The distribution of the indexes does not influence the indexing.
- the audio object metadata encoder 111 comprises an ISM ratio vector generator 205 which is configured to receive the quantized ISM ratio values and generate a vector representation of the ISM ratios for the subband.
- the vectors to be indexed are ⁇ % G Z n
- i i K ⁇ - In other words, they are the lattice Zn vectors of the pyramidal layer of norm K.
- the Vector Quantized ISM ratio index values 206 can then be passed to the ISM vector index generator 207.
- step 303 the following operation is performed of generating ISM ratio values from the independent streams with metadata as shown in Figure 3 by step 303.
- ISM ratio values Having determined the ISM ratio values, they can be quantized to generate quantized ISM ratio values as shown in Figure 3 by step 305.
- ISM vector index generator 207 is shown in further detail.
- the ISM vector index generator 207 comprises a vector component selector 401 which is configured to receive the vector quantized ISM ratio index values 206 and select vector components and pass these to a number generator 403.
- the sum of the ISM ratios across all objects is 1.
- the quantization indexes sum up to 1
- the vector component selector is configured to select and forward the first N-1 components of a N length vector.
- a complexity reduction can be employed by (in the while loops) favouring the lowest value in the beginning.
- the number generator can then pass the generated number to a number to index generator 405.
- the ISM vector index generator 207 comprises a number to index generator 405.
- ⁇ index ratio_ism_idx[0]
- the function valid() is verifying if a given number corresponds to a valid (n- 1 )-dimensional array of integers, i.e. having the Laplacian norm less or equal to K.
- the valid vector may be one in which one of: a sum of vector element values may be less than or equal to seven where the number of iterations for the valid vector loop is seven for two objects (as only one value is checked), seventy for three objects (as one two values are checked) or seven-hundred for four objects (as three values are checked) in the 3 bit quantization example.
- there are eight (0,1 ,2,3, ... 7) valid vectors for two objects 36 valid vectors for 3 objects (such as shown in the table above) and 120 valid vectors for 4 objects.
- the index value is decremented (by 1 ) if the vector corresponding to J is valid as shown in Figure 8 by step 805.
- the last component is then generated based on the difference between the sum of the n-1 components and the Laplacian norm value K as shown in Figure 8 by step 813.
- the deindexing function can furthermore be defined by the following pseudo-code:
- ratio_idx_ism[n - 1 ] K - sum
- the decoded ISM vector values 606 can then be passed to an ISM ratio generator 607.
- the metadata decoder 603 can in some embodiments comprise an ISM ratio generator 607 which is configured to receive the decoded ISM vector values 606 and generate decoded ISM ratios 608 in a manner employing the opposite methods to those described above.
- step 707 The operation of generating the ISM ratios from the ISM vector values is shown in Figure 7 by step 707.
- the ISM ratio values can then be output as shown in Figure 7 by step 709.
- the device 1400 comprises at least one processor or central processing unit 1407.
- the processor 1407 can be configured to execute various program codes such as the methods such as described herein.
- the device 1400 comprises at least one memory 1411.
- the at least one processor 1407 is coupled to the memory 1411.
- the memory 1411 can be any suitable storage means.
- the memory 1411 comprises a program code section for storing program codes implementable upon the processor 1407.
- the memory 1411 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein.
- the implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1407 whenever needed via the memory-processor coupling.
- the device 1400 comprises a user interface 1405.
- the user interface 1405 can be coupled in some embodiments to the processor 1407.
- the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405.
- the user interface 1405 can enable a user to input commands to the device 1400, for example via a keypad.
- the user interface 1405 can enable the user to obtain information from the device 1400.
- the user interface 1405 may comprise a display configured to display information from the device 1400 to the user.
- the user interface 1405 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1400 and further displaying information to the user of the device 1400.
- the user interface 1405 may be the user interface for communicating.
- the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
- some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
- firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
- While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
- Embodiments of the inventions may be practiced in various components such as integrated circuit modules.
- the design of integrated circuits is by and large a highly automated process.
- Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
- Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules.
- the resultant design in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
- circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
- non-transitory is a limitation of the medium itself (i.e., tangible, not a signal ) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Mathematical Physics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP26159862.7A EP4723108A3 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2217884.2A GB2624869A (en) | 2022-11-29 | 2022-11-29 | Parametric spatial audio encoding |
| PCT/EP2023/080907 WO2024115052A1 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP26159862.7A Division EP4723108A3 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
Publications (3)
| Publication Number | Publication Date |
|---|---|
| EP4627572A1 true EP4627572A1 (en) | 2025-10-08 |
| EP4627572B1 EP4627572B1 (en) | 2026-03-04 |
| EP4627572C0 EP4627572C0 (en) | 2026-03-04 |
Family
ID=84889624
Family Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP26159862.7A Pending EP4723108A3 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
| EP23805488.6A Active EP4627572B1 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
Family Applications Before (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP26159862.7A Pending EP4723108A3 (en) | 2022-11-29 | 2023-11-07 | Parametric spatial audio encoding |
Country Status (9)
| Country | Link |
|---|---|
| EP (2) | EP4723108A3 (en) |
| JP (1) | JP2025540763A (en) |
| KR (1) | KR20250088634A (en) |
| CN (2) | CN120226075B (en) |
| AU (1) | AU2023405234B2 (en) |
| CO (1) | CO2025006784A2 (en) |
| GB (1) | GB2624869A (en) |
| MX (1) | MX2025006029A (en) |
| WO (1) | WO2024115052A1 (en) |
Family Cites Families (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPWO2013118476A1 (en) * | 2012-02-10 | 2015-05-11 | パナソニック インテレクチュアル プロパティ コーポレーション オブアメリカPanasonic Intellectual Property Corporation of America | Acoustic / speech encoding apparatus, acoustic / speech decoding apparatus, acoustic / speech encoding method, and acoustic / speech decoding method |
| WO2014108738A1 (en) * | 2013-01-08 | 2014-07-17 | Nokia Corporation | Audio signal multi-channel parameter encoder |
| KR101868252B1 (en) * | 2013-12-17 | 2018-06-15 | 노키아 테크놀로지스 오와이 | Audio signal encoder |
| GB2578603A (en) * | 2018-10-31 | 2020-05-20 | Nokia Technologies Oy | Determination of spatial audio parameter encoding and associated decoding |
| CN114270437B (en) * | 2019-06-14 | 2025-05-30 | 弗劳恩霍夫应用研究促进协会 | Parameter encoding and decoding |
| GB2585187A (en) * | 2019-06-25 | 2021-01-06 | Nokia Technologies Oy | Determination of spatial audio parameter encoding and associated decoding |
| GB2586214A (en) * | 2019-07-31 | 2021-02-17 | Nokia Technologies Oy | Quantization of spatial audio direction parameters |
| WO2021053266A2 (en) * | 2019-09-17 | 2021-03-25 | Nokia Technologies Oy | Spatial audio parameter encoding and associated decoding |
| GB2592896A (en) * | 2020-01-13 | 2021-09-15 | Nokia Technologies Oy | Spatial audio parameter encoding and associated decoding |
| WO2022200666A1 (en) * | 2021-03-22 | 2022-09-29 | Nokia Technologies Oy | Combining spatial audio streams |
| WO2022223133A1 (en) * | 2021-04-23 | 2022-10-27 | Nokia Technologies Oy | Spatial audio parameter encoding and associated decoding |
-
2022
- 2022-11-29 GB GB2217884.2A patent/GB2624869A/en not_active Withdrawn
-
2023
- 2023-11-07 AU AU2023405234A patent/AU2023405234B2/en active Active
- 2023-11-07 CN CN202380080216.0A patent/CN120226075B/en active Active
- 2023-11-07 CN CN202610236132.7A patent/CN122050399A/en active Pending
- 2023-11-07 KR KR1020257017666A patent/KR20250088634A/en active Pending
- 2023-11-07 JP JP2025531242A patent/JP2025540763A/en active Pending
- 2023-11-07 EP EP26159862.7A patent/EP4723108A3/en active Pending
- 2023-11-07 EP EP23805488.6A patent/EP4627572B1/en active Active
- 2023-11-07 WO PCT/EP2023/080907 patent/WO2024115052A1/en not_active Ceased
-
2025
- 2025-05-22 MX MX2025006029A patent/MX2025006029A/en unknown
- 2025-05-23 CO CONC2025/0006784A patent/CO2025006784A2/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| EP4723108A2 (en) | 2026-04-08 |
| JP2025540763A (en) | 2025-12-16 |
| EP4627572B1 (en) | 2026-03-04 |
| MX2025006029A (en) | 2025-06-02 |
| EP4627572C0 (en) | 2026-03-04 |
| WO2024115052A1 (en) | 2024-06-06 |
| GB2624869A (en) | 2024-06-05 |
| KR20250088634A (en) | 2025-06-17 |
| CN122050399A (en) | 2026-05-15 |
| CO2025006784A2 (en) | 2025-06-06 |
| CN120226075A (en) | 2025-06-27 |
| AU2023405234B2 (en) | 2025-09-25 |
| AU2023405234A1 (en) | 2025-05-29 |
| CN120226075B (en) | 2026-03-13 |
| EP4723108A3 (en) | 2026-05-06 |
| GB202217884D0 (en) | 2023-01-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112639966B (en) | Determination of spatial audio parameter encoding and associated decoding | |
| WO2021130404A1 (en) | The merging of spatial audio parameters | |
| US20240185869A1 (en) | Combining spatial audio streams | |
| CN116762127A (en) | Quantize spatial audio parameters | |
| WO2022223133A1 (en) | Spatial audio parameter encoding and associated decoding | |
| WO2024115050A1 (en) | Parametric spatial audio encoding | |
| AU2024249186A1 (en) | Low coding rate parametric spatial audio encoding | |
| AU2023405234B2 (en) | Parametric spatial audio encoding | |
| CN116508098A (en) | Quantize Spatial Audio Parameters | |
| AU2024224308A1 (en) | Combined input format spatial audio encoding | |
| WO2024175320A1 (en) | Priority values for parametric spatial audio encoding | |
| AU2023405231A1 (en) | Parametric spatial audio encoding | |
| WO2025078226A1 (en) | Parametric spatial audio decoding with pass-through mode | |
| CA3208666A1 (en) | Transforming spatial audio parameters |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| GRAP | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOSNIGR1 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: GRANT OF PATENT IS INTENDED |
|
| 17P | Request for examination filed |
Effective date: 20250528 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| INTG | Intention to grant announced |
Effective date: 20250930 |
|
| GRAS | Grant fee paid |
Free format text: ORIGINAL CODE: EPIDOSNIGR3 |
|
| GRAA | (expected) grant |
Free format text: ORIGINAL CODE: 0009210 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE PATENT HAS BEEN GRANTED |
|
| AK | Designated contracting states |
Kind code of ref document: B1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: F10 Free format text: ST27 STATUS EVENT CODE: U-0-0-F10-F00 (AS PROVIDED BY THE NATIONAL OFFICE) Effective date: 20260304 Ref country code: GB Ref legal event code: FG4D |
|
| REG | Reference to a national code |
Ref country code: DE Ref legal event code: R096 Ref document number: 602023013139 Country of ref document: DE |
|
| REG | Reference to a national code |
Ref country code: IE Ref legal event code: FG4D |
|
| REG | Reference to a national code |
Ref country code: HK Ref legal event code: DE Ref document number: 40130346 Country of ref document: HK |
|
| U01 | Request for unitary effect filed |
Effective date: 20260309 |
|
| U07 | Unitary effect registered |
Designated state(s): AT BE BG DE DK EE FI FR IT LT LU LV MT NL PT RO SE SI Effective date: 20260313 |