WO2025201411A1 - Audio processing method and apparatus, electronic device, computer readable storage medium and computer program product - Google Patents
Audio processing method and apparatus, electronic device, computer readable storage medium and computer program productInfo
- Publication number
- WO2025201411A1 WO2025201411A1 PCT/CN2025/085050 CN2025085050W WO2025201411A1 WO 2025201411 A1 WO2025201411 A1 WO 2025201411A1 CN 2025085050 W CN2025085050 W CN 2025085050W WO 2025201411 A1 WO2025201411 A1 WO 2025201411A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sound field
- audio
- orientation
- audio processing
- adjusting
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S3/00—Systems employing more than two channels, e.g. quadraphonic
- H04S3/008—Systems employing more than two channels, e.g. quadraphonic in which the audio signals are in digital form, i.e. employing more than two discrete digital channels
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
- H04S7/304—For headphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/01—Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/03—Application of parametric coding in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
Definitions
- This disclosure relates to the field of audio processing technology, in particular to an audio processing method, an audio processing apparatus, a computer-readable storage medium and a computer program product.
- the IVAS (Immersive Voice and Audio Services) codec is an extension of the 3GPP (Third Generation Partnership Project) EVS (Enhanced Voice Services) codec. It provides full and bit exact EVS codec functionality for mono speech/audio signal input. It further provides: 1. Encoding and decoding of stereo and immersive audio formats such as multi-channel audio, scene-based audio (Ambisonics) , metadata-assisted spatial audio (MASA) , object-based audio (ISM) , and their combination; 2. VAD/DTX/CNG for rate efficient stereo and immersive conversational voice transmissions; 3. Error concealment mechanisms to combat the effects of transmission errors and lost packets. Jitter buffer management is also provided; 4. The IVAS codec operates on 20-ms audio frames.
- 3GPP Third Generation Partnership Project
- EVS Enhanced Voice Services
- the codec for Immersive Voice and Audio Services is part of a framework comprising of an encoder, decoder, and renderer.
- an audio processing method comprising: obtaining a corresponding rotation angle for each sound field of a plurality of sound fields of an audio; adjusting an orientation of the each sound field independently, according to the corresponding rotation angle; and rendering the audio according to each sound field adjusted.
- the adjusting the orientation of the each sound field independently, according to the corresponding rotation angle comprises: adjusting the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
- the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: rotating the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
- SBA Scene-Based Audio
- the each sound field comprises a plurality of ambisonic signals
- the rotating the sound field as a whole in response to the sound field type being the Scene-Based Audio (SBA) comprises: determining a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals; and determining a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
- SBA Scene-Based Audio
- the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel (MC) .
- MC Multi-Channel
- the rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being the Multi-Channel (MC) comprises: adding a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated.
- the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: converting the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio; and adjusting the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
- the rendering the audio according to each sound field adjusted comprising: calculating a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
- the adjusting unit rotates the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
- SBA Scene-Based Audio
- the rendering unit renders each channel rotated as Independent Streams with Metadata.
- the rendering unit calculates a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
- an audio processing method comprising: determining a collection of rotation matrices; and applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
- an audio processing apparatus comprising: determining module for determining a collection of rotation matrices; and applying module for applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
- the applying module determining rotated ambisonics signals, according to a column-vector of ambisonics signals and the collection of rotation matrices.
- the sound field comprises at least one of SBA, MASA or Multi-Channel.
- a computer program product comprising: instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of the above embodiments.
- FIG. 8a, 8b show schematic diagrams of the locations of possible nonzero elements of the rotation matrix for an arbitrary yaw rotation according to some embodiments of the present disclosure
- Fig. 9 shows a block diagram of an audio processing apparatus according to some embodiments of the present disclosure
- FIG. 10 shows a block diagram of the electronic device according to other embodiments of the present disclosure
- FIG. 11 shows a block diagram of the electronic device according to further embodiments of the present disclosure.
- Application an application is a service enabler deployed by service providers, manufacturers or users. Individual applications will often be enablers for a wide range of services; Receiver side: in end to end media system, 3 steps should be implemented, capturing and encoding, network transmission, decoding and rendering, the 3rd step is usually called receiver side; Loudspeaker reproduction: Use one or more speakers for audio playback; Headphone reproduction: Use headphone for audio playback; Room acoustics: Room acoustics is the simulation of how sound behaves in an enclosed space, influencing sound quality through early reflections and late reflections; Head tracking: Dynamic tracking of the user's head posture when playing back audio using headphones; Sound field: Sound field refers to the distribution and characteristics of sound waves in a given space or environment. It includes factors such as sound intensity, directionality, reflections, and reverberation within the area of interest.
- the IVAS codec supports coding of parametric spatial audio format called metadata-assisted spatial audio (MASA) .
- This format is specifically optimized for the direct immersive audio capture from smartphones and other form factors that can be unsuitable for dedicated spherical microphone arrays.
- Digital audio decoded by IVAS decoder can be rendered for loudspeaker reproduction.
- the process of rendering depends on the decoded audio format.
- the decoded format can match the loudspeaker configuration or output channels can be generated by application of multi-channel conversion gains from conversion tables.
- MASA and ISM formats the spatial audio needs to be mapped to the loudspeaker positions of the loudspeaker setup.
- amplitude panning is employed, with either the vector-base amplitude panning (VBAP) scheme (using triangles) , or an improved edge-fading amplitude panning (EFAP) scheme (using polygons) .
- VBAP vector-base amplitude panning
- EFAP improved edge-fading amplitude panning
- SBA audio the AllRAD loudspeaker decoding scheme is used with EFAP.
- the renderer when rendering binaural output obtains a room effect control indication TODO: ADD REF TO CONFIG, and based on this indication, it is determined whether to apply a room effect to the input spatial audio signal or to not apply the room effect.
- a room effect control indication TODO For rendering binaural audio with a room effect, two data sets related to binaural rendering are obtained. The first is a pre-defined data set containing the HRTFs as spherical harmonics to binaural conversion matrices. The second data set contains binaural room responses as reverberation early part energy correction gains (which are used for modifying the resulting spectrum that is obtained from the rendering according to the first data set) , and late reverberation energy correction gains and reverberation times.
- the binaural output signals are generated from the transport audio signal (s) and associated metadata based on a combination of these two data sets.
- the weights b and c control input and output of the feedback-delay network.
- Interaural coherence is controlled using u (z) and v (z) filter coefficients, while ear-dependent coloration using hL (z) and hR(z) .
- the coloration filters are pre-computed based on reverb characteristics, and on HRIR used for binauralization.
- the late reverb is driven by the set of parameters comprising of: RT60 –indicating the time that it takes for the reflections to drop 60 dB in energy level; DSR –diffuse to source signal energy ratio; Pre-delay –delay at which the computation of DSR values was made. Can be interpreted as the threshold between early reflections and late reverberation phase.
- FIG. 5 shows a flow diagram of an audio processing method according to some embodiments of the present disclosure.
- step 110 obtaining a corresponding rotation angle for each sound field of a plurality of sound fields of an audio.
- the each sound field comprises a plurality of ambisonic signals, in response to the sound field type being a Scene-Based Audio. determining a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals; and determining a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
- step 130 rendering the audio according to each sound field adjusted.
- rotated sound fields can be SBA, MASA, Multi-Channel and their combination.
- Fig. 8b the locations of possible nonzero elements (indicated by the symbols) of this matrix may be illustrated in Fig. 8b
- the pitch rotation and roll rotation are similar to yaw rotation, and it's just that they have some rules of their own.
- this format is specifically optimized for direct immersive audio capture from smartphones and other form factors that can be unsuitable for dedicated spherical microphone arrays.
- ⁇ n channel direction of MC, represents rotation angle.
- the input is MASA format, that should be converted to SBA or MC, then rotate the converted sound field by above known rules.
- renderer control parameters include rotation parameters of input sound field, there need to rotate it separately according to the type of input audio.
- ⁇ n channel direction of MC, represents rotation angle.
- the sound field comprises at least one of SBA, MASA or Multi-Channel.
- a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements the audio processing method according to any one of the above embodiments.
- the rendering unit 93 calculates a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
- FIG. 10 shows a block diagram of the electronic device according to other embodiments of the present disclosure.
- the electronic device 6 comprises: a memory 61 and a processor 62 coupled to the memory 61, the processor 62 configured to, based on instructions stored in the memory 61, carry out the audio processing method according to any one of the embodiments of the present disclosure.
- FIG. 11 shows a block diagram of the electronic device according to further embodiments of the present disclosure.
- the apparatus 7 for generating the fitness regimen information of this embodiment comprises: a memory 710 and a processor 720 coupled to the memory 710, the processor 720 configured to, based on instructions stored in the memory 710, carry out the audio processing method according to any one of the embodiments of the present disclosure.
- the memory 710 may comprise, for example, system memory, a fixed non-transitory storage medium, or the like.
- the system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
- the apparatus 7 for generating the fitness regimen information may further comprise an input-output interface 730, a network interface 740, a storage interface 750, and the like. These interfaces 730, 740, 750, the memory 710 and the processor 720 may be connected through a bus 760, for example.
- the input-output interface 730 provides a connection interface for input-output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, a loudspeaker, etc.
- the network interface 740 provides a connection interface for various networked devices.
- the storage interface 750 provides a connection interface for external storage devices such as an SD card and a USB flash disk.
- the method and system of the present disclosure may be implemented in many ways.
- the method and system of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware.
- the above sequence of steps of the method is merely for the purpose of illustration, and the steps of the method of the present disclosure are not limited to the above-described specific order unless otherwise specified.
- the present disclosure may also be implemented as programs recorded in a recording medium, which comprise machine-readable instructions for implementing the method according to the present disclosure.
- the present disclosure also covers a recording medium storing programs for executing the method according to the present disclosure.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Multimedia (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Stereophonic System (AREA)
Abstract
An audio processing method, an audio processing apparatus, an electronic device, a computer-readable storage medium and a computer program product. The audio processing method comprises: Obtaining a corresponding rotation angle(110); Adjusting an orientation of each sound field independently(120); Rendering the audio(130).
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
This application claims priority to PCT Patent Application No. PCT/CN2024/084739, and filed on March 29, 2024. The entire disclosure of the prior application is hereby incorporated by reference in its entirety.
This disclosure relates to the field of audio processing technology, in particular to an audio processing method, an audio processing apparatus, a computer-readable storage medium and a computer program product.
The IVAS (Immersive Voice and Audio Services) codec is an extension of the 3GPP (Third Generation Partnership Project) EVS (Enhanced Voice Services) codec. It provides full and bit exact EVS codec functionality for mono speech/audio signal input. It further provides:
1. Encoding and decoding of stereo and immersive audio formats such as multi-channel audio, scene-based
audio (Ambisonics) , metadata-assisted spatial audio (MASA) , object-based audio (ISM) , and their combination;
2. VAD/DTX/CNG for rate efficient stereo and immersive conversational voice transmissions;
3. Error concealment mechanisms to combat the effects of transmission errors and lost packets. Jitter buffer
management is also provided;
4. The IVAS codec operates on 20-ms audio frames. In addition, rendering is possible with 5-ms granularity;
5. Support for bit rate switching upon command;
6. Stereo and immersive audio coding at the following discrete bit rates [kbps] : 13.2, 16.4, 24.4, 32, 48, 64,
80, 128, 160, 192, 256, 384, and 512, with supported bit rate ranges listed in table1.
Table 1 Ranges of supported bitrates for stereo and immersive coding of the IVAS codec
1. Encoding and decoding of stereo and immersive audio formats such as multi-channel audio, scene-based
audio (Ambisonics) , metadata-assisted spatial audio (MASA) , object-based audio (ISM) , and their combination;
2. VAD/DTX/CNG for rate efficient stereo and immersive conversational voice transmissions;
3. Error concealment mechanisms to combat the effects of transmission errors and lost packets. Jitter buffer
management is also provided;
4. The IVAS codec operates on 20-ms audio frames. In addition, rendering is possible with 5-ms granularity;
5. Support for bit rate switching upon command;
6. Stereo and immersive audio coding at the following discrete bit rates [kbps] : 13.2, 16.4, 24.4, 32, 48, 64,
80, 128, 160, 192, 256, 384, and 512, with supported bit rate ranges listed in table1.
Table 1 Ranges of supported bitrates for stereo and immersive coding of the IVAS codec
13.2 kbps-128 kbps for 1 ISM, 16.4 kbps-256 kbps for 2 ISMs, 24.4 kbps-384 kbps for 3 ISMs, 24.4 kbps-512 kbps for 4 ISMs.
The codec for Immersive Voice and Audio Services is part of a framework comprising of an encoder, decoder, and renderer.
According to some embodiments of the present disclosure, there is provided an audio processing method, comprising: obtaining a corresponding rotation angle for each sound field of a plurality of sound fields of an audio; adjusting an orientation of the each sound field independently, according to the corresponding rotation angle; and rendering the audio according to each sound field adjusted.
In some embodiments, the adjusting the orientation of the each sound field independently, according to the corresponding rotation angle comprises: adjusting the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
In some embodiments, the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: rotating the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
In some embodiments, the each sound field comprises a plurality of ambisonic signals, and the rotating the sound field as a whole in response to the sound field type being the Scene-Based Audio (SBA) comprises: determining a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals; and determining a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
In some embodiments, the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel (MC) .
In some embodiments, the rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being the Multi-Channel (MC) comprises: adding a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated.
In some embodiments, the rendering the audio according to the each sound field adjusted comprises: rendering each channel rotated as Independent Streams with Metadata.
In some embodiments, the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises: converting the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio; and adjusting the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
In some embodiments, the rendering the audio according to each sound field adjusted comprising: calculating a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, comprising: obtaining unit, configured to obtain a corresponding rotation angle for each sound field of a plurality of sound fields of an audio; adjusting unit, configured to adjust an orientation of the each sound field independently, according to the corresponding rotation angle; and rendering unit, configured to render the audio according to each sound field adjusted.
In some embodiments, the adjusting unit adjusts the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
In some embodiments, the adjusting unit rotates the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
In some embodiments, the each sound field comprises a plurality of ambisonic signals, and the adjusting unit determines a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals and determines a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
In some embodiments, the adjusting unit rotates the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel (MC) .
In some embodiments, the adjusting unit adds a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated.
In some embodiments, the rendering unit renders each channel rotated as Independent Streams with Metadata.
In some embodiments, the adjusting unit converts the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio and adjusts the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
In some embodiments, the rendering unit calculates a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
According to some embodiments of the present disclosure, there is provided an audio processing method, comprising: determining a collection of rotation matrices; and applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
In some embodiments, the applying the collection of rotation matrices to a set of ambisonics signals comprises: determining rotated ambisonics signals, according to a column-vector of ambisonics signals and the collection of rotation matrices.
In some embodiments, the sound field comprises at least one of SBA, MASA or Multi-Channel.
According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, comprising: determining module for determining a collection of rotation matrices; and applying module for applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
In some embodiments, the applying module determining rotated ambisonics signals, according to a column-vector of ambisonics signals and the collection of rotation matrices.
In some embodiments, the sound field comprises at least one of SBA, MASA or Multi-Channel.
According to still other embodiments of the present disclosure, there is provided an electronic device, comprising: a memory; a processor coupled to the memory, the processor configured to, based on instructions stored in the memory, carry out the audio processing method according to any one of the above embodiments.
According to still other embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements the audio processing method according to any one of the above embodiments.
According to still other embodiments of the present disclosure, there is provided a computer program product, comprising: instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of the above embodiments.
The accompanying drawings, which are incorporated in and constitute a portion of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
The present disclosure will be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
FIG. 1 shows schematic diagrams of overview of IVAS audio processing functions -receiver side;
FIG. 2 shows schematic diagrams of overview of TD binaural renderer;
FIG. 3 shows schematic diagrams of feedback-delay network reverberator;
FIG. 4 shows schematic diagrams of coordinate system used by the Early Reflections mode;
FIG. 5 shows a flow diagram of an audio processing method according to some embodiments of the present
disclosure;
FIG. 6 shows schematic diagrams of IVAS rendering multiple sound fields with fixed orientation;
FIG. 7 shows schematic diagrams of rendering multiple sound fields with dynamic orientation according
to some embodiments of the present disclosure;
FIG. 8a, 8b show schematic diagrams of the locations of possible nonzero elements of the rotation matrix
for an arbitrary yaw rotation according to some embodiments of the present disclosure;
Fig. 9 shows a block diagram of an audio processing apparatus according to some embodiments of the
present disclosure;
FIG. 10 shows a block diagram of the electronic device according to other embodiments of the present
disclosure;
FIG. 11 shows a block diagram of the electronic device according to further embodiments of the present
disclosure.
FIG. 1 shows schematic diagrams of overview of IVAS audio processing functions -receiver side;
FIG. 2 shows schematic diagrams of overview of TD binaural renderer;
FIG. 3 shows schematic diagrams of feedback-delay network reverberator;
FIG. 4 shows schematic diagrams of coordinate system used by the Early Reflections mode;
FIG. 5 shows a flow diagram of an audio processing method according to some embodiments of the present
disclosure;
FIG. 6 shows schematic diagrams of IVAS rendering multiple sound fields with fixed orientation;
FIG. 7 shows schematic diagrams of rendering multiple sound fields with dynamic orientation according
to some embodiments of the present disclosure;
FIG. 8a, 8b show schematic diagrams of the locations of possible nonzero elements of the rotation matrix
for an arbitrary yaw rotation according to some embodiments of the present disclosure;
Fig. 9 shows a block diagram of an audio processing apparatus according to some embodiments of the
present disclosure;
FIG. 10 shows a block diagram of the electronic device according to other embodiments of the present
disclosure;
FIG. 11 shows a block diagram of the electronic device according to further embodiments of the present
disclosure.
Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Notice that, unless otherwise specified, the relative arrangement, numerical expressions and numerical values of the components and steps set forth in these examples do not limit the scope of the disclosure.
At the same time, it should be understood that, for ease of description, the dimensions of the plurality of parts shown in the drawings are not drawn to actual proportions.
The following description of at least one exemplary embodiment is in fact merely illustrative and is in no way intended as a limitation to the disclosure, its application or use.
Techniques, methods, and apparatus known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, these techniques, methods, and apparatuses should be considered as part of the specification.
Of all the examples shown and discussed herein, any specific value should be construed as merely illustrative and not as a limitation. Thus, other examples of exemplary embodiments may have different values.
Notice that, similar reference numerals and letters are denoted by the like in the accompanying drawings, and therefore, once an item is defined in a drawing, there is no need for further discussion in the accompanying drawings.
The terms used in this disclosure are described as following:
Application: an application is a service enabler deployed by service providers, manufacturers or users.
Individual applications will often be enablers for a wide range of services;
Receiver side: in end to end media system, 3 steps should be implemented, capturing and encoding,
network transmission, decoding and rendering, the 3rd step is usually called receiver side;
Loudspeaker reproduction: Use one or more speakers for audio playback;
Headphone reproduction: Use headphone for audio playback;
Room acoustics: Room acoustics is the simulation of how sound behaves in an enclosed space, influencing
sound quality through early reflections and late reflections;
Head tracking: Dynamic tracking of the user's head posture when playing back audio using headphones;
Sound field: Sound field refers to the distribution and characteristics of sound waves in a given space or
environment. It includes factors such as sound intensity, directionality, reflections, and reverberation within the area of interest.
Application: an application is a service enabler deployed by service providers, manufacturers or users.
Individual applications will often be enablers for a wide range of services;
Receiver side: in end to end media system, 3 steps should be implemented, capturing and encoding,
network transmission, decoding and rendering, the 3rd step is usually called receiver side;
Loudspeaker reproduction: Use one or more speakers for audio playback;
Headphone reproduction: Use headphone for audio playback;
Room acoustics: Room acoustics is the simulation of how sound behaves in an enclosed space, influencing
sound quality through early reflections and late reflections;
Head tracking: Dynamic tracking of the user's head posture when playing back audio using headphones;
Sound field: Sound field refers to the distribution and characteristics of sound waves in a given space or
environment. It includes factors such as sound intensity, directionality, reflections, and reverberation within the area of interest.
The abbreviations used in this disclosure are described as following:
3GPP Third Generation Partnership Project;
IVAS Immersive Voice and Audio Services;
FOA First-Order Ambisonics;
HOA Higher-Order Ambisonics;
HRIR Head-Related Impulse Response ;
HRTF Head-Related Transfer Function;
ITD Inter-Channel Time Delay;
ISM Independent Streams with Metadata;
LFE Low-Frequency Effects;
MASA Metadata-Assisted Spatial Audio;
McMASA Multi-Channel MASA;
MC Multi-Channel;
OSBA Objects with SBA;
SBA Scene-Based Audio;
VBAP Vector Base Amplitude Panning VBR Variable Bit Rate.
3GPP Third Generation Partnership Project;
IVAS Immersive Voice and Audio Services;
FOA First-Order Ambisonics;
HOA Higher-Order Ambisonics;
HRIR Head-Related Impulse Response ;
HRTF Head-Related Transfer Function;
ITD Inter-Channel Time Delay;
ISM Independent Streams with Metadata;
LFE Low-Frequency Effects;
MASA Metadata-Assisted Spatial Audio;
McMASA Multi-Channel MASA;
MC Multi-Channel;
OSBA Objects with SBA;
SBA Scene-Based Audio;
VBAP Vector Base Amplitude Panning VBR Variable Bit Rate.
An overview of the audio processing functions of the receive side of the codec is shown in Fig. 1, with rendering features highlighted.
In Fig. 1, the interfaces are marked consistently with 5 using the following numbers:
3: Encoded audio frames (50 frames/s) , number of bits depending on IVAS codec mode;
4: Encoded Silence Insertion Descriptor (SID) frames;
5: RTP Payload packets/
6: Lost Frame Indicator (BFI) ;
7: Renderer config data;
8: Head-tracker pose information and scene orientation control data;
9: Audio output channels (16-bit linear PCM, sampled at 8 (only EVS) , 16, 32, or 48 kHz) , 10: Metadata
associated with output audio.
3: Encoded audio frames (50 frames/s) , number of bits depending on IVAS codec mode;
4: Encoded Silence Insertion Descriptor (SID) frames;
5: RTP Payload packets/
6: Lost Frame Indicator (BFI) ;
7: Renderer config data;
8: Head-tracker pose information and scene orientation control data;
9: Audio output channels (16-bit linear PCM, sampled at 8 (only EVS) , 16, 32, or 48 kHz) , 10: Metadata
associated with output audio.
For Multi-Channel (MC) Operation, coding of multi-channel inputs is available for the channel layouts 5.1, 7.1, 5.1+2, 5.1+4, and 7.1+4. The coding technique is selected from a set of coding modes based on the available bitrate and specified channel layout. The general principle in technique selection is to aim for best possible quality given the allowed bitrate. For all techniques, LFE channel coding is also offered either separately or within the technique. The multi-channel operation supports output to mono, stereo, multi-channel (at the same or any other layout with up to 16 speakers) , Ambisonics (at up to order 3) , and binaural.
For Scene-based Audio (Ambisonics) Operation, coding of ambisonics signals is supported for 1st-to 3rd-order inputs throughout the full bitrate range. The decoder's output can be SBA (of order 1, 2, or 3) , mono, stereo, binaural or multi-channel. This flexibility of input order, bitrate and output format combinations is in part achieved by the combination of covariance-based and directional analysis at different frequencies.
For the lowest 8 bands, covariance analysis with a 20ms stride is performed at the encoder and corresponding reconstruction is performed at the decoder. For the highest 4 bands, an estimation of the parameters of a psychoacoustic model is implemented with a time resolution of 5ms.
At the encoder, the covariance-analysis metadata for the higher bands are estimated from these model parameters and combined with the directly calculated metadata for the lower bands. Based on these metadata, a downmix to 1 to 4 channels (dependent on bitrate) is obtained. The downmix channels are then coded with the appropriate core coder.
At the decoder, the downmix channels plus the metadata are received. The latter comprise the transmitted covariance-analysis metadata and model parameters for the lower and higher bands, respectively. These metadata are used to reconstruct the HOA signal and render to the requested output format. In this, the psychoacoustic model parameters for the 8 lower frequency bands are estimated from the reconstructed audio channels. The model-based reconstruction allows for the output SBA order on the decoder side to be higher than the input order on the encoder side.
Low-latency operation (less than or equal to 38ms end-to-end) is achieved by using a very-low-latency 1-ms MDFT-based filterbank at the encoder and a 5-ms CLDFB-based filterbank at the decoder, which additionally enables the type of signal modifications that are necessary for rendering to a broad variety of output configuration.
For Metadata-assisted Spatial Audio (MASA) Operation, the IVAS codec supports coding of parametric spatial audio format called metadata-assisted spatial audio (MASA) . This format is specifically optimized for the direct immersive audio capture from smartphones and other form factors that can be unsuitable for dedicated spherical microphone arrays.
The MASA format is based on 1-2 audio channels and associated metadata that is provided for each audio frame. The MASA spatial metadata describes the spatial audio characteristics of the captured immersive audio using several spatial parameters including spatial direction information, directional and non-directional energy ratios, and two types of coherence information. The spatial metadata is provided in each frame according to a time-frequency resolution of 4 subframes and 24 frequency bands. The MASA descriptive metadata provides additional information relating to the creation and understanding of the MASA audio signal.
The coding in this operation is based on compression of the metadata exploiting detected redundancies and prioritization of selected parameters at each bitrate. The 1 or 2 audio transport channels are coded using the SCE and CPE coding capabilities. A joint bitrate allocation between these two coding blocks is based on a metadata analysis and simplification processing.
The MASA format can be flexibly rendered for binaural or loudspeaker reproduction, including mono and stereo playback. Rendering to Ambisonics is also supported. Furthermore, decoded MASA format bitstream can be directly output from the decoder without rendering as a fully compliant MASA format output for further processing.
Rendering is the process of generating digital audio output from the decoded digital audio signal. Rendering is used when output format is different than input format. In case output format is the same as input format, the decoded audio channels are simply passed through to the output channels. Binaural rendering is a special case, where binaural output channels are prepared for headphone reproduction. This process includes head-tracking and scene orientation control, head-related transfer function processing, and room acoustic synthesis. IVAS rendering is integrated with IVAS decoder but can also be operated standalone as external rendering while bypassing the internal renderer. The external renderer can be applied e.g., in the case of rendering outputs originating from multiple sources, such as decoders or audio streams.
The internal IVAS renderer is integrated into the IVAS decoder. In case of specific operating points, this integration allows for combining decoding and rendering processes, resulting in efficient processing.
The external IVAS renderer supports all the functionality of the internal renderer. However, since the external renderer operates stand-alone, combined decoding and rendering processing is not available.
Digital audio decoded by IVAS decoder can be rendered for loudspeaker reproduction. The process of rendering depends on the decoded audio format. In case of multi-channel formats, the decoded format can match the loudspeaker configuration or output channels can be generated by application of multi-channel conversion gains from conversion tables. For SBA, MASA and ISM formats the spatial audio needs to be mapped to the loudspeaker positions of the loudspeaker setup. Depending on the decoded format amplitude panning is employed, with either the vector-base amplitude panning (VBAP) scheme (using triangles) , or an improved edge-fading amplitude panning (EFAP) scheme (using polygons) . Specifically for SBA audio, the AllRAD loudspeaker decoding scheme is used with EFAP.
The time domain (TD) renderer operates on signals in time domain. In the IVAS decoder it is used for binaural rendering of discrete ISM, where each audio signal is encoded and decoded with a dedicated SCE module. This covers all ISM bit rates, except 3-4 objects for bit rates 24.4 kbps and 32 kbps. Further it is used in the decoder for binaural rendering of 5.1 and 7.1 signals when headtracking is enabled. In the external renderer it is used for all ISM configurations and all multichannel loudspeaker configurations, both with and without headtracking enabled. An overview of the TD binaural renderer is found in Fig. 2 below. An HRIR model accepts the object position metadata along with the headtracking data and generates an HRIR filter pair. The ITD may be modelled as a part of the HRIR, or it may be modelled as a separate parameter. In case an ITD parameter is output, the ITD synthesis is performed in the ITD synthesis stage. The time aligned signals are then convolved with the HRIR filter pair to form a binauralized signal.
The parametric binauralizer and stereo renderer operates on the following IVAS formats and operations: MASA, OMASA, multi-channel (in McMASA mode) , SBA, OSBA, and ISM, i.e., the input to the encoder has been audio signals (and potentially spatial metadata) in one of these formats, and it is now being rendered to binaural or stereo output. The IVAS format being processed (i.e., whether operating on MASA, OMASA, multi-channel (in McMASA mode) , SBA, OSBA, or ISM format) is obtained.
The parametric binaural (and stereo) rendering is performed in subframes, where m denotes the subframe index. A subframe contains Nslots CLDFB slots (in non-JBM operation, Nslots = 4, in JBM operation Nslots = 1 . . . 7) . The data determined at previous calls (i.e., subframes m-1 and earlier) affects the rendering of the present subframe m due to temporal averaging and interpolation. The binaural renderer system described in the following is also capable to render a stereo signal instead of a binaural signal. Also, binaural sound with and without room effect can be reproduced.
As an input, the renderer obtains (or receives) a spatial audio signal containing one or two transport audio signals and associated spatial metadata. The number of transport audio signals is one, when TODO: ADD REF TO CONFIG. The number of transport audio signals is two, when TODO: ADD REF TO CONFIG. The spatial metadata obtained (or received) by the renderer contains the following parameters: azimuth θ (b, m, i) , elevation φ (b, m, i) , direct-to-total energy ratio rdir (b, m, i) , spread coherence ζ (b, m, i) , and surround coherence γ (b, m) . In case of the SBA and OSBA formats, the spatial metadata contains SPAR metadata. The audio signals and the spatial metadata are used for providing spatial audio reproduction (i.e., to enable the rendering of spatial audio) .
In addition, the input to the renderer includes a separated centre channel audio signal when the IVAS format is multi-channel and operating in McMASA “separate channel” mode. In addition, the input to the renderer contains a separated object audio signal and associated object metadata when the IVAS format is OMASA and operating in “MASA one object coding mode” or “parametric one object coding mode” . In addition, the input to the renderer contains all object audio signals and associated object metadata when the IVAS format is OMASA and operating in “discrete coding mode” .
Furthermore, head orientation data and external orientation data may be received. Head tracking is described in clause 3.6.1, external orientation input is described in clause 3.6.2. Head tracking (and external orientation information) should be used if they are available for improved spatial audio experience. In the following it is referred to head orientation and head tracking for simplicity regardless of which components the orientation data is originally composed of, and which processing has been applied to it prior to the rendering step.
In addition, the renderer (when rendering binaural output) obtains a room effect control indication TODO: ADD REF TO CONFIG, and based on this indication, it is determined whether to apply a room effect to the input spatial audio signal or to not apply the room effect. For rendering binaural audio with a room effect, two data sets related to binaural rendering are obtained. The first is a pre-defined data set containing the HRTFs as spherical harmonics to binaural conversion matrices. The second data set contains binaural room responses as reverberation early part energy correction gains (which are used for modifying the resulting spectrum that is obtained from the rendering according to the first data set) , and late reverberation energy correction gains and reverberation times. Thus, the binaural output signals are generated from the transport audio signal (s) and associated metadata based on a combination of these two data sets.
The fetching of the temporally correct spatial metadata parameter values for the current subframe m so that they are in sync with the audio signals is handled in clause XXX. TODO: ADD REF. It operates differently for the JBM and non-JBM use. Fetching the correct spatial metadata values is not discussed in the following, it is assumed that it has already been correctly performed, as described in the aforementioned clause.
The binaural (or stereo) rendering is based on using covariance matrices. The transport signal covariance matrices are measured, and the target covariance matrices are determined. Based on these covariance matrices, processing matrices are determined to process the transport audio signals to a determined target. The details of this operation are described in the following.
IVAS rendering supports synthesis of room acoustics for realistic immersive effect. The room acoustics can be synthesized using room impulse response convolution or late reverb, optionally combined with early reflections. The room impulse response data (BRIRs) , late reverb and early reflection synthesis are driven by the set of parameters that are discussed in detail in rendering control section.
The feedback-delay network reverberator provided in IVAS decoder/renderer is based on the Jot reverberator. Additionally, filters have been added to control interaural correlation and ear-dependent coloration. A schematic depiction of the modified Jot reverberator is shown in Fig. 3.
In Fig. 3, The weights b and c control input and output of the feedback-delay network. Interaural coherence is controlled using u (z) and v (z) filter coefficients, while ear-dependent coloration using hL (z) and hR(z) . To match with the direct path binaural filter characteristics, the coloration filters are pre-computed based on reverb characteristics, and on HRIR used for binauralization.
The early reflections part of the room acoustics immersive effect can be enabled when using any multichannel input with binaural output modes. The reflection parameters described in section [xx] define a virtual 3D room with absorption coefficients representing the average broadband acoustic reflection characteristics of each surface in a geometric rectangular room model. The “shoebox” model computes first order early reflections for each source in the multichannel configuration using the image source method, described in [ref? ] . This is done by considering the cartesian position of the listener probe and the relative position of the multichannel emitter array within the virtual room. The listener position defaults to the centre of the room (at a standard height) but it can optionally be defined in the renderer parameters to be in a different location than the centre. Once computed, the resulting reflection gains and delay times are used by a process loop to create the reflection signals, totalling to six reflections per source. Each reflection signal is then mixed into the channel buffer that is closest to the direction of arrival, according to the configuration layout. The resulting mix of early and direct sound is thus jointly sent to the orientation rotation processing creating an interactive directional sound effect.
The virtual room is defined by three dimensions, respectively representing the length, width, and height of a rectangular room, in meter units. The resulting room model of dimensions R (x, y, z) is placed at a centre of a 3D coordinate system where the origin is the centre of the room floor. As seen in the figure below, the room can extend in the positive and negative directions of the x and y axis, but only in the positive direction for the z axis. Each wall surface is referenced by an index following the order of the axis from positive to negative (e.g. W: 5 and W: 6 reference the surfaces along the z axis, in this case ceiling and floor) , this permits the definition of independent absorption coefficients for each surface.
Coordinate system used by the Early Reflections model is shown in Fig. 4.
For head tracking, in addition to supporting head rotation processing, the IVAS renderer supports a number of modes for listener orientation tracking. The listener orientation tracking refers to a set of methods used to provide or estimate listener’s frontal orientation (being listener’s head neutral position or torso position) . Such a listener’s frontal orientation is further referred to as reference orientation. The following orientation tracking modes are supported:
External reference orientation,
External reference vector orientation,
External reference levelled vector orientation,
Adaptive long-term average reference orientation.
External reference orientation,
External reference vector orientation,
External reference levelled vector orientation,
Adaptive long-term average reference orientation.
For external reference orientation External reference vector orientation, the coordinate system of IVAS needs to be properly referenced and aligned, i.e. pivasforward: needs The input parameters to the reference vector orientation (identified as HEAD_ORIENT_TRK_REF_VEC) orientation tracking modes are:
The absolute position of the listeners head in Cartesian 3D coordinates (plistenerabs) ;
The absolute position of an acoustic reference (prefabs) . In case of camera-based head tracking by a UE,
where the UE should act as the acoustic reference direction, this would be the position of the phone in Cartesian 3D coordinates;
The absolute head orientation of the listener (rlistenerabs) ;
The position of the listener and the reference must refer to a common coordinate system, e.g., an Earth-
fixed coordinate system.
The absolute position of the listeners head in Cartesian 3D coordinates (plistenerabs) ;
The absolute position of an acoustic reference (prefabs) . In case of camera-based head tracking by a UE,
where the UE should act as the acoustic reference direction, this would be the position of the phone in Cartesian 3D coordinates;
The absolute head orientation of the listener (rlistenerabs) ;
The position of the listener and the reference must refer to a common coordinate system, e.g., an Earth-
fixed coordinate system.
For room acoustics parameters, the late reverb is driven by the set of parameters comprising of:
RT60 –indicating the time that it takes for the reflections to drop 60 dB in energy level;
DSR –diffuse to source signal energy ratio;
Pre-delay –delay at which the computation of DSR values was made. Can be interpreted as the threshold
between early reflections and late reverberation phase.
RT60 –indicating the time that it takes for the reflections to drop 60 dB in energy level;
DSR –diffuse to source signal energy ratio;
Pre-delay –delay at which the computation of DSR values was made. Can be interpreted as the threshold
between early reflections and late reverberation phase.
Spatialized, rotation-responsive, first-order early reflections can be added when using multichannel input (any configuration accepted) . The early reflection rendering is determined by several parameters that drive a shoebox model using the image-source method. The set of parameters consists of:
3D rectangular virtual room dimensions;
Broadband energy absorption coefficient per wall;
Listener origin coordinates within room (optional) ;
Low-complexity mode (optional) –favours efficient early reflection rendering over spatial accuracy.
3D rectangular virtual room dimensions;
Broadband energy absorption coefficient per wall;
Listener origin coordinates within room (optional) ;
Low-complexity mode (optional) –favours efficient early reflection rendering over spatial accuracy.
Room acoustics parameters are provided to the renderer as metadata. Two metadata formats are supported in the IVAS decoder/renderer implementation: binary renderer config metadata format, and text renderer config metadata format. Regardless of the metadata format, the general metadata processing is shared. Both metadata formats support multiple acoustic environment datasets, allowing for selecting between such acoustic environments.
In 3GPP TS 26.253 V1.0.0 (2023-12) , an immersive audio rendering system is mentioned as specification, that covers almost all 3DoF rendering techniques, such as listener head tracking, object position panning and room acoustic synthesis, aims at leading next generation audio technology. IVAS renderer consumes 3D audio like multi-channel, ISM, SBA, MASA and their combination as input signals, and then rendering these signals into output audio that includes binaural and loudspeaker format.
IVAS can be used not only in call areas, but also suitable for XR, games and broadcast industries. However, IVAS renderer just supports some common rendering abilities. When sound fields are rendered by IVAS, each item should be fixed in original orientation except for ISM. In single screen rendering scenario, if there is only one input sound field or all sound fields are rotated for the same angle, users can get rotated sound effects by reversely rotating listener head instead. But as far as multiple inputs, when each sound field has a different rotation angle, they must reproduce input signals to render the correct effect.
As mentioned above, in single screen rendering scenario, it tentatively believes that reproducing the content can meet the needs of sound field rotation. But in general dynamic scenarios, especially in mixed reality applications, input sound fields may be changing in real-time. Thus, rendering by reproducing content is no longer applicable. Furthermore, in multiple screen rendering scenarios, there does not exist a unified transformation relationship to meet different rotation angles of each sound field for each screen. Obviously, IVAS does not have the ability to rotate multiple sound fields independently to match these scenarios.
This disclosure, Audio Rendering Technique with Independent Rotation of Multiple Sound Fields, makes the problems above much easier. Users render rotated sound fields just need to change the orientation parameters, so users can render 3D audio by IVAS more flexible without reproducing content. Moreover, multiple screen rendering scenarios seem easily achievable.
In view of the above technical problems, the present disclosure provides a technical solution for audio processing that is capable of independently adjusting the direction of each sound field, thereby improving the effectiveness of audio rendering.
For example, the technical solution of the present disclosure can be realized by the embodiments in FIG. 5.
FIG. 5 shows a flow diagram of an audio processing method according to some embodiments of the present disclosure.
As shown in FIG. 5, in step 110, obtaining a corresponding rotation angle for each sound field of a plurality of sound fields of an audio.
In step 120, adjusting an orientation of the each sound field independently, according to the corresponding rotation angle.
In some embodiments, adjusting the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
For examples, rotating the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
For examples, rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel; or rotating the sound field as a whole in response to the sound field type being a Multi-Channel.
For examples, converting the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio; and adjusting the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
In some embodiments, the each sound field comprises a plurality of ambisonic signals, in response to the sound field type being a Scene-Based Audio. determining a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals; and determining a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
In some embodiments, adding a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated in response to the sound field type being a Multi-Channel. For examples, rendering each channel rotated as Independent Streams with Metadata.
In the above embodiments, it is possible to adaptively select a suitable adjusting method according to the type of sound field, thereby improving the performance of audio rendering.
In step 130, rendering the audio according to each sound field adjusted.
In some embodiments, calculating a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
In the following, some embodiments are used to illustrate, exemplarily, the rotation of each sound field independently of the other.
In this disclosure, each input sound field can be rotated independently, users can change input sound fields orientation during real-time rendering. In some embodiments, the schematic diagrams of differences between IVAS renderer and the renderer with independent rotation of multiple sound fields shown as Fig. 6 and Fig. 7.
In above figures, rotated sound fields can be SBA, MASA, Multi-Channel and their combination.
In some embodiments, quaternion or Euler angle is used to represent the orientation of sound field. For example, Euler angle is used to denote it, the orientation is denoted atand the relationship between original sound field and rotated sound field described below.
In the following, some embodiments are used to illustrate, exemplarily, that each sound field is rotated independently, according to the type of sound field.
In some embodiments, for SBA, a collection of rotation matrices is derived , denoted by Q, which, when applied to a set of ambisonics signals, rotate the sound field around the listener. That is, given a column-vector of ambisonics signals S, the rotated ambisonics signals are given by:
S′=Qn (Ω) ·S
S′=Qn (Ω) ·S
Where Q represents ambisonic rotation matrices, n represents ambisonic signal order. S represents original ambisonic signal, S'represents rotated ambisonic signal.
The elements of Q are denoted byand the arrangement of them is shown in Fig. 8a.
In some embodiments, for variable yaw rotation, the first rotation matrix derived is for an arbitrary azimuthal (yaw) rotation around the z-axis. Given a desired rotation angle α, the corresponding rotation matrix may be denoted by Q (α) , with elements given by
In some embodiments, the locations of possible nonzero elements (indicated by the symbols) of this matrix may be illustrated in Fig. 8b
In some embodiments, , the pitch rotation and roll rotation are similar to yaw rotation, and it's just that they have some rules of their own.
In some embodiments, for multi-channel, there are two ways to achieve sound field rotation, one is rotate the whole sound field by a panning matrix like Q in SBA rotation. The other one is rotating each channel forand the rendering each channel as ISM. Take 5.1 layout audio for example, the second method may be used to rotate sound field, the new orientation of each channel equals their original orientation plus the rotation angle that needs to be rotated, then send each channel audio into renderer as ISM. The new angle of each channel is given by:
Where Ωn represents channel direction of MC, represents rotation angle.
In some embodiments, for MASA, a metadata-assisted spatial audio format, this format is specifically optimized for direct immersive audio capture from smartphones and other form factors that can be unsuitable for dedicated spherical microphone arrays.
In some embodiments, MASA assists audio processing and optimization through the use of metadata, and MASA can be converted to SBA or MC. So rotating MASA is equal to rotate SBA or MC. On the basis of MC and SBA rotation, it is only needed to render MASA to SBA or MC. By the way, the IVAS renderer can cover MASA format conversion.
In some embodiments, in a case where several sound field should be rotated to different angles rendered by IVAS, each one corresponds to a new sound field by the above transform relationship, and then the new work may be used to continue rendering without changing other processes.
In some embodiments, VBAP gain determination may be performed as following: the input to the processing is a target panning direction comprising a target azimuth angle θ and target elevation anglerenderer control parameters include rotation parameters of input sound field, there need to rotate it separately according to the type of input audio; If the input is SBA format, the new input is given by:
S′=Qn (Ω) ·S
S′=Qn (Ω) ·S
Q represents ambisonic rotation matrices, n represents ambisonic signal order. S represents original ambisonic signal, S′ represents rotated ambisonic signal.
If the input is MC format, the new input is given by:
Ωn represents channel direction of MC, represents rotation angle.
If the input is MASA format, that should be converted to SBA or MC, then rotate the converted sound field by above known rules.
Then use the rotated sound fields as input to continue VBAP gain determination.
Thus, the best virtual surface triplet from the determined virtual surface arrangement is selected and the corresponding panning gains are generated based on it. In other words, triplet gains gtmp and the triplet to be used with respect to the present azimuth θ and elevationare formulated. The triplet node indices (a, b, c) for which the triplet gains gtmp are associated with are thus also determined.
In some embodiments, the render technique of headphone reproduction is formulated; the determining direction part gains is described and the new text after incorporating the content is as follows.
In the following, it is described how direct part gains are formulated for binaural or stereo output for a given direction. For determining the gains, the direction parameter is denoted azimuth θ and elevationand the output gains are denoted as a 2x1 column vector g. In this clause, the time and frequency indices, and most of the subscripts, are omitted for clarity of the equations.
If renderer control parameters include rotation parameters of input sound field, there need to rotate it separately according to the type of input audio.
If the input is SBA format, the new input is given by:
S′=Qn (Ω) ·S
S′=Qn (Ω) ·S
Q represents ambisonic rotation matrices, n represents ambisonic signal order. S represents original ambisonic signal, S′ represents rotated ambisonic signal.
If the input is MC format, the new input is given by:
Ωn represents channel direction of MC, represents rotation angle.
If the input is MASA format, that should be converted to SBA or MC, then rotate the converted sound field by above known rules.
Then use the rotated sound fields as input to continue the next processing.
First, the determining of stereo gains is described. In that operating mode, the azimuth and elevation are first mapped to a horizontal azimuth by:
If azimuth is between -θref < θmap < θref , where θref is 30 degrees, then the panning gains are formulated by:
a1 =tan (θmap) /tan (θref)
a2 = (a1 -1) / max (0.001, a1 + 1)
a3 = 1 / (a2 ^2 + 1)
a1 =tan (θmap) /tan (θref)
a2 = (a1 -1) / max (0.001, a1 + 1)
a3 = 1 / (a2 ^2 + 1)
When θmap ≥θref , then g= [1] . When θmap ≤ -θref , then g= [0] . 01.
In IVAS reference code, independent multiple sound field rotation techniques can be implemented by 3 functions interfaces. Corresponding API IVAS_REND_SetXXXOrientation should be added to source file lib_rend. h:
ivas_error IVAS_REND_SetMCOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) ;
ivas_error IVAS_REND_SetSBAOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) ;
ivas_error IVAS_REND_SetMASAOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) .
ivas_error IVAS_REND_SetMCOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) ;
ivas_error IVAS_REND_SetSBAOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) ;
ivas_error IVAS_REND_SetMASAOrientation (IVAS_REND_HANDLE hIvasRend,
const
IVAS_REND_InputId inputId,
const
IVAS_QUATERNION inputRotation) .
According to some embodiments of the present disclosure, there is provided an audio processing method, comprising: determining a collection of rotation matrices; and applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
In some embodiments, the applying the collection of rotation matrices to a set of ambisonics signals comprises: determining rotated ambisonics signals, according to a column-vector of ambisonics signals and the collection of rotation matrices.
In some embodiments, the sound field comprises at least one of SBA, MASA or Multi-Channel.
According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, comprising: determining module for determining a collection of rotation matrices; and applying module for applying the collection of rotation matrices to a set of ambisonics signals to rotate a sound field around a listener.
In some embodiments, the applying module determining rotated ambisonics signals, according to a column-vector of ambisonics signals and the collection of rotation matrices.
In some embodiments, the sound field comprises at least one of SBA, MASA or Multi-Channel.
According to still other embodiments of the present disclosure, there is provided an electronic device, comprising: a memory; a processor coupled to the memory, the processor configured to, based on instructions stored in the memory, carry out the audio processing method according to any one of the above embodiments.
According to still other embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements the audio processing method according to any one of the above embodiments.
According to still other embodiments of the present disclosure, there is provided a computer program product, comprising: instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of the above embodiments.
Fig. 9 shows a block diagram of an audio processing apparatus according to some embodiments of the present disclosure.
As shown in Fig. 9, an audio processing apparatus 9, comprising: obtaining unit 91, configured to obtain a corresponding rotation angle for each sound field of a plurality of sound fields of an audio; adjusting unit 92, configured to adjust an orientation of the each sound field independently, according to the corresponding rotation angle; and rendering unit 93, configured to render the audio according to each sound field adjusted.
In some embodiments, the adjusting unit 92 adjusts the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
In some embodiments, the adjusting unit 92 rotates the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
In some embodiments, the each sound field comprises a plurality of ambisonic signals, and the adjusting unit 92 determines a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals and determines a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
In some embodiments, the adjusting unit 92 rotates the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel (MC) .
In some embodiments, the adjusting unit 92 adds a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated.
In some embodiments, the rendering unit 93 renders each channel rotated as Independent Streams with Metadata.
In some embodiments, the adjusting unit 92 converts the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio and adjusts the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
In some embodiments, the rendering unit 93 calculates a panning gain according to an orientation of the each sound field adjusted; and rendering the audio according to the panning gain.
FIG. 10 shows a block diagram of the electronic device according to other embodiments of the present disclosure.
As shown in FIG. 10, the electronic device 6 comprises: a memory 61 and a processor 62 coupled to the memory 61, the processor 62 configured to, based on instructions stored in the memory 61, carry out the audio processing method according to any one of the embodiments of the present disclosure.
Where the memory 61 may comprise, for example, system memory, a fixed non-transitory storage medium, or the like. The system memory stores, for example, an operating system, applications, a boot loader, a database, and other programs.
FIG. 11 shows a block diagram of the electronic device according to further embodiments of the present disclosure.
As shown in FIG. 11, the apparatus 7 for generating the fitness regimen information of this embodiment comprises: a memory 710 and a processor 720 coupled to the memory 710, the processor 720 configured to, based on instructions stored in the memory 710, carry out the audio processing method according to any one of the embodiments of the present disclosure.
The memory 710 may comprise, for example, system memory, a fixed non-transitory storage medium, or the like. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
The apparatus 7 for generating the fitness regimen information may further comprise an input-output interface 730, a network interface 740, a storage interface 750, and the like. These interfaces 730, 740, 750, the memory 710 and the processor 720 may be connected through a bus 760, for example. Where the input-output interface 730 provides a connection interface for input-output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, a loudspeaker, etc. The network interface 740 provides a connection interface for various networked devices. The storage interface 750 provides a connection interface for external storage devices such as an SD card and a USB flash disk.
Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, embodiments of the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment containing both hardware and software elements. Moreover, the present disclosure may take the form of a computer program product embodied on one or more computer-usable non-transitory storage media (comprising but not limited to disk storage, CD-ROM, optical memory, etc. ) having computer-usable program code embodied therein.
Heretofore, the method, apparatus and generation system of fitness regimen information and the non-transitory computer-readable storage medium according to the present disclosure have been described in detail. In order to avoid obscuring the concepts of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can understand how to implement the technical solutions disclosed herein.
The method and system of the present disclosure may be implemented in many ways. For example, the method and system of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above sequence of steps of the method is merely for the purpose of illustration, and the steps of the method of the present disclosure are not limited to the above-described specific order unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which comprise machine-readable instructions for implementing the method according to the present disclosure. Thus, the present disclosure also covers a recording medium storing programs for executing the method according to the present disclosure.
Although some specific embodiments of the present disclosure have been described in detail by way of example, those skilled in the art should understand that the above examples are only for the purpose of illustration and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure.The scope of the disclosure is defined by the following claims.
Claims (13)
- An audio processing method, comprising:obtaining a corresponding rotation angle for each sound field of a plurality of sound fields of an audio;adjusting an orientation of the each sound field independently, according to the corresponding rotation angle; andrendering the audio according to each sound field adjusted.
- The audio processing method according to claim 1, wherein the adjusting the orientation of the each sound field independently, according to the corresponding rotation angle comprises:adjusting the orientation of the each sound field according to an adjusting method corresponding to a sound field type of the each sound field, the adjusting method comprising rotating the each sound field as a whole, and rotating each channel of a plurality of channels of the each sound field independently.
- The audio processing method according to claim 2, wherein the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises:rotating the sound field as a whole in response to the sound field type being a Scene-Based Audio (SBA) .
- The audio processing method according to claim 3, wherein the each sound field comprises a plurality of ambisonic signals, andthe rotating the sound field as a whole in response to the sound field type being the Scene-Based Audio (SBA) comprises:determining a corresponding rotation matrix of the plurality of ambisonic signals according to the corresponding rotation angle and an order of the plurality of ambisonic signals; anddetermining a plurality of ambient acoustic signals rotated according to the plurality of ambisonic signals and the corresponding rotation matrix.
- The audio processing method according to claim 2, wherein the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises:rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being a Multi-Channel (MC) .
- The audio processing method according to claim 5, wherein the rotating the each channel of the plurality of channels of the each sound field independently in response to the sound field type being the Multi-Channel (MC) comprises:adding a current angle of the each channel to the corresponding rotation angle to obtain each channel rotated.
- The audio processing method according to claim 5, wherein the rendering the audio according to the each sound field adjusted comprises:rendering each channel rotated as Independent Streams with Metadata.
- The audio processing method according to claim 2, wherein the adjusting the orientation of the each sound field according to the adjusting method corresponding to the sound field type of the each sound field comprises:converting the each sound field to a SBA or a MC in response to the sound field type of the each sound field being a Metadata-Assisted Spatial Audio; andadjusting the orientation of the each sound field according to an adjusting method corresponding to the SBA or the MC.
- The audio processing method according to any one of claims 1 to 8, wherein the rendering the audio according to each sound field adjusted comprising:calculating a panning gain according to an orientation of the each sound field adjusted; andrendering the audio according to the panning gain.
- An audio processing apparatus, comprising:obtaining unit, configured to obtain a corresponding rotation angle for each sound field of a plurality of sound fields of an audio;adjusting unit, configured to adjust an orientation of the each sound field independently, according to the corresponding rotation angle; andrendering unit, configured to render the audio according to each sound field adjusted.
- An electronic device, comprising:a processor;a memory for storing processor executable instructions;wherein the processor is used to read the executable instructions from the memory and execute the instructions to implement an audio processing method of any one of claims 1 to 9.
- A computer readable storage medium storing thereon a computer program that, when executed by a processor, causes the processor to implement an audio processing method of any one of claims 1 to 9.
- A computer program product, comprising:instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of claims 1 to 9.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN2024084739 | 2024-03-29 | ||
| CNPCT/CN2024/084739 | 2024-03-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025201411A1 true WO2025201411A1 (en) | 2025-10-02 |
Family
ID=97217260
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/085050 Pending WO2025201411A1 (en) | 2024-03-29 | 2025-03-26 | Audio processing method and apparatus, electronic device, computer readable storage medium and computer program product |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025201411A1 (en) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105120421A (en) * | 2015-08-21 | 2015-12-02 | 北京时代拓灵科技有限公司 | Method and apparatus of generating virtual surround sound |
| US20160036987A1 (en) * | 2013-03-15 | 2016-02-04 | Dolby Laboratories Licensing Corporation | Normalization of Soundfield Orientations Based on Auditory Scene Analysis |
| CN105376691A (en) * | 2014-08-29 | 2016-03-02 | 杜比实验室特许公司 | Direction-aware surround sound playback |
| CN109302525A (en) * | 2017-07-25 | 2019-02-01 | 西安中兴新软件有限责任公司 | A kind of method and multi-screen terminal playing sound |
| CN116193196A (en) * | 2023-02-16 | 2023-05-30 | 阿里巴巴(中国)有限公司 | Virtual surround sound rendering method, device, equipment and storage medium |
-
2025
- 2025-03-26 WO PCT/CN2025/085050 patent/WO2025201411A1/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160036987A1 (en) * | 2013-03-15 | 2016-02-04 | Dolby Laboratories Licensing Corporation | Normalization of Soundfield Orientations Based on Auditory Scene Analysis |
| CN105376691A (en) * | 2014-08-29 | 2016-03-02 | 杜比实验室特许公司 | Direction-aware surround sound playback |
| CN105120421A (en) * | 2015-08-21 | 2015-12-02 | 北京时代拓灵科技有限公司 | Method and apparatus of generating virtual surround sound |
| CN109302525A (en) * | 2017-07-25 | 2019-02-01 | 西安中兴新软件有限责任公司 | A kind of method and multi-screen terminal playing sound |
| CN116193196A (en) * | 2023-02-16 | 2023-05-30 | 阿里巴巴(中国)有限公司 | Virtual surround sound rendering method, device, equipment and storage medium |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11950085B2 (en) | Concept for generating an enhanced sound field description or a modified sound field description using a multi-point sound field description | |
| RU2759160C2 (en) | Apparatus, method, and computer program for encoding, decoding, processing a scene, and other procedures related to dirac-based spatial audio encoding | |
| US9552819B2 (en) | Multiplet-based matrix mixing for high-channel count multichannel audio | |
| KR102540642B1 (en) | A concept for creating augmented sound field descriptions or modified sound field descriptions using multi-layer descriptions. | |
| CN101490743B (en) | Dynamic decoding of binaural audio signals | |
| US20080298610A1 (en) | Parameter Space Re-Panning for Spatial Audio | |
| JP7818660B2 (en) | Spatial Audio Representation and Rendering | |
| US10764709B2 (en) | Methods, apparatus and systems for dynamic equalization for cross-talk cancellation | |
| CN111295896A (en) | Virtual rendering of object-based audio on arbitrary sets of speakers | |
| US11483669B2 (en) | Spatial audio parameters | |
| EP4128824A1 (en) | Spatial audio representation and rendering | |
| JP2022552474A (en) | Spatial audio representation and rendering | |
| US20210250717A1 (en) | Spatial audio Capture, Transmission and Reproduction | |
| US12300215B2 (en) | Spatial audio reproduction by positioning at least part of a sound field | |
| KR20240152893A (en) | Parametric spatial audio rendering | |
| WO2025201411A1 (en) | Audio processing method and apparatus, electronic device, computer readable storage medium and computer program product | |
| KR20190060464A (en) | Audio signal processing method and apparatus | |
| WO2025232856A1 (en) | Audio processing method and apparatus | |
| GB2627482A (en) | Diffuse-preserving merging of MASA and ISM metadata | |
| WO2025232857A1 (en) | Audio processing method and apparatus | |
| RU2809609C2 (en) | Representation of spatial sound as sound signal and metadata associated with it | |
| GB2639905A (en) | Rendering of a spatial audio stream | |
| CN121312155A (en) | Audio rendering method, apparatus and non-volatile computer-readable storage medium | |
| CN117917901A (en) | Generating a parametric spatial audio representation | |
| EA053181B1 (en) | AUDIO ENCODING AND DECODING USING REPRESENTATION TRANSFORMATION PARAMETERS |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25778318 Country of ref document: EP Kind code of ref document: A1 |