EP4736159A1 - Apparatus, methods and computer program for encoding spatial audio content - Google Patents
Apparatus, methods and computer program for encoding spatial audio contentInfo
- Publication number
- EP4736159A1 EP4736159A1 EP24732620.0A EP24732620A EP4736159A1 EP 4736159 A1 EP4736159 A1 EP 4736159A1 EP 24732620 A EP24732620 A EP 24732620A EP 4736159 A1 EP4736159 A1 EP 4736159A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- encoding
- alternative
- format
- input format
- spatial audio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S5/00—Pseudo-stereo systems, e.g. in which additional channel signals are derived from monophonic signals by means of phase shifting, time delay or reverberation
- H04S5/005—Pseudo-stereo systems, e.g. in which additional channel signals are derived from monophonic signals by means of phase shifting, time delay or reverberation of the pseudo five- or more-channel type, e.g. virtual surround
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/173—Transcoding, i.e. converting between two coded representations avoiding cascaded coding-decoding
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Stereophonic System (AREA)
Abstract
Examples of the disclosure relate to encoding spatial audio content using immersive voice and audio services (IVAS) codec. In examples an apparatus is configured to obtain a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content. The apparatus is also configured to receive an indication of an output format configured to be used by a playback device. The apparatus is also configured to upmix the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
Description
TITLE
Apparatus, Methods and Computer Program for Encoding Spatial Audio Content
TECHNOLOGICAL FIELD
Examples of the disclosure relate to apparatus, methods and computer programs for encoding spatial audio content. Some relate to apparatus, methods and computer programs for encoding spatial audio content using immersive voice and audio services (IVAS) codec.
BACKGROUND
Spatial audio content can be used for immersive voice and audio services such as immersive voice and audio for virtual reality or mediated reality. The methods used to encode the audio content can affect the audio quality perceived by a listener.
BRIEF SUMMARY
According to various, but not necessarily all, examples of the disclosure there is provided an apparatus comprising means for: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
The means may be for switching the encoding option to an alternative encoding option and wherein the switching is based, at least in part, on the output format used by the playback device and the alternative input format; and encoding spatial audio content using the alternative encoding option.
The alternative encoding option may enable a higher bit rate to be supported.
The means may be for receiving a list of supported modes for encoding spatial audio content and sorting the list into a preferred order and adding one or more alternative modes to the list where the alternative modes comprise multi-channel formats that support a higher bit rate.
A mode for encoding spatial audio content may comprise an input/output format and an encoding option.
The indication of an output format used by the playback device may comprise a value of a parameter.
The means may be for mixing the spatial audio content such that additional channels of the alternative input format do not contain any content.
The means may be for mixing the spatial audio content such that additional channels of the alternative input format do contain some content.
The means may be for mixing the spatial audio content such that no change is made to content of channels of the selected input format when upmixing to the alternative input format.
The means may be for mixing the spatial audio content such that at least some changes are made to content of the channels of the selected input format when upmixing to the alternative input format.
The alternative input format may comprise a format that provides at least two additional channels.
The output format used by the playback device may be a binaural format and the updated encoding option may comprise multichannel coding using metadata-assisted spatial audio.
The means may be for enabling switching to the alternative encoding option if it is determined that the alternative encoding option supports a higher bit rate.
The means may be for determining the alternative input format and encoding option to be used for spatial audio encoding through negotiation with a playback device.
The means may be for determining the alternative input format and encoding option to be used for spatial audio encoding based on the indication of the output format used by the playback device.
According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and
upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all of the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all of the features, in any combination, may be implemented by/comprised in/performable by an apparatus, a method, and/or computer program instructions as desired, and as appropriate.
BRIEF DESCRIPTION
Some examples will now be described with reference to the accompanying drawings in which:
FIG. 1 shows an example of channel-based audio input coding;
FIG. 2 shows an example audio output system;
FIG. 3 shows an example method;
FIG. 4 shows an example use case scenario;
FIG. 5 shows an example use case scenario;
FIG. 6 shows an example use case scenario;
FIG. 7 shows an example method;
FIG. 8 shows an example system; and
FIG. 9 shows an example apparatus.
The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures.
DETAILED DESCRIPTION
Examples of the disclosure can be implemented in audio systems that make use of the immersive voice and audio services (IVAS) codec.
Fig. 1 schematically shows an example of channel-based audio input coding.
Fig. 1 shows a first input 101 A and a second input 101 B. The different inputs 101 A, 101 B can comprise different input formats. The input formats are the formats that are used for signals that are provided to an encoder. The inputs 101A, 101 B could be provided as inputs to an IVAS encoder.
Different input formats and combinations of input formats can be supported by IVAS. The different input formats and combinations of input formats can comprise stereo, multichannel (MC), object-based audio, Scene-Based Audio (SBA), Metadata Assisted Spatial Audio (MASA), Objects with MASA (OMASA) or any other suitable format or combination of formats.
Stereo input refers to audio representation, where a first channel of audio is assigned to a left channel and a second channel of audio is assigned to a right channel.
Multichannel (MC) input refers to audio representation, where each transported channel represents an audio signal for a loudspeaker positioned around a listener. Example MC formats that can be supported by IVAS are surround formats 5.1 and 7.1 and surround formats with elevated speaker positions 5.1.2, 5.1.4 and 7.1.4.
Object-based audio, or Independent Streams with Metadata (ISM), input refers to audio representation, where individual mono audio object streams are transmitted. In addition to the transported audio, metadata describing the audio objects is transmitted. The metadata can comprise information such as the azimuth and elevation of the audio object or any other suitable information.
Scene-based audio (SBA) input refers to Ambisonics-based audio representation.
Ambisonics signals carry a representation of the audio scene, where the transport
channels refer to capturing directions in a spherical domain. The first channel (W) represents the omnidirectional capture. The omnidirectional capture is the incoming sound field from all directions. The next three channels (X, Y, Z) represent the incoming sound from the corresponding spatial axes. These four channels form the first order Ambisonics (FOA) representation. A higher spatial accuracy can be achieved by increasing the number of capturing directions with more channels. This increases the order of the Ambisonics representation, referred to as higher order Ambisonics (HOA). Second order Ambisonics (HOA2) comprises nine channels, and third order (HOA3) comprises sixteen channels. IVAS supports first, second and third order Ambisonics.
MASA refers to a parametric spatial audio representation. MASA uses audio signal(s) together with corresponding spatial metadata. The spatial metadata can comprise information such as directions and direct-to-total energy ratios in frequency bands or any other suitable information. A MASA stream can be obtained by capturing spatial audio with microphones and then estimating the spatial metadata based on the microphone signals. In some examples a MASA stream can be obtained from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (such as 5.1 mix) or other content by means of a suitable format conversion.
OMASA refers to an input comprising of MASA with additional object-based audio. The object based audio can comprise between one and four objects. The object-based audio streams can be provided to an encoder as separate streams from the MASA stream, and these can be encoded together.
Different encoding options are available to be used for encoding audio content with the respective input formats. The different encoding options can comprise different coding technologies. As shown in Fig. 1 different encoding options can be used based on factors such as the input format and the bit rate and any other suitable factor or combination of factors. In the example of Fig. 1 a first encoding option 103A is indicated by the hashed portion and a second encoding option is indicated by the dotted portion 103B.
Examples of encoding options that can be supported by IVAS comprise McMASA, ParamMC, MC Param Uprnix, MCT and any other suitable technology. Multichannel
coding using metadata-assisted spatial audio (McMASA) refers to multichannel encoding and format perceptually represented using a MASA stream. In some configurations, McMASA can also include a separated channel such as the center channel C. ParamMC refers to parametric multichannel which encodes a reduced set of transport channels and parameters describing channel relations. Multichannel Parametric Uprnix (MC Param Uprnix) refers to another type of parametric multichannel encoding which encodes a reduced set of transport channels and parameters, which are used to uprnix back to original configuration. MCT refers to multichannel coding tool and represents discrete channel coding using single channel or channel pair coding tools. The different MC coding technologies for different surround input formats and their active bitrates are presented in Table 1.
Table 1
Fig. 2 shows an example audio output system 200 that could be used to implement some examples of the disclosure. In this example the audio output system 200 comprises a 7.1.4 surround system. The 7.1.4 surround system comprises three front channels 202, four surround channels 204, a Low Frequency Effect (LFE) channel 206, and four height channels 208. The front channels 202 comprise a left front channel 202_L, a center front channel 202_C, and right front channel 202_R. The surround channels 204 comprise a left surround channel 202_LS, a left rear surround channel 202_LRS, a right surround channel 202_RS, and right rear surround channel 202_RRS. The height channels 208 comprise a left front channel 208_LFH, a left rear channel 208_LFR, a right front channel 208_RFH, a right rear channel 208_RRH.
The LFE channel 206 is not restricted to a particular location in space but could be positioned anywhere within the system 200. In the example of Fig. 2, and in other examples in this disclosure the LFE channel 206 is positioned next to the right front channel 202_R. The LFE channel 206 could be in other locations in other examples.
Fig. 2 also shows a listener 210 that is listening to the audio content that can be output by the audio output system 200.
Other types of audio output systems 200 can also be used. For instance, a 7.1 system comprises the three front channels 202, four surround channels 204, and an LFE channel 206 but would not comprise the height channels 208.
For a bitrate of 80 kbps ParamMC encoding options could be used for 5.1 input formats. However, for binaural outputs, the McMASA encoding options would be preferable because binaural rendering does not aim to truthfully reconstruct all the channels of the audio content but instead aims to reproduce the corresponding perception via headphones. McMASA has the additional benefit of being computationally more efficient, and thus saves battery life. Examples of the disclosure address this issue.
Fig. 3 shows an example method that can be used to implement examples of the disclosure. The method could be implemented by an apparatus within an encoding device or within any suitable type of device.
At block 300 the method comprises obtaining a selected input format and an encoding option for encoding spatial audio content. The encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content.
An input format specifies the format that is used for signals that are provided to an encoder. The input formats can comprise mono, stereo, multichannel, MASA, FOA, HOA2, HOA2, HOA3, binaural or any other suitable format.
An encoding option is used for encoding audio content. An encoding option can comprise a coding technology. Different encoding options could be McMASA, ParamMC, MC Param Uprnix, MCT or any other suitable encoding option.
The selected input format and the encoding option can be obtained by being negotiated between an encoding device and a playback device. A playback device is a device that is configured to playback audio for the listener 210. The playback device can be configured to generate an audio output that the listener can hear. The playback device can comprise an audio output system or can be coupled to an audio output system. The encoding device can be a device that encodes the audio into an audio signal that can be transmitted to the playback device. The negotiation can occur during establishment of a session between the respective devices.
The selected input format can comprise a number of channels.
At block 302 the method comprises receiving an indication of an output format configured to be used by the playback device. The indication of an output format used by the playback device can comprise a value of a parameter or any other suitable information.
An output format specifies the format that is used for signals that are provided from a decoder. The output formats can comprise mono, stereo, 5.1. 7.1 , 7.1.4, multichannel, FOA, HOA2, HOA2, HOA3, binaural or any other suitable format. The output formats can comprise rendered formats that can be listened to (e.g., binaural) or formats that require rendering before the output can be listened to, such as FOA.
In some examples the type of playback device that is to be used might determine the output format that is used or the output formats that are available for use.
At block 304 the method comprises upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format. In some examples the alternative input format comprises a format that provides at least two additional channels. The alternative input format is selected based, at least in part, on the output format.
In some examples the method can also comprise switching the encoding option to an alternative encoding option. That is, both the input format and the encoding option can be changed from the originally selected input format and encoding option in some examples of the disclosure.
In some examples the switching to the alternative encoding option can be enabled if it is determined that one or more conditions are satisfied by the alternative encoding option. For instance, in some examples the switching to an alternative encoding option can be enabled if it is determined that the alternative encoding option supports a higher bit rate. Other scenarios in which a switch to an alternative encoding option could be made could be, if it is determined that the alternative encoding option results in lower complexity or if it is determined that the alternative encoding option results in a lower transmission bitrate.
The alternative input format and/or the encoding option to be used for spatial audio encoding can be determined using any suitable means or processes. In some examples the alternative input format and/or the encoding option to be used for spatial audio encoding can be determined through negotiation with a playback device. In some examples the alternative input format and/or the encoding option to be used for spatial audio encoding can be determined based on the indication of the output format used by the playback device.
In some examples the output format used by the playback device could be a binaural format and the updated encoding option could comprise multichannel coding using
metadata-assisted spatial audio. Other formats and encoding options could be used in other examples.
The switching of the encoding option can be based, at least in part, on the output format used by the playback device and the alternative input format. The spatial audio content can then be encoded using the alternative encoding option. The audio content that has been encoded using the alternative encoding option can then be sent to the playback device for playback to a listener 210. Before the audio is played back there would be decoding and rendering. The decoding and rendering could be performed by different device or by a single device.
In some examples the method can comprise receiving a list of supported modes for encoding spatial audio content. A mode for encoding spatial audio content can comprise an input/output format and an encoding option. The method can comprise sorting the list into a preferred order. The order of preference can be determined using any suitable parameters. The method also comprises adding one or more alternative modes to the list where the alternative modes comprise multi-channel formats that support a higher bit rate than the originally selected combination of input format and encoding option.
In examples of the disclosure the upmixing to the selected input format and/or the switching to an alternative encoding option can enable a higher bit rate to be supported by the same encoding option compared to the originally selected input format and encoding option.
When the spatial audio content is being upmixed to the selected input format there are different ways in which the mixing can be performed. In some examples the mixing of the spatial audio content can be such that additional channels of the alternative input format do not contain any content. For instance, if the alternative input format comprises additional channels then these additional channels could comprise only zeroes (i.e., “digital silence").
In other examples the mixing of the spatial audio content can be such that additional channels of the alternative input format do contain some content. For instance, if the
alternative input format comprises additional channels then these additional channels could comprise some non-zero values (i.e. “non-silent channels”).
In some examples the mixing of the spatial audio content can be such that no change is made to content of channels of the selected input format when upmixing to the alternative input format. That is the content that is mixed to specific channels in the original input format is still mixed to those channels in the alternative input format.
In other examples the mixing of the spatial audio content can be such that at least some changes are made to content of the channels of the selected input format when upmixing to the alternative input format. That is, at least some of the content that is mixed to specific channels in the original input format would be mixed to different channels in the alternative input format.
Fig. 4 shows an example use case scenario that illustrates problems that can be addressed by examples of the disclosure. In this example the input format used for encoding spatial audio content is 7.1 . The input format comprises three front channels 202, four surround channels 204, and a Low Frequency Effect (LFE) channel 206. The front channels 202 comprise a left front channel 202_L, a center front channel 202_C, and right front channel 202_R. The surround channels 204 comprise a left surround channel 202_LS, a left rear surround channel 202_LRS, a right surround channel 202_RS, and right rear surround channel 202_RRS.
In this example the bitrate can be 64kbps. Using the data in table 1 the encoding option selected would be ParamMC.
This would therefore provide a mode for encoding of 7.1 with ParamMC.
In the use case scenario shown in Fig. 4 the listener 210 is using headphones 400 rather than, e.g., a stereo or multichannel loudspeaker system. The headphones 400 provide a playback device for playing back the spatial audio to the user. In this example a binaural output format is used for the spatial audio content. Other output formats could be used in other examples.
The headphones 400 can provide a receiving device that receives the encoded spatial audio content. In other examples there could be an intermediate receiving device such as a mobile phone or other similar device. The receiving device can be configured to receive the encoded content and decode the content. In some examples the receiving device can also be configured to render the decoded content. The decoded content or the rendered content can then be provided from the receiving device to the headphones 400. The connection between the receiving device and the headphones 400 could be a wired or wireless connection. In case of wireless connection, there could be a second encoding/decoding for local connectivity purposes. This could use a high bit rate and some other codec such as a codec part of Bluetooth profile.
The use of 7.1 with ParamMC for encoding the audio content and then using binaural format for the playback of the audio content would result in a lower perceptual quality for the listener 210 using the headphones 400 than what the IVAS codec is capable of. In other words, the IVAS codec capability is not utilized optimally by encoding. This is indicated by the listener 210 being unhappy in Fig. 4.
Fig. 5 shows another example use case scenario that provides a higher perceptual quality for the listener 210 using the headphones 400.
In the example of Fig. 5 the input format used for encoding spatial audio content is 7.1.4. The input format comprises three front channels 202, four surround channels 204, a Low Frequency Effect (LFE) channel 206, and four height channels 208. These can be arranged as shown in Fig. 2.
In this example the bitrate can also be 64kbps. Using the data in table 1 the encoding option selected would now be McMASA. This would therefore provide a mode for encoding of 7.1.4 with McMASA. This provides an improved perceptual quality for the listener 210 using the headphones 400 compared to the example shown in Fig. 4. The use of the McMASA encoding option can help to improve the perceptual quality because the binaural rendering aims to reproduce the perception of the audio content rather than reconstruct the original channels. McMASA is also a perceptually motivated model with no physical dependency relating to the number of input channels and so can provide the improved perceptual quality for a listener 210. McMASA can
also be more computationally efficient and so could also provide the additional benefit of saving power.
Fig. 6 shows the audio content being upmixed from a first input format to an alternative input format with additional channels. In this example, the audio content is upmixed from 7.1 to 7.1.4.
In this example the upmixing comprises adding the additional height channels 208. The additional height channels 208 are indicated in dashed lines in Fig. 6 to indicate that these channels are additional. The additional channels can comprise no content or only a small amount of content. However, by changing to the alternative input format with the additional channels a different encoding option can be selected at given higher bit rate which can improve the perceptual audio quality for the listener 210.
In examples of the disclosure an initial input format and/or encoding option can be switched to an alternative input format and/or encoding option. For example, and initial input format of 7.1 could be switched to an alternative input format such as 7.1.4. The additional channels of 7.1.4 help to improve the perceptual quality for the listener 210 using a binaural output. The input format can be upmixed to an alternative format by adding additional channels. The additional channels can be empty or can include some content.
Other types of input formats and alternative input formats can also be used. For example, a 5.1 input format or a 5.1.2 input format could be upmixed to a 5.1.4 input format. The 5.1.4 input format has two or four additional channels, respectively. The additional channels can be empty or can include some content.
In examples of the disclosure the selection of the alternative input format and/or an alternative encoding option can be made based on the bitrate or any other suitable factor.
The selection of the alternative input format can enable an encoding device to select a mode for encoding that is better suited for the particular output format used by the playback device. For example, the mode for encoding can be better suited for binaural
outputs. In some examples the mode for encoding can be selected because it has lower computation complexity or for any other suitable reason.
The selection of the alternative input format can occur in response to the encoding device receiving an indication of the output format used by a playback device. For example, a playback device can indicate to the encoding device that it is consuming audio in a binaural output format or in any other suitable format. The indication can be made during session negotiation or at any other suitable time.
In some examples the selection of the alternative input format can be made based on session negotiation offer and answer. The session negotiation and answer can be made using Session Description Protocol (SDP) or any other suitable protocol. In such session negotiations and answer the playback device can indicate its output capability to the encoding device. The input format and/or the encoding technology can then be updated or switched using an audio input pre-processor or any other suitable means.
Fig. 7 shows an example method that can be used for the session negotiation and answer in some examples of the disclosure. In the example method of Fig. 7 blocks 700 to 706 and blocks 712 to 716 can be performed by the encoding device and blocks 708 and 710 can be performed by the receiving device or a playback device.
At block 700 the encoding device obtains the one or more audio input formats that are supported. The supported input formats can comprise formats with different amounts of channels. For example, the supported input formats can comprise 5.1 , 5.1.2, 5.1.4, 7.1 , 7.1.4 multichannel formats or any other suitable format.
At block 702 the one or more supported input formats are sorted into a preferred order. Any suitable criteria can be used to define the order of preference of the list. The criteria can comprise the encoding computation complexity, the audio capture capabilities or any other suitable criteria.
In some examples additional input formats can be added to the list. For instance, one or more multichannel input formats can be added. One or more multichannel input formats can be added if it corresponds to a higher bit rate being supported for a
particular encoding option. The particular encoding option could be McMASA which could be used if the output format is to be binaural. Other particular encoding options could be used in other examples.
At block 704 the supported input formats are included in the session description file. Any suitable means can be used to include such information in the session description file. For instance, an input format attribute can be included in the session description file. The input format attribute can then be populated with the sorted list of supported audio input formats.
The session description file can comprise the input format indication, output format indication, format switching for bitrate adaptation as an attribute or media format parameter.
At block 706 a session negotiation offer can then be generated and transmitted. The session negotiation description can be represented as a session description protocol (SDP) file or in any other suitable file. The session negotiation can be performed as session description offer answer model or using any other suitable procedure.
At block 708 the receiving device receives the session negotiation offer. At block 710 the receiving device responds to the session negotiation offer by sending an indication of an output format capability of the receiving device. The indication of the output format capability can be provided in a session negotiation answer. In some examples the session negotiation answer can comprise a modified list of preferred input formats. The list can be modified into the preferred order of the receiving device. This modified list can be taken into account by the encoding device when choosing the input format.
At block 712 the encoding device receives the output format capability in the session negotiation answer and parses it.
If the session negotiation answer indicates a particular output format or type of output formats then the encoding device that receives the answer can select an alternative input format at block 714. The particular output format could be binaural output or any other suitable type of output format. The alternative audio input format can be a
different to the audio input format that was initially preferred or selected by the encoding device. The alternative input format can comprise more channels than the initially preferred or selected input format.
In some examples the alternative input format can be created by inserting a suitable number of empty channels into the initially preferred or selected audio input format.
In some examples the encoding options used for encoding the audio content can also be switched.
The alternative input format and/or the alternative encoding options can be selected based on any suitable criteria or combination of criteria. In some examples the alternative input format and/or the alternative encoding options can be selected to provide improved perceptual quality. In some examples the alternative input format can be selected to provide lower computational complexity.
For instance, if the output format of the receiving device is a binaural output then the encoding option can be switched to McMASA so as to improve the perceptual quality for the listener 210. If the output format of the receiving device is a loudspeaker presentation then the encoding option can be maintained as ParamMC, because it maintains individual channels more faithfully.
At block 716 the alternative input format and/or the alternative encoding option is used to encode the audio content. The use of the alternative input format and/or the alternative encoding option can result in perceptually better audio quality for rendering/presentation at a receiving device/playback device. In some examples the alternative input format and/or the alternative encoding option can result in lower computational complexity in encoding and/or decoding/rendering the audio content. In some examples the alternative input format and/or the alternative encoding option can result in a lower bitrate.
Table 2 shows an example Multichannel format that can be used to select alternative input formats and encoding options in some examples. In this example there are two possible multichannel formats to which the original multichannel format can be
promoted to allow McMASA encoding at a higher bitrate: 5.1.4 and 7.1.4. Otherwise, McMASA coding is not available above 32 kbps, which can compromise binaural output quality.
Table 2
Using this table in some examples, instead of transforming 5.1 and 5.1 .2 into 5.1.4 and 7.1 into 7.1.4, all other multichannel formats can be transformed into 7.1.4. In other examples, 5.1 , 5.1.2, and 7.1 can be transformed into 7.1.4, while 5.1.4 can maintained as they are.
Fig. 8 shows an example system 800 that can be used to implement examples of the disclosure. The example system 800 comprises an encoding device 802, a receiving device 804 and a network 806.
In this example the encoding device 802 is configured to capture spatial audio content and transmit it to a receiving device 804 via the network 806. The network can be a wireless communications network or any other suitable type of network. The receiving device 804 can be configured to receive the spatial audio content and render the spatial audio content so that it can be played back via a playback device. In this example the play back device is a head set 400 that a listener 210 is wearing over their
ears. Other types of playback devices, such as speaker devices, could be used in other examples.
The encoding device 802 comprises a microphone array 810, a processing module 812, a session controller module 814, an encoder/decoder module 818, a format adaptation module 816, and a transport stream module 820. In other examples the encoding device 802 can comprise different modules and/or combinations of modules.
The microphone array 810 can comprise multiple microphones. The microphones are configured to capture sound and produce electrical output signals. The microphones within the microphone array 810 can be spatially arranged so as to enable spatial audio to be captured.
The encoding device 802 is configured so that the output signals from the microphone array 810 are provided to a processing module 812. The processing module 812 is configured to process the microphone signals into a selected input format. The processing module 812 can be configured to process the microphone signals into the selected encoding option of the selected input format.
The encoding device 802 is configured so that the output signals from the processing module 812 is provided as an input to a format adaptation module 816. The format adaptation module can be configured to upmix the audio content to a different input format. The different input format can be selected based on information received from the receiving device 804. The information received from the receiving device 804 can comprise information indicating an output format used by the receiving device 804 or used by a playback device.
For example, the format adaptation module 816 can be configured to add additional channels to an input format. The additional channels can be empty or can comprise some content.
The output of the format adaptation module 816 can be provided to the encoder/decoder module 918. The encoder/decoder module 918 is configured to
encode the processed microphone signals for sending to the receiving device 804 via the network 806.
The encoded signals are provided to a transport stream module 820. The transport stream module 820 can be configured to generate a transport audio stream. The transport stream can use any suitable protocol such as Real-time Transport Protocol (RTP). The transport stream module 820 provides an encoded stream as a payload. The encoded stream can be transmitted to the receiving device 804 via the network 806.
The session controller module 814 can be configured to establish a communication session between the encoding device 802 and the receiving device 804. The session controller module 814 can be configured to transmit session negotiation signaling to the receiving device 804. The session controller module 814 can be configured to negotiate parameters for a multimedia session between the encoding device 802 and the receiving device 804.
The receiving device 804 also comprises a microphone array 810, a processing module 812, a session controller module 814, an encoder/decoder module 818, and a transport stream module 820. In other examples the receiving device 804 can comprise different modules and/or combinations of modules.
The transport stream module 820 of the receiving device can be configured to receive the encoded stream from the encoding device 802. The encoded stream can be provided to the encoder/decoder modules 818 for decoding. The decoded stream can be rendered and played back via the headphones 400 or via any other suitable playback device.
In this example the encoder/decoder module 818 can also receive head tracking information from the headphones 400. This information can provide an indication of the head orientation of the listener 210. This information can be used for the spatial rendering of the audio content.
The receiving device 804 also comprises a session controller module 814 that can be configured to establish a communication session between the encoding device 802 and the receiving device 804. The session controller module 814 can be configured to transmit responses to session negotiation signaling to the encoding device 802. The session controller module 814 can be configured to negotiate parameters for a multimedia session between the encoding device 802 and the receiving device 804.
The system 800 shown in Fig. 8 therefore enables the receiving device 804 to provide an indication of the output format used by a playback device to the encoding device 802. This information can be provided during session negotiation. This information can be used by the encoding device to adapt the input format and/or the encoding option used by the encoding device 802.
As an example the receiving device 804 indicates to the encoding device 802 during the session negotiation that binaural outputs are used. The encoding device 802 takes this into account and adapts the selected input format of the encoding device 802 to an alternative input format. The adaptation could be promoting a 7.1 audio input to a 7.1.4 input by inserting four empty height channels.
This adaptation of the input format can enable a different encoding option to be used. For example, it can enable McMASA encoding to be used at a significantly higher bitrate. In this example, in the affected bitrate range of 48-96 kbps, the perceptual quality of the binaurally rendered audio output is improved and the computational complexity is reduced. On the other hand, if the output format were a loudspeaker setup, such as a 7.1 or 7.1.4 or any other suitable setup, it might be preferable to use ParamMC as the encoding option so as to maintain as close match per channel between the original input and the decoded/rendered audio as possible at these bit rates.
The output format used by the playback device can be indicated by an output format parameter (outf) in a session negotiation communication. outf: Indicates the output format capability. If multiple output formats are supported in a continuous range, the range can be indicated by the first output format in the range
and the last in the range separated by a hyphen (outf1-outf2). If the multiple output formats that are supported are not in a contiguous range then these can be listed as comma separated values (out1 , out2). Comma separated values can also be used, when the output formats are within a range, but the preferred order of the formats is not the default contiguous range. In both cases, hyphen or comma separated list, the output formats can be listed in a preferred order from the most preferred to the least preferred output format, outf-send and outf-recv can be used in case of different output formats are used in both the send and receive directions respectively. If outf parameter is not present, all possible output formats are supported.
Table 3 shows an example association between parameter values and output formats.
Table 3
Fig. 9 schematically illustrates an apparatus 901 that can be used to implement examples of the disclosure. In this example the apparatus 901 comprises a controller 903. The controller 903 can be a chip or a chipset. The apparatus 901 can be provided within an encoding device 802 or any other suitable type of device.
In the example of Fig. 9 the implementation of the controller 903 can be as controller circuitry. In some examples the controller 903 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware).
As illustrated in Fig. 9 the controller 903 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 909 in a general-purpose or special-purpose processor 905 that may be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 905.
The processor 905 is configured to read from and write to the memory 907. The processor 905 can also comprise an output interface via which data and/or commands are output by the processor 905 and an input interface via which data and/or commands are input to the processor 905.
The memory 907 stores a computer program 909 comprising computer program instructions (computer program code 911) that controls the operation of the controller 903 when loaded into the processor 905. The computer program instructions, of the computer program 909, provide the logic and routines that enables the controller 903. to perform the methods illustrated in the accompanying Figs. The processor 905 by reading the memory 907 is able to load and execute the computer program 909.
The apparatus 901 comprises: at least one processor 905; and at least one memory 907 storing instructions that, when executed by the at least one processor 905, cause the apparatus 901 at least to perform: obtaining 300 a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving 302 an indication of an output format configured to be used by a playback device; and upmixing 304 the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
As illustrated in Fig. 9, the computer program 909 can arrive at the controller 903 via any suitable delivery mechanism 913. The delivery mechanism 913 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory
computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid-state memory, an article of manufacture that comprises or tangibly embodies the computer program 909. The delivery mechanism can be a signal configured to reliably transfer the computer program 909. The controller 903 can propagate or transmit the computer program 909 as a computer data signal. In some examples the computer program 909 can be transmitted to the controller 903 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.
The computer program 909 comprises computer program instructions for causing an apparatus 901 to perform at least the following or for performing at least the following: obtaining 300 a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving 302 an indication of an output format configured to be used by a playback device; and upmixing 304 the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
The computer program instructions can be comprised in a computer program 909, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 909.
Although the memory 907 is illustrated as a single component/circuitry it can be implemented as one or more separate components/circuitry some or all of which can be integrated/removable and/or can provide permanent/semi-permanent/ dynamic/cached storage.
Although the processor 905 is illustrated as a single component/circuitry it can be implemented as one or more separate components/circuitry some or all of which can be integrated/removable. The processor 905 can be a single core or multi-core processor.
References to ‘computer-readable storage medium’, ‘computer program product’, ‘tangibly embodied computer program’ etc. or a ‘controller’, ‘computer’, ‘processor’ etc. should be understood to encompass not only computers having different architectures such as single /multi- processor architectures and sequential (Von Neumann)/parallel architectures but also specialized circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc.
As used in this application, the term ‘circuitry’ may refer to one or more or all of the following:
(a) hardware-only circuitry implementations (such as implementations in only analog and/or digital circuitry) and
(b) combinations of hardware circuits and software, such as (as applicable):
(i) a combination of analog and/or digital hardware circuit(s) with software/firmware and
(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory or memories that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and
(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (for example, firmware) for operation, but the software may not be present when it is not needed for operation.
This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their)
accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
The blocks illustrated in Figs. 3 and 7 can represent steps in a method and/or sections of code in the computer program 909. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted.
The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to “comprising only one...” or by using “consisting”.
In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected/coupled/in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., so as to provide direct or indirect connection/coupling/communication. Any such intervening components can include hardware and/or software components.
As used herein, the term "determine/determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, " determine/determining" can include resolving, selecting, choosing, establishing, and the like.
In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions
are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’ or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all of the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example.
Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
Features described in the preceding description may be used in combinations other than the combinations explicitly described above.
Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not.
The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a/an/the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning.
The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and also to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result.
In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described.
The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure.
Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance it should be understood that the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and/or shown in the drawings whether or not emphasis has been placed thereon. l/we claim:
Claims
1 . An apparatus comprising means for: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
2. An apparatus as claimed in claim 1 , wherein the means are for at least one of: switching the encoding option to the alternative encoding option and wherein the switching is based, at least in part, on the output format used by a playback device and the alternative input format; and encoding the spatial audio content using the alternative encoding option.
3. An apparatus as claimed in any preceding claim, wherein the alternative encoding option enables a higher bit rate to be supported.
4. An apparatus as claimed in any preceding claim, wherein the means are for at least one of: receiving a list of supported modes for encoding the spatial audio content; sorting the list into a preferred order; and adding one or more alternative modes to the list where the alternative modes comprise multi-channel formats that support a higher bit rate.
5. An apparatus as claimed in claim 4, wherein a mode for encoding the spatial audio content comprises an input/output format and an encoding option.
6. An apparatus as claimed in any preceding claim, wherein the indication of the output format used by a playback device comprises a value of a parameter.
7. An apparatus as claimed in any preceding claim, wherein the means are for mixing the spatial audio content such that additional channels of the alternative input format do not contain any content.
8. An apparatus as claimed in any of claims 1 to 6, wherein the means are for mixing the spatial audio content such that additional channels of the alternative input format do contain some content.
9. An apparatus as claimed in any of claims 1 to 6, wherein the means are for mixing the spatial audio content such that no change is made to content of channels of the selected input format when upmixing to the alternative input format.
10. An apparatus as claimed in any of claims 1 to 7, wherein the means are for mixing the spatial audio content such that at least some changes are made to content of the channels of the selected input format when upmixing to the alternative input format.
11. An apparatus as claimed in any preceding claim, wherein the alternative input format comprises a format that provides at least two additional channels.
12. An apparatus as claimed in any preceding claim, wherein the output format used by a playback device is a binaural format and the updated encoding option comprises multichannel coding using metadata-assisted spatial audio.
13. An apparatus as claimed in any preceding claim, wherein the means are for enabling switching to the alternative encoding option when it is determined that the alternative encoding option supports a higher bit rate.
14. An apparatus as claimed in any preceding claim, wherein the means are for determining the alternative input format and the encoding option to be used for spatial audio encoding through negotiation with a playback device.
15. An apparatus as claimed in any of claims 1 to 13, wherein the means are for determining the alternative input format and the encoding option to be used for spatial audio encoding based on the indication of the output format used by a playback device.
16. A method comprising: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
17. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform: obtaining a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receiving an indication of an output format configured to be used by a playback device; and upmixing the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
18. An apparatus comprises: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain a selected input format and an encoding option for encoding spatial audio content wherein the encoding option is configured to be switched to an alternative encoding option for encoding the spatial audio content; receive an indication of an output format configured to be used by a playback device; and
upmix the selected input format to an alternative input format wherein the alternative input format comprises more channels than the selected input format, and the alternative input format is selected based, at least in part, on the output format.
19. An apparatus as claimed in claim 18, wherein the apparatus is caused to: switch the encoding option to the alternative encoding option, wherein the switching is based, at least in part, on the output format used by the playback device and the alternative input format; and encoding spatial audio content using the alternative encoding option.
20. An apparatus as claimed in any of claim 18 or 19, wherein the alternative encoding option enables a higher bit rate to be supported.
21. An apparatus as claimed in any of claims 18 to 20, wherein the apparatus is caused to at least one of: receive a list of supported modes for encoding spatial audio content; sort the list into a preferred order; and add one or more alternative modes to the list where the alternative modes comprise multi-channel formats that support a higher bit rate.
22. An apparatus as claimed in any of claims 18 to 21 , wherein the indication of thew output format used by the playback device comprises a value of a parameter.
23. An apparatus as claimed in any of claims 18 to 22, wherein the output format used by the playback device is a binaural format and the updated encoding option comprises multichannel coding using metadata-assisted spatial audio.
24. An apparatus as claimed in any of claims 18 to 23, wherein the apparatus is caused to at least one of: determine the alternative input format and the encoding option to be used for spatial audio encoding through negotiation with a playback device; determine the alternative input format and the encoding option to be used for spatial audio encoding based on the indication of the output format used by a playback device; and
switch to the alternative encoding option when it is determined that the alternative encoding option supports a higher bit rate
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2310063.9A GB2631478A (en) | 2023-06-30 | 2023-06-30 | Apparatus, methods and computer program for encoding spatial audio content |
| PCT/EP2024/065999 WO2025002779A1 (en) | 2023-06-30 | 2024-06-11 | Apparatus, methods and computer program for encoding spatial audio content |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4736159A1 true EP4736159A1 (en) | 2026-05-06 |
Family
ID=87556946
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24732620.0A Pending EP4736159A1 (en) | 2023-06-30 | 2024-06-11 | Apparatus, methods and computer program for encoding spatial audio content |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4736159A1 (en) |
| CN (1) | CN121464479A (en) |
| GB (1) | GB2631478A (en) |
| WO (1) | WO2025002779A1 (en) |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7212872B1 (en) * | 2000-05-10 | 2007-05-01 | Dts, Inc. | Discrete multichannel audio with a backward compatible mix |
| TWI618051B (en) * | 2013-02-14 | 2018-03-11 | 杜比實驗室特許公司 | Audio signal processing method and apparatus for audio signal enhancement using estimated spatial parameters |
| IL319278A (en) * | 2018-07-02 | 2025-04-01 | Dolby Laboratories Licensing Corp | Methods and devices for generating or decoding a bitstream comprising immersive audio signals |
| SG11202007627RA (en) * | 2018-10-08 | 2020-09-29 | Dolby Laboratories Licensing Corp | Transforming audio signals captured in different formats into a reduced number of formats for simplifying encoding and decoding operations |
| JP7316384B2 (en) * | 2020-01-09 | 2023-07-27 | パナソニック インテレクチュアル プロパティ コーポレーション オブ アメリカ | Encoding device, decoding device, encoding method and decoding method |
-
2023
- 2023-06-30 GB GB2310063.9A patent/GB2631478A/en active Pending
-
2024
- 2024-06-11 WO PCT/EP2024/065999 patent/WO2025002779A1/en not_active Ceased
- 2024-06-11 CN CN202480043825.3A patent/CN121464479A/en active Pending
- 2024-06-11 EP EP24732620.0A patent/EP4736159A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025002779A1 (en) | 2025-01-02 |
| GB202310063D0 (en) | 2023-08-16 |
| CN121464479A (en) | 2026-02-03 |
| GB2631478A (en) | 2025-01-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20230179941A1 (en) | Audio Signal Rendering Method and Apparatus | |
| WO2019229299A1 (en) | Spatial audio parameter merging | |
| US12167220B2 (en) | Audio representation and associated rendering | |
| US20240312469A1 (en) | Apparatus, Methods and Computer Programs for Encoding Spatial Metadata | |
| JP7639053B2 (en) | Processing of mono signals in a 3D audio decoder delivering binaural content | |
| WO2020152394A1 (en) | Audio representation and associated rendering | |
| US20250330760A1 (en) | Methods and systems for immersive 3dof/6dof audio rendering | |
| US11950080B2 (en) | Method and device for processing audio signal, using metadata | |
| US12089028B2 (en) | Presentation of premixed content in 6 degree of freedom scenes | |
| EP4736159A1 (en) | Apparatus, methods and computer program for encoding spatial audio content | |
| US20230188924A1 (en) | Spatial Audio Object Positional Distribution within Spatial Audio Communication Systems | |
| US20250315207A1 (en) | Power Saving for Audio Streams | |
| WO2025119539A1 (en) | Apparatus and method for encoding a combined input format spatial audio signal | |
| WO2025078226A1 (en) | Parametric spatial audio decoding with pass-through mode | |
| WO2025223950A1 (en) | Signalling of pass-through mode in spatial audio coding | |
| WO2024245695A1 (en) | Apparatus, methods and computer program for selecting a mode for an input format of an audio stream | |
| GB2640555A (en) | Immersive communication sessions | |
| CN119559954A (en) | Spatial Audio | |
| GB2633673A (en) | Method and system for coding audio data |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |