EP4674142A1 - Split binaural rendering - Google Patents
Split binaural renderingInfo
- Publication number
- EP4674142A1 EP4674142A1 EP24715359.6A EP24715359A EP4674142A1 EP 4674142 A1 EP4674142 A1 EP 4674142A1 EP 24715359 A EP24715359 A EP 24715359A EP 4674142 A1 EP4674142 A1 EP 4674142A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- pose
- binaural
- probing
- metadata
- axis
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
- H04S7/304—For headphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/01—Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]
Definitions
- Immersive audio is an essential media component of extended reality (XR) applications, which includes augmented reality (AR), mixed reality (MR) and virtual reality (VR).
- XR extended reality
- AR augmented reality
- MR mixed reality
- VR virtual reality
- immersive audio may support adjusting the presented immersive audio/visual scene in response to motion of the user. For example, it may be desirable to track a user’s head position and head movement during audio rendering and to adjust the audio accordingly.
- an immersive audio experience may process head movements using models with three degrees of freedom (3DoF) or six degrees of freedom (6DoF).
- 3DoF three degrees of freedom
- 6DoF six degrees of freedom
- Various immersive audio services e.g., immersive voice and audio services (IVAS), may be used to render high quality audio renditions at the XR device that include awareness of pose information, which may include metadata for head positions with relative or absolute movements of the user.
- IVAS immersive voice and audio services
- One potential solution is to reduce audio rendering requirements at the end-device (e.g., the AR device operated by the user) with a split-rendering topology that leverages processing from some other entity of the mobile/wireless network (e.g., a network based device) to which the end-device is connected or tethered (e.g., via a network or cloud-based connection).
- some other entity of the mobile/wireless network e.g., a network based device
- a powerful network entity such as mobile user equipment (e.g., UE, a device used by an end-user, a portable multi-function device, a gaming console, a cloud-based resource, etc.) may be connected to the end-device to assist in split-rendering of immersive audio.
- Pose information based on the user movement may be gathered at the end-device and transmitted to the network entity.
- the end-device may then receive the already rendered audio from the network entity; where the high complexity calculations such as processing 3DoF/6DoF pose information (e.g., head-tracking metadata) may be performed by the rendering entity (e.g., network entity).
- 3DoF/6DoF pose information e.g., head-tracking metadata
- the rendering entity e.g., network entity.
- One problem with the described split-rendering topology is the latency for transmissions between end-device and network entity may be on the order of 100ms; which means the network entity may be relying on outdated pose/head-tracking information. Because of this delay, the rendered audio from the network entity may not match the current head pose/head position of the user at the end-device.
- U.S. Prov. Appl. No. US 63/340,181 discloses a novel approach to interactive headtracking.
- the described approach generates multiple binaural representations (pre- renditions) corresponding to various head poses at the main device or pre-renderer and computes metadata which can be used along with a reference binaural representation to reconstruct binaural output corresponding to any given pose at the post-renderer.
- the reference binaural representation and the metadata are sent to a post-rendering device.
- the post-renderer determines binaural audio corresponding to the current head pose.
- U.S. Prov. Appl. No. US 63/386,465 describes a system relying on a binaural rendering for a reference pose ⁇ ′ obtained upstream from the post-renderer device, and a number of ⁇ pre-renditions for ‘probing’ poses ⁇ ⁇ which are close to reference pose ⁇ ′.
- complexity constraints at the pre-renderer device and metadata bit rate limitations on the transmission interface between the pre- and post-renderer devices may limit the number of pre-renditions (or binaural representations) to be computed at the pre-renderer device and also may limit the amount of metadata to be transmitted to the post renderer for pose correction.
- One cause of numerical complexity at the pre-renderer device is the required number of pre-renditions.
- the present disclosure describes a low-complexity split rendering technique based on a limited number of pre-renditions, thereby significantly reducing the required number of computations on the pre-rendering side, as well as reducing the amount of transmitted metadata.
- the present disclosure further describes low complexity solutions for post-renderer corrections around one, two or three rotation axes, e.g., for deviations of yaw, pitch and roll.
- this and other objects are achieved by a method of rendering audio (in a main device) to enable split rendering techniques with pose correction around multiple rotational axes (in a lightweight device), the method comprising obtaining an immersive audio content, obtaining a reference pose, rendering the immersive audio content into a first number of binaural pre-renditions, wherein the binaural pre-renditions correspond to a set of probing poses, wherein the set of probing poses include poses equal to the reference pose (P’) and/or poses that deviate from the reference pose by rotation around at least one of the rotational axes, calculating a second number of approximate binaural representations based on the binaural pre-renditions, wherein the approximate binaural representations correspond to a set of virtual probing poses, wherein the virtual probing poses deviate from the probing poses by rotation around at least one of the rotational axes, determining a reference binaural representation based on one or more of the binaural pre-ren
- this and other objects are achieved by a method of rendering audio (in a main device) to enable split rendering with pose correction around yaw axis and pitch axis (in a lightweight device), the method comprising obtaining an immersive audio content, obtaining a reference pose, rendering the immersive audio content into a reference binaural representation corresponding to a reference pose, rendering the immersive audio content into one or more binaural pre-renditions, wherein the binaural pre-renditions correspond to one or more probing poses deviating from the reference pose by rotation deviate from the reference pose by rotation around both yaw and pitch axes, computing, for each probing pose, yaw metadata representing a deviation around the yaw axis, and pitch metadata representing deviation around the pitch axis, encoding the reference binaural representation, the yaw metadata and the pitch metadata in an output bitstream, and outputting the output bitstream [017]
- this and other objects are achieved by obtaining an immersive audio content, obtaining a reference pose, rendering the
- this and other objects are achieved by a method of audio processing with pose correction around multiple rotational axes, the method comprising receiving a bitstream from a main device, decoding the bitstream to obtain a reference binaural representation and first reconstruction metadata associated with a set of probing poses representing deviation from a reference pose by rotation around the multiple rotational axes, detecting a current head-pose, for each of the rotational axes, selecting a probing pose closest to the detected pose along the rotation axis, determining axis-specific reconstruction metadata based on the first metadata associated with the selected probing pose and a difference between the reference pose and the current head pose along the rotational axis, and determining a binaural output corresponding to the current head pose based on the reference binaural representation and the axis-specific reconstruction metadata for each rotational axis.
- a main processing device comprising a decoder configured to decode a first bitstream to obtain decoded immersive audio content, a renderer configured to obtain a reference pose, render the immersive audio content into a reference binaural representation based on the reference pose, and render the immersive audio content into a number of binaural pre-renditions, wherein the binaural pre- renditions correspond to a set of probing poses associated with a reference pose, wherein the set of probing poses include poses that deviate from the reference pose by rotation about at least one of the rotational axes, a metadata generator configured to compute reconstruction metadata to enable reconstruction of the binaural pre-renditions from the reference binaural representation, an encoder configured to encode the reference binaural representation and the reconstruction metadata into an output bitstream, and an interface configured to output the output bitstream.
- a lightweight processing device comprising a decoder configured to decode a bitstream to obtain a reference binaural representation and first reconstruction metadata associated with a set of probing poses representing deviation from a reference pose by rotation around multiple rotational axes, a head- tracker configured to detect a current head-pose, a binaural reconstruction block configured to, for each of the rotational axes, select a probing pose closest to the detected pose along the rotation axis, and determine axis-specific reconstruction metadata based on the first metadata associated with the selected probing pose and a difference between the reference pose and the current head pose along the rotational axis, and determine a binaural output corresponding to the current head pose based on the reference binaural representation, and the axis-specific reconstruction metadata.
- FIG.1 shows a user with smartphone and a set of headphones.
- FIG.2 is a diagram illustrating a head pose and rotation around the three axis yaw, pitch and roll.
- FIG.3 is a schematic block diagram showing split rendering in a main processing device and a lightweight processing device.
- FIG.4 is a flow chart illustrating processing in a main processing device, in accordance with embodiments of a first aspect of the invention.
- FIG.5 is a flow chart illustrating processing in a main processing device, in accordance with embodiments of a second aspect of the invention.
- FIG.6 is a flow chart illustrating processing in a lightweight processing device, in accordance with embodiments of a further aspect of the invention.
- FIG.7 is a flow chart illustrating processing in a main processing device, in accordance with embodiments of a yet further aspect of the invention.
- FIG.8 is a flow chart illustrating processing in a main processing device, in accordance with embodiments of a still further aspect of the invention.
- FIG.9 illustrates a schematic block diagram of an example device or architecture that may be used to implement embodiments of the invention.
- FIG.1 shows schematically a user 1 having a smartphone 2 and wearing a headset 3.
- the smartphone could in the context of the present invention serve as the main, pre-rendering device, while the headset could serve as the lightweight, user-held, post-rendering device. It is the post rendering device that has the most recent information about the users head pose P.
- a user head pose P is in the present context defined by three degrees of freedom, namely rotation ⁇ around the yaw axis 101, rotation ⁇ around the pitch axis 102, and rotation ⁇ around the roll axis 103.
- the head pose will also be associated with a position in the room, defined by three additional degrees of freedom, spatial coordinates x, y, z.
- spatial translation will not be relevant for the purposes of the present disclosure.
- FIG.3 shows an example of some of the functional blocks that may be implemented in the main device 2 and lightweight device 3 in an example split rendering system.
- the main device, or pre-rendering device, 2 includes a decoder 11, a binaural renderer 12, a metadata generator 13, a first encoder 14, a second encoder 15, and a multiplexer 16.
- the main device 2 may also include a pose decoder 17.
- the decoder 11, e.g., an IVAS decoder is configured to receive and decode a bitstream b1, and decode an immersive audio content A.
- the binaural renderer 12 is configured to receive (or obtain) the immersive audio content A and a reference pose P’, and responsively provide a reference binaural representation (rendition) Bin ref , associated with the reference pose P’.
- the reference pose may be an assumed head pose, or be determined based on head pose information received from the lightweight device 3 via pose decoder 17.
- the binaural renderer 12 is further configured to responsively render N (N>0) binaural representations (pre-renditions) Bin n corresponding to a set of probing poses P n associated with the reference pose P’.
- the metadata generator 13 is configured to receive the reference binaural representation Bin ref and pre-renditions Bin n , and responsively generate reconstruction metadata M to enable reconstruction of the pre-renditions from the reference binaural representation.
- the reference pose P’ may be received by metadata generator 13, and responsively encoded in the reconstruction metadata M.
- Pose decoder 17 is an optional block that is not required for all implementations. When pose decoder 17 is present, pose decoder 17 is configured to receive head pose information from the lightweight device 3 via bitstream bp, and responsively generate the reference pose P’.
- the first encoder 14 is configured to receive the reference binaural representation Bin ref , and responsively encode the reference binaural representation Bin ref as encoded bitstream b11.
- the second encoder 15 is configured to receive reconstruction metadata M, and responsively encode the reconstruction metadata (and optionally pose information) as encoded bitstream b 12 .
- the multiplexer 16 is configured to receive the encoded bitstreams b11 and b12 from the outputs of the two encoders 14 and 15, and responsively combine the encoded bitstreams b 11 and b 12 into a bitstream b2.
- the main device may also include an interface to output the bitstream b2, whereby the bitstream may be subsequently transmitted or otherwise made available to another device that is external to the main device 2, here the lightweight device 3.
- the encoder 15 is further configured to encode pose information into bitstream b 12 , where the encoded pose information is indicative of the reference pose P’ and/or the probing poses Pn.
- the lightweight device, or post-renderer device, 3 here includes a demultiplexer 21, a first decoder 22, a second decoder 23, a binaural reconstruction block 24 and a head-tracker 25.
- the lightweight device 3 also includes a pose information encoder 26.
- the demultiplexer 21 is configured to receive bitstream b2 from the main device 2 and responsively separate the received bitstream b2 into two encoded bitstreams b21 and b22.
- the decoder 22 is configured to receive encoded bitstream b21, and responsively decode bitstream b21 into a reference binaural signal Bin ref .
- the decoder 23 is configured to receive encoded bitstream b22, and responsively decode bitstream b22 into metadata M’ (and, if present, information about the reference pose P’ and/or the probing poses P n ).
- the binaural reconstruction block 24 is configured to receive a current user head P detected by the head tracker 25, and responsively determine a binaural output on the reference binaural signal Bin ref , and the metadata M’, and the current head pose P in relation to the reference pose P’.
- the reference pose P’ and/or the probing poses P n may be included in the bitstream b22 received from the main device 2. This is especially useful when the reference pose is based on pose information received by the main device 2 from the lightweight device 3.
- the reference pose P’ is an assumed pose and thus the lightweight device is already aware of pose information P’.
- the reference pose may be a “straight ahead” pose, e.g., a pose looking straight at a display device.
- information about the probing poses may be received in the bitstream, but may alternatively be predefined and known by the lightweight device.
- the probing poses may be pre-defined deviations from the reference pose.
- the encoder 26 is an optional block that is not required in all implementations. If encoder 26 is present, encoder 26 is configured to receive pose information P from the head- tracker 25, and responsively encode the pose information in a bitstream b P , which is sent to the main device 2. [048] In an example implementation, a heavy weight device 2 uses a pose ⁇ ⁇ to generate a reference binaural signal '() ⁇ and metadata (MD) such that the light weight post renderer can do the pose correction from ⁇ ⁇ to the actual pose ⁇ and generate '() ⁇ from '() ⁇ using metadata M, wherein '() ⁇ has all the spatial cues as per Pose ⁇ .
- MD metadata
- pose ⁇ ⁇ at the pre-renderer is an assumed pose without any information from light weight device. In some other implementations, pose ⁇ ⁇ at the pre-renderer is received from light weight device through a back channel.
- Computing metadata corresponding to various probing poses such that the lightweight post renderer can do the pose correction from ⁇ ⁇ to the actual pose ⁇ can require multiple binaural renditions at the pre-renderer, also referred to as pre-renditions in this document, and these pre-renditions may be complexity intensive.
- computing metadata corresponding to multiple probing poses can increase the metadata bitrate significantly. Hence, it is desired to carefully select the probing pose points for pre-renditions such that the total number of pre-renditions can be limited.
- ⁇ and ⁇ ′ may be (yaw) angles * and ⁇ *.
- the deviation from the reference pose is not necessarily equal ( ⁇ *) but is assumed here for simplifying the description.
- One potential way to save pre-renderer complexity is to skip one rendition. This is a workable solution as long as the trend in time of the yaw angle is known. In that case, the probing rendition may be done for just either +* or ⁇ * towards which the pose is expected to evolve. However, in general, such a trend may be unknown. For example, the trend is unknown in cases where the current pose is static, since it is unknown whether the user will next turn the head to the right or the left.
- a first solution to reducing the number of renditions to two is to give up pre- rendering to the reference pose ⁇ ′. Instead, pre-renditions are generated for poses ⁇ ′ ⁇ * and ⁇ ⁇ + * and one of these binaural renditions, e.g., for pose ⁇ ′ ⁇ * is transmitted to the post- renderer device. In case, pose ⁇ at the post-renderer is static and thus identical to ⁇ ′, the post renderer will thus have to do a correction by +*.
- a further potential disadvantage is the bias of the solution with potentially less accurate post-renderer output signal for pose ⁇ ⁇ + * compared to the (perfect) post-renderer output signal for pose ⁇ ⁇ ⁇ *. This bias may be overcome by transmitting one of the11inaurall channels (e.g. left channel) of the rendition for pose ⁇ ⁇ ⁇ * and one of the binaural channels (e.g. right channel) of the rendition for pose ⁇ ⁇ + *.
- the post rendering for pose ⁇ will consequently involve adjusting the left channel using renderer metadata relative to the rendition of that channel for pose ⁇ ⁇ ⁇ * and adjusting the right channel using renderer metadata relative to the rendition for pose ⁇ ⁇ + *.
- Another solution for that problem is in the pre-renderer to firstly generate an approximation of the binaural rendition for the reference pose ⁇ ⁇ through interpolation between the available renditions for poses ⁇ ′ ⁇ * and ⁇ ⁇ + *. This may simply involve averaging the two available binaural renditions to generate a reference binaural rendition for virtual reference pose ⁇ ′.
- split renderer metadata can be calculated as described in U.S.63/340,181 based on the binaural renditions for ⁇ ⁇ , ⁇ ′ ⁇ * and ⁇ ⁇ + *, whereby it is notable that the fact that the reference rendition is obtained through interpolation creates metadata symmetries which alleviates the need to calculate and transmit metadata associated with one of the probing positions.
- the reference rendition along with the metadata are transmitted to the post renderer device where operations can take place as described in U.S.63/340,181, hereby incorporated by reference.
- split renderer metadata for post-renderer corrections for pose deviations around 2 and 3 axes, e.g., for corrections of yaw and pitch deviations or for corrections of yaw, pitch and roll deviations.
- 2-AXES SPLIT RENDERING METADATA CALCULATION Another example case is described where split rendered metadata is calculated with two-axis (e.g., yaw and pitch) correction. For this example, four probing poses may be considered with a reference pose, where techniques suggested by U.S.63/340,181 may be carried out.
- Two probing poses may be employed to probe first axis (e.g., yaw axis) deviations from the reference pose, e.g., by varying the pose relative to the reference pose by deviations of ⁇ * about the first axis while keeping a second axis (e.g., pitch) unchanged.
- Two other probing poses may be employed to probe pitch deviations from the reference pose, e.g., by varying the pose relative to the reference pose by pitch deviations of ⁇ , while keeping the yaw unchanged.
- the total number of renditions is five for this example.
- FIG.4 is a flow chart illustrating processing in the main device 2 in accordance with embodiments of a first aspect of the invention, relating to a method of rendering audio in the main device 2 to enable split rendering with pose correction around multiple rotational axes.
- the flow chart may be broken into various blocks or partitions, such as blocks S11 – S17.
- step S11 (obtain audio content)
- a first bitstream is received and decoded (e.g., by decoder 11) to obtain an immersive audio content A.
- step S11 may be followed by step S12.
- step S12 (obtain reference pose)
- a reference pose P’ is obtained.
- the reference pose may be an assumed pose (e.g. straight ahead) or may be based on pose information received from the lightweight device 3.
- step S12 may be followed by step S13.
- step S13 pre-rendering
- a first number of binaural pre-renditions are rendered (e.g., by renderer 12), wherein the binaural pre-renditions correspond to a set of probing poses Pn, including poses deviating from the reference pose P' by rotation around at least one of the rotational axes.
- the set of probing poses also includes the reference pose.
- Step S13 may be followed by step S14.
- step S14 calculate approximate representations
- a second number of approximate binaural representations Bin'm are calculated based on the binaural pre-renditions Bin n , wherein the approximate binaural representations Bin' m correspond to a set of virtual probing poses Pm, each virtual probing pose (Pm) deviating from the probing poses Pn by rotation around at least one of the rotational axes.
- Step S14 may be followed by step S15.
- step S15 determine Binref
- a reference binaural representation, Binref is determined.
- the reference binaural representation Binref may be equal to one of the binaural pre-renditions Binn or one of the approximate binaural representations Bin'm.
- the reference binaural representation Binref may correspond to the reference pose P' and may then be a pre-rendition corresponding to the reference pose.
- a reference binaural representation Binref corresponding to the reference pose P' may also be obtained by linearly combining several pre-renditions.
- Step S15 may be followed by step S16. [063]
- step S16 (generate M), reconstruction metadata M, which enables reconstruction of the binaural pre-renditions Bin n and the approximate binaural representations Bin' m from the reference binaural representation Binref, is computed.
- Steps S14 – S16 may all be performed by metadata generator 13 in figure 3. If the reference binaural representation is rendered, such rendering may be performed by renderer 12, and the reference binaural representation will be one of the binaural pre-renditions. Step S16 may be followed by step S17. [065] In step S17 (encode and output bitstream), the reference binaural representation Binref and the reconstruction metadata M are encoded (e.g., by encoders 14, 15 or a single encoder) in an output bitstream (b 2 ), which is subsequently outputted on an appropriate communication channel. The reconstruction metadata may be quantized and encoded based on symmetries in reconstruction metadata.
- the reconstruction metadata may be encoded using differential coding between metadata relating to different (symmetrical) probing poses.
- the step of computing reconstruction metadata may include computing axis-specific metadata for each rotational axis.
- the method comprises, for each rotational axis, selecting a first set of representations from the binaural pre-renditions Binn and the approximate binaural representations Bin' m , this first set of representations corresponding to probing poses deviating from each other by rotation around the axis (e.g.
- the pre-renderer may render binaural presentations for three probing poses ⁇ ⁇ , ⁇ # and ⁇ -. 1.
- * denotes a probing angle for deviations around the yaw axis, herein referred to as yaw deviations
- pitch deviations a corresponding probing angle for deviations around the pitch axis.
- pitch correction metadata H can be calculated based on pre-renditions Bin1 and Bin 2 and approximate rendition Bin' 1 , using a technique described in U.S.63/340,181 and with Bin'1 as reference representation.
- yaw correction metadata M can be calculated based the pre-rendition for probing pose ⁇ - and the approximate rendition for pose ⁇ . , using a technique described in U.S.63/340,181 and again with Bin'1 as reference representation.
- approximate rendition Bin'1 was used as reference for the calculation of both yaw and pitch metadata, and it will be appropriate to encode and transmit this representation.
- the binaural reference rendition to be transmitted to the post- renderer may be any of those available for the probing poses ⁇ ⁇ , ⁇ # and ⁇ - and the virtual probing pose ⁇ . .
- This reference rendition may be obtained through low-complex post-renderer operations based on any (or a combination) of the available pre-renderings for the probing positions and using a technique described in U.S.63/340,181. It is also possible to obtain the reference rendition directly based on linear or triangular interpolation. Let 0 ⁇ , 0 # and 0- denote the binaural pre-renditions for probing poses ⁇ ⁇ , ⁇ # and ⁇ -, an interpolated reference rendition 0 1 2 for virtual pose ⁇ ′ can be obtained by the following weighted averaging: 1 1 .
- the pre-renderer may render binaural presentations for only four probing poses ⁇ ⁇ , ⁇ # , ⁇ - and ⁇ . . 1.
- * denotes a probing angle for yaw deviations, , a corresponding probing angle for pitch deviations and 5 a probing angle for roll deviations.
- an approximate rendition Bin' 1 for a virtual probing pose ⁇ 6 ⁇ ⁇ + ( ⁇ *, ⁇ ,, 0 is calculated, for instance by interpolating the renditions for ⁇ ⁇ and ⁇ # . 3.
- roll correction metadata can be calculated based on pre-renditions Bin1 and Bin 2 and approximate rendition Bin' 1 , using a technique described in U.S. 63/340,181, and using Bin'1 for pose ⁇ 6 as reference presentation. 4.
- pitch correction metadata can be calculated based the pre-rendition Bin3 for pose ⁇ - and the approximate rendition Bin'1 for virtual probing pose ⁇ 6 using a technique described in U.S.63/340,181 and again using Bin'1 for pose ⁇ 6 as reference presentation. 5.
- yaw correction metadata can be calculated based on the pre-rendition Bin4 for probing pose ⁇ . and the approximate rendition Bin'2 for pose ⁇ 7 , using a technique described in U.S.63/340,181 and using either one of the renditions as reference presentation. [076] As described above, it is possible to do the operation steps to obtain yaw, pitch and roll correction metadata in different orders. The order may also be adapted based on properties of the immersive audio signal. Some of the steps and correction metadata calculations may even be omitted based on such immersive audio signal properties.
- pitch pose correction can be approximated with a table that contains gain parameters corresponding to various pitch angles.
- a table can be computed once during initialization time and both pre-renderer and post renderer can have prior knowledge about these tables.
- pitch correction metadata is not necessary in the bitstream, and the above steps could be adapted to calculate yaw and roll correction metadata only.
- roll correction metadata is not necessary in the bitstream, and roll pose correction can be approximated with a table that contains gain parameters corresponding to various roll angles.
- the binaural reference rendition (representation) to be transmitted to the post-renderer may be any of the pre-renditions for the exercised probing poses or any of the approximated pre-renditions at the virtual probing poses.
- any other approximated binaural pre-rendition can be used as reference rendition based on the available pre-renditions.
- a reference rendition may be obtained through low-complex post-renderer operations based on any (or a combination) of the available pre-renderings for the probing positions and using a technique described in U.S.63/340,181.
- An approximation of the pre- rendition for the reference pose can be obtained through interpolation between the available pre- renditions.
- the approximated rendition for reference pose P’ can be obtained from the available pre-renditions for probing poses ⁇ ⁇ through ⁇ . .
- 0 ⁇ through 0 . denote the binaural pre-renditions for probing poses ⁇ ⁇ through ⁇ .
- ITERATIVE SPLIT RENDERER METADATA ENHANCEMENT [080]
- the above examples of complexity-reduced split renderer metadata calculation for post-renderer corrections have a certain bias. For instance, in the 3-axes case, roll correction metadata is calculated for yaw and pitch angle deviations from the reference pose ⁇ ⁇ of ⁇ * and ⁇ ,. This makes the obtained roll correction metadata less precise for the more likely case that the yaw and pitch angles of the pose corresponds to those of the reference pose ⁇ ⁇ . It would thus be more correct to calculate the roll correction metadata for yaw and pitch angle deviations equal to 0.
- the pitch correction metadata is biased since it is calculated for a yaw deviation angle of ⁇ * rather than 0.
- the bias in the metadata calculations may in turn cause inaccuracies in the renditions obtained by the post-renderer using that biased metadata.
- an iterative enhancement technique is described that can mitigate the described bias and the resulting post-renderer inaccuracies. It is assumed that a binaural rendition for the reference pose ⁇ ⁇ is available. Reference is made to the above procedural description of the 3 axes case with yaw, pitch and roll correction.
- Part of this procedure is the calculation of approximate renditions for these virtual probing poses, for instance by carrying out low-complexity post-renderer operations of US 63/386,465 or U.S.63/340,181 using the previously calculated pitch and yaw correction metadata from the steps above and the pre-renditions for probing poses ⁇ ⁇ and ⁇ # .
- roll correction metadata is re- calculated, e.g., using a technique described in U.S.63/340,181.
- Part of this procedure is the calculation of approximate renditions for these virtual probing poses, for instance by carrying out low-complexity post-renderer operations of US 63/386,465 or U.S.63/340,181 using the previously calculated roll and yaw correction metadata and/or the previously performed pre-renditions.
- the interpolating operations between the pre-renditions for poses ⁇ ⁇ and ⁇ # may also involve applying post-renderer techniques using the previously enhanced roll correction metadata.
- An approximate rendition for pose P ⁇ > can be calculated using the pre- rendition for probing pose ⁇ - applying post-rendering techniques using the previously calculated yaw correction metadata.
- yaw correction metadata is enhanced using the pre-renditions for reference pose ⁇ ⁇ and probing pose ⁇ . and an approximate rendition for virtual probing pose ⁇ 7 .
- This enhancement step may involve applying post-renderer operations using the previously enhanced roll and pitch metadata and the available pre-renditions for probing poses ⁇ ⁇ , ⁇ # , and/or ⁇ -.
- Each of the above-described metadata enhancement steps relies on previously calculated metadata. It is thus possible to achieve even more enhancements by carrying out multiple iterations.
- FIG.5 is a flow chart illustrating processing in the main device 2 in accordance with embodiments of a second aspect of the invention, relating to a method of rendering audio to facilitate split rendering with pose correction around yaw axis and pitch axis.
- the flow chart may be broken into various blocks or partitions, such as blocks S21 – S26. Processing for the various blocks of FIG.5, which may be described as operations, processes, methods, steps, acts or functions, may commence at block S21.
- step S21 (obtain audio content) a first bitstream is received and decoded (e.g., by decoder 11) to obtain an immersive audio content A and in step S22 (obtain reference pose) a reference pose P' is obtained.
- the reference pose may be an assumed pose (e.g. straight ahead) or may be based on pose information received from the lightweight device 3.
- step S21 may be followed by step S22.
- step S23 render Binref
- the immersive audio content A is rendered (e.g., by renderer 12) into a reference binaural representation, Binref, corresponding to a reference pose P'.
- Step S23 may be followed by step S24.
- the immersive audio content A is rendered (e.g., by renderer 12) into two binaural pre-renditions Bin n , wherein the binaural pre-renditions correspond to two probing poses Pn deviating from the reference pose by rotation around both yaw and pitch axes.
- each probing pose deviates form the reference pose by rotation around a probing axis 104 (see figure 2) with the same origin as the yaw and pitch axes, and extending between the yaw axis 101 and the pitch axis 102.
- step S24 may be followed by step S25.
- step S25 (compute M and H)
- yaw metadata M representing a deviation around the yaw axis
- pitch metadata H representing deviation around the pitch axis are computed (e.g. by metadata generator 13) for each probing pose Pn.
- step S25 may be followed by step S26.
- step S26 encode and output bitstream
- the reference binaural representation Bin ref and the yaw metadata M and the pitch metadata H are encoded (e.g., by encoders 14, 15) in an output bitstream b2, which is subsequently outputted on an appropriate communication channel.
- the yaw and pitch metadata may be quantized and encoded based on symmetries in reconstruction metadata.
- the metadata may be encoded using differential coding between metadata relating to different (symmetrical) probing poses.
- Such complete reconstruction metadata may include, for each time-frequency tile, a complex or real 2x2 transformation matrix ? @ .
- the yaw metadata may include, for each time-frequency tile, a complex or real 2x2 yaw correction matrix M.
- the pitch metadata may include, for each time-frequency tile, a real 2x2 diagonal pitch correction matrix H.
- the light-weight post renderer device 3 sends the reference head pose ⁇ ⁇ to the heavy weight pre renderer device 2 through a back channel.
- the heavy weight device uses the ⁇ ⁇ pose to generate a reference binaural signal '() ⁇ and metadata such that the post renderer can do the pose correction from ⁇ ⁇ to the actual pose ⁇ and generate '() ⁇ from '() ⁇ using the metadata, wherein '() ⁇ has all the spatial cues as per Pose ⁇ .
- the deviation between ⁇ ⁇ and ⁇ depends on the motion-to-sound latency as described in this document.
- $ %, ⁇ 2 , ⁇ 2 is the covariance binaural signal '() ⁇
- $ %, ⁇ F, ⁇ 2 is the covariance and binaural signal '() ⁇ generated with pose ⁇ ⁇
- $ %, ⁇ , ⁇ is the covariance matrix of left and right channels of binaural signal '() ⁇ that is generated with pose ⁇ ⁇ .
- the post renderer decodes O ⁇ , E ⁇ F , O ⁇ # , E ⁇ e and '() ⁇ [#M ⁇ ] . Furthermore, if the actual pose P at the post renderer is not equal to either ⁇ ⁇ or ⁇ # or ⁇ ⁇ then the parameters corresponding to pose ⁇ ⁇ or ⁇ # or both are interpolated or extrapolated using linear interpolation that includes choosing two pose points out of ⁇ ⁇ , ⁇ ⁇ and ⁇ # that are closest to pose P, where in the two pose points may be different for yaw and pitch interpolation or extrapolation. Then the parameters E :)D O are interpolated or extrapolated between the two chosen pose points using linear interpolation.
- the number of pre-renditions may be controlled by choosing the pose points based on perceptual importance as follows. Let the 3DOF pose angles along yaw, pitch and roll axes in pose ⁇ ⁇ be * ⁇ AB , , ⁇ AB , 5 ⁇ AB respectively and the deviations in angles along yaw, pitch and roll axes between P and ⁇ ⁇ be * C , , C , 5 C .
- following probing pose points are selected to generate side information for rotations along yaw, pitch and roll axes.
- Side information corresponding to ⁇ ⁇ and ⁇ # can be computed as per U.S.63/340,181.
- ITD internal time difference
- O ⁇ - ⁇ h ⁇ , ⁇ - 0 0 h ⁇ , ⁇ - ⁇ , here where, $ %, ⁇ 2 , ⁇ 2 is the is the covariance matrix of reference binaural signal '() ⁇ - that is generated with with pose ⁇ -.
- O ⁇ - is quantized and coded and multiplexed into bitstream along with the coded bits for yaw and roll related side information and coded '() ⁇ signal.
- Deviation in roll angle may change ITD cues and hence it may be desired to model roll deviation with complex gain parameters in low frequencies (e.g., 0-2kHz) and with real only gain parameters in high frequencies (e.g., above 2 kHz). Side information corresponding to roll probing pose ⁇ .
- Prediction parameters may be computed as per U.S.63/340,181 with modifications as shown below. In an example implementation, same modifications are applied to the side information corresponding to yaw probing poses, E ⁇ F and E ⁇ e .
- FIG.6 is a flow chart illustrating processing in the lightweight device 3 in accordance with embodiments of a further aspect of the invention, relating to a method of split rendering with pose correction around multiple rotational axes.
- the flow chart may be broken into various blocks or partitions, such as blocks S41 – S45. Processing for the various blocks of FIG.6, which may be described as operations, processes, methods, steps, acts or functions, may commence at block S41.
- step S41 receive and decode bitstream
- a bitstream is received and decoded (e.g., by decoders 22, 23) from a main device (e.g., main device 2) to obtain a reference binaural representation Bin ref and first reconstruction metadata M, H associated with a set of probing poses Pn representing deviation from a reference pose P' by rotation around the multiple rotational axes.
- Step S41 may be followed by step S42.
- step S42 detect current head pose
- a current head pose P is detected (e.g., by head-tracker 25).
- step S42 may be followed by step S43.
- Step S43 (for each axis) is the beginning of a loop that includes one or more of steps S44 – S45.
- the loop is performed for each of the rotational axes, e.g., for yaw, pitch and roll, respectively.
- Step S43 may be followed by step S44 when additional processing is required for additional rotational axis. Otherwise step S43 may be followed by step S46 when processing is not required for any additional rotational axis.
- step S44 select probing pose
- Step S44 may be followed by step S45.
- step S45 (generate second metadata) second, axis-specific reconstruction metadata M ⁇ , M ⁇ , M ⁇ is determined based on the first metadata associated with the selected probing pose and a difference between the reference pose and the current head pose along the particular rotational axis.
- Step S46 may be followed by step S43 or step S46 when the processing loop is complete.
- step S46 (determine Binout) a binaural output Binout corresponding to the current head pose is determined based on the reference binaural representation Binref and the second, axis-specific reconstruction metadata for each rotational axis.
- An indication of the reference pose (P') may be obtained from the bitstream.
- the reference pose (P') can be determined based on an expected delay of transmission to the main device.
- the set of probing poses may be obtained from the bitstream, but may also be obtained by adding a set of offsets to the reference pose. Such offsets may be pre-defined (e.g., known before-hand) or may be obtained from the bitstream.
- E l is computed by performing linear interpolation or extrapolation on O ⁇ - , :)D E ⁇ based on the pitch angle in pose P and pitch angle in ⁇ - and ⁇ ⁇ .
- E k is computed by performing linear interpolation or extrapolation on E ⁇ h , :)D E ⁇ based on the roll angle in pose P and roll angle in ⁇ . and ⁇ ⁇ .
- ⁇ C, ⁇ is not transmitted to the post renderer and M matrix is computed with an additional gain matrix G as mentioned above.
- the additional gain matrix G which here was computed with respect to pose P4, may be computed for any pre-rendition including P1 and P2.
- FIG.7 is a flow chart illustrating processing in the main device 2 in accordance with embodiments of a yet further aspect of the invention.
- the flow chart may be broken into various blocks or partitions, such as blocks S51 – S58. Processing for the various blocks of FIG. 7, which may be described as operations, processes, methods, steps, acts or functions, may commence at block S51.
- step S51 (obtain audio content), a first bitstream is received and decoded (e.g., by decoder 11) to obtain an immersive audio content A.
- step S51 may be followed by step S52.
- step S52 receive head pose info
- a second bitstream is received and decoded (e.g., by decoder 17) to receive head pose information P associated with a user of a lightweight processing device.
- step S53 may be followed by step S53.
- step S33 determine reference pose
- a reference pose P' is determined (e.g. in decoder 17) based on the received head pose information.
- step S53 may be followed by step S54.
- step S54 the immersive audio content A is rendered into a reference binaural representation Binref corresponding to a reference pose (e.g., by renderer 12).
- step S54 may be followed by step S55.
- step S55 pre-rendering
- the immersive audio content A is rendered (e.g., by rendered 12) into one or more binaural pre-renditions Bin n the pre-renditions corresponding to one or more probing poses Pn deviating from the reference pose about a rotational axis.
- step S55 may be followed by step S56.
- step S56 generate metadata
- reconstruction metadata is computed (e.g., by metadata generator 13) to enable reconstruction of the binaural pre-renditions Bin n from the reference binaural representation Binref.
- the reconstruction metadata includes, for each time- frequency tile, a transformation matrix ? @ .
- Step S56 may be followed by step S57.
- step S57 enhanced metadata M is computed (e.g., by metadata generator 13) by multiplying each reconstruction matrix ?
- the rotational axis may be the yaw axis and/or the roll axis.
- step S58 encode and output bitstream
- the reference binaural representation Binref and the enhanced metadata M are encoded (e.g., by encoders 14, 15) in an output bitstream b2, which is subsequently outputted on an appropriate communication channel.
- the enhanced metadata may be quantized and encoded based on symmetries in reconstruction metadata.
- the metadata may be encoded using differential coding between metadata relating to different (symmetrical) probing poses.
- the set of probing poses includes pitch probing poses deviating from the reference pose only by rotation around the pitch axis.
- pitch reconstruction metadata is calculated, to enable reconstruction of binaural pre-renditions Bin n corresponding to the pitch probing poses from the reference binaural representation Binref , wherein the pitch reconstruction metadata includes, for each time-frequency tile, a diagonal real 2x2 pitch correction matrix H.
- FIG.8 is a flow chart illustrating processing in the main device 2 in accordance with embodiments of a still further aspect of the invention.
- the flow chart may be broken into various blocks or partitions, such as blocks S31 – S37. Processing for the various blocks of FIG. 8, which may be described as operations, processes, methods, steps, acts or functions, may commence at block S31.
- step S31 (obtain audio content), a first bitstream is received and decoded (e.g., by decoder 11) to obtain an immersive audio content A.
- Step S31 may be followed by step S32.
- step S32 receive head pose info
- a second bitstream is received and decoded (e.g., by decoder 17) to receive head pose information (P, ⁇ P, ⁇ P) associated with a user of a lightweight processing device.
- step S32 may be followed by step S33.
- step S33 determine reference pose
- a reference pose P' and at least one of a head pose rotation axis ⁇ P and a head pose rate of rotation ⁇ P is determined (e.g. in decoder 17) based on the received head pose information.
- step S33 may be followed by step S34.
- step S34 (render Bin ref )
- the immersive audio content A is rendered (e.g., by renderer 12) into a reference binaural representation, Binref, corresponding to a reference pose P'.
- step S34 may be followed by step S35.
- step S35 pre-rendering
- the immersive audio content (A) is rendered (e.g., by renderer 12) into a set binaural pre-renditions Binn, wherein the binaural pre-renditions correspond to a set of probing poses P n rotated with respect to the reference pose, wherein the probing poses are selected based on the head pose information (P, ⁇ P, ⁇ P).
- step S35 may be followed by step S62.
- step S36 compute M
- reconstruction metadata M is computed (e.g. by metadata generator 13), to enable reconstruction of the binaural pre-renditions Bin n from the reference binaural representation Binref.
- Step S36 may be followed by step S37.
- step S37 encode and output bitstream
- the reference binaural representation Binref and the reconstruction metadata M are encoded (e.g., by encoders 14, 15) in an output bitstream b 2 , which is subsequently output on an appropriate communication channel.
- the reconstruction metadata may be quantized and encoded based on symmetries in reconstruction metadata.
- the reconstruction metadata may be encoded using differential coding between metadata relating to different (symmetrical) probing poses.
- the probing poses P n may deviate from the reference pose P' by rotation around this head pose rotation axis ⁇ P.
- the probing poses Pn may be symmetrically distributed around the reference pose P'.
- the probing poses P n may include only one probing pose around each rotational degree of freedom.
- a head pose rate of rotation ⁇ P when the head pose rate of rotation ⁇ P is below a predefined threshold value the probing poses Pn may be selected to deviate from the reference pose by less than a first angle ⁇ lower , and when the head pose rate of rotation ⁇ P is above the threshold value the probing poses Pn may deviate from the reference pose by more than a second angle ⁇ upper , wherein the first angle ⁇ lower is smaller than the second angle ⁇ upper.
- a light-weight post renderer device 3 sends the reference head pose ⁇ ⁇ to heavy weight pre renderer device 2 through a back channel.
- the heavy weight device uses the ⁇ ⁇ to generate a reference binaural signal '() ⁇ and metadata M such that the post renderer can do the pose correction from ⁇ ⁇ to the actual pose ⁇ and generates '() ⁇ from '() ⁇ using the metadata, wherein '() ⁇ has all the spatial cues as per pose ⁇ .
- the deviation between ⁇ ⁇ and ⁇ depends on the motion-to-sound latency.
- the number of pre-renditions are controlled by choosing the pose points based on an estimation of the head movement velocity (rate of rotation, ⁇ P ) or direction of head movement (head pose rotation axis, ⁇ P) or both.
- the head movement velocity and direction of movement can be computed at the pre- renderer 2.
- post-renderer may provide the velocity and direction of movement along with pose information. With this information, the pre-renderer can significantly reduce the number of probing pose points for pre-renditions by choosing the probing pose points along the axis of head movement.
- ⁇ ′ is the reference pose from post renderer
- ⁇ ⁇ + (*, ,, 5 is a pose rotated around the head movement axis ⁇ P, in the direction of head movement.
- velocity and acceleration of head movement is used to further limit the number of pose points to one to generate side information.
- *, , :)D 5 in the probing pose points can be set to a lower value (e.g. smaller than a lower boundary ⁇ lower ) and would be sufficient to extrapolate the side information corresponding to ⁇ *, ⁇ , :)D ⁇ 5.
- the value of *, , :)D 5 is controlled based on head velocity and acceleration.
- Side information for ⁇ ⁇ , ⁇ # and ⁇ - can be computed as per the above sections.
- the computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware.
- PC personal computer
- PDA personal digital assistant
- a cellular telephone a smartphone
- web appliance a web appliance
- network router switch or bridge
- the present disclosure shall relate to any collection of computer hardware that individually or jointly executes instructions to perform any one or more of the concepts discussed herein.
- FIG.9 shows a schematic block diagram of an example electronic device or architecture 200 (e.g., an apparatus 200) suitable for implementing example embodiments of the present disclosure.
- Architecture 200 includes but is not limited to main processing devices and lightweight processing devices as described in relation to FIG.3.
- the architecture 200 includes central processing unit (CPU) 201 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 202 or a program loaded from, for example, storage unit 208 to random access memory (RAM) 203.
- the CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples the processor 201 may be multiple processors.
- I/O interface 205 input unit 206, that may include a keyboard, a mouse, or the like; output unit 207 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208 including a hard disk, or another suitable storage device; and communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).
- input unit 206 that may include a keyboard, a mouse, or the like
- output unit 207 that may include a display such as a liquid crystal display (LCD) and one or more speakers
- communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).
- input unit 206 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
- output unit 207 include systems with various number of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
- communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I/O interface 205, as required.
- Removable medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 210, so that a computer program read therefrom is installed into storage unit 208, as required.
- Removable medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 210, so that a computer program read therefrom is installed into storage unit 208, as required.
- apparatus 200 is described as including the above-described components, in real applications, it is possible to add, remove, and/or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure. [166]
- the processes described above may be implemented as computer software programs or on a computer-readable storage medium.
- embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods.
- the computer program may be downloaded and mounted from the network via the communication unit 209, and/or installed from the removable medium 211, as shown in FIG.9.
- various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof.
- control circuitry e.g., CPU 201 in combination with other components of FIG.9
- the control circuitry may be performing the actions described in this disclosure.
- Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and/or other computing device(s), which may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non- limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
- various blocks shown in the flowcharts may be viewed as method steps, and/or as operations that result from operation of computer program code, and/or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s).
- embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
- Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages.
- These computer program codes may be provided to one or more processors of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by one or more processors of the computer or other programmable data processing apparatus, cause the functions/operations specified in the flowcharts and/or block diagrams to be implemented.
- the program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and/or servers.
- the one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s).
- Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
- the software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media).
- computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.
- Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as ROM, PROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.
- communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Stereophonic System (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363448830P | 2023-02-28 | 2023-02-28 | |
| PCT/US2024/017570 WO2024182457A1 (en) | 2023-02-28 | 2024-02-27 | Split binaural rendering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4674142A1 true EP4674142A1 (en) | 2026-01-07 |
Family
ID=97323105
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24715359.6A Pending EP4674142A1 (en) | 2023-02-28 | 2024-02-27 | Split binaural rendering |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4674142A1 (en) |
| CN (1) | CN120814251A (en) |
-
2024
- 2024-02-27 CN CN202480014532.2A patent/CN120814251A/en active Pending
- 2024-02-27 EP EP24715359.6A patent/EP4674142A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN120814251A (en) | 2025-10-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN101490743B (en) | Dynamic decoding of binaural audio signals | |
| US20230370803A1 (en) | Spatial Audio Augmentation | |
| CN111527760B (en) | Method and system for processing global transitions between listening locations in a virtual reality environment | |
| CN112771479B (en) | 6DOF and 3DOF backward compatibility | |
| US20250220384A1 (en) | Method and Apparatus for Efficient Delivery of Edge Based Rendering of 6DOF MPEG-I Immersive Audio | |
| JP2023533414A (en) | Adaptive audio delivery and rendering | |
| KR20210071972A (en) | Signal processing apparatus and method, and program | |
| WO2024182457A1 (en) | Split binaural rendering | |
| JP2025531871A (en) | Head-tracked split rendering and head-related transfer function personalization | |
| US12604152B2 (en) | Binarual rendering | |
| KR20210055278A (en) | Method and system for hybrid video coding | |
| EP4674142A1 (en) | Split binaural rendering | |
| WO2025136874A1 (en) | Pose correction metadata for interactive headtracking | |
| WO2018190151A1 (en) | Signal processing device, method, and program | |
| TWI822032B (en) | Video display systems, portable video display apparatus, and video enhancement method | |
| KR20250103678A (en) | Efficient time delay synthesis | |
| HK40130038A (en) | Binarual rendering | |
| US20260088036A1 (en) | Audio rendering method, system, and electronic device | |
| CN116670758A (en) | Sound component rotation for orientation-dependent coding schemes | |
| WO2025259685A1 (en) | Partitioned processing for rendering of audio scenes | |
| WO2026006293A1 (en) | Transmission of interactive audio content | |
| JP2023550934A (en) | Immersive media compatibility | |
| CN115966216A (en) | Audio stream processing method and device | |
| TW201528251A (en) | Apparatus and method for efficient object metadata coding |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250820 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| P01 | Opt-out of the competence of the unified patent court (upc) registered |
Free format text: CASE NUMBER: UPC_APP_0004278_4674142/2026 Effective date: 20260206 |