EP4642055A1 - Rendering audio - Google Patents

Rendering audio

Info

Publication number
EP4642055A1
EP4642055A1 EP25166950.3A EP25166950A EP4642055A1 EP 4642055 A1 EP4642055 A1 EP 4642055A1 EP 25166950 A EP25166950 A EP 25166950A EP 4642055 A1 EP4642055 A1 EP 4642055A1
Authority
EP
European Patent Office
Prior art keywords
audio
loudspeakers
dimensional scene
audio object
rendering
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP25166950.3A
Other languages
German (de)
French (fr)
Inventor
Jussi Artturi LEPPÄNEN
Miikka Tapani Vilermo
Arto Juhani Lehtiniemi
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Technologies Oy
Original Assignee
Nokia Technologies Oy
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Technologies Oy filed Critical Nokia Technologies Oy
Publication of EP4642055A1 publication Critical patent/EP4642055A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R5/00Stereophonic arrangements
    • H04R5/02Spatial or constructional arrangements of loudspeakers
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/11Positioning of individual sound objects, e.g. moving airplane, within a sound field

Definitions

  • the present specification relates to rendering audio, particularly in rendering audio at loudspeaker(s).
  • this specification provides an apparatus comprising: means for rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises: means for determining a number of users consuming the three-dimensional scene; and means for configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
  • Some examples include means for determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises means for configuring, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, an apparent position of the second audio object to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • the means for configuring the apparent positions of the first and second audio objects to be mapped on to the positions of the first and second loudspeakers comprises means for performing one or more of shifting, rotating and/or scaling the three-dimensional scene.
  • the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped on to the first and second loudspeakers.
  • Some examples include: means for determining whether the three-dimensional scene comprises more than a threshold number of audio objects; wherein if the three-dimensional scene comprises more than the threshold number of audio objects, the means for rendering comprises means for configuring apparent position(s) of at least one of the audio objects to be mapped outside of the first loudspeaker arrangement.
  • the threshold number is two.
  • the apparent positions of each of at least two of the audio objects is configured to be mapped on positions of each of at least two loudspeakers respectively.
  • rendering audio corresponding to audio objects to be mapped outside of the first loudspeaker arrangement comprises rendering said audio with spatial extent.
  • the means for rendering comprises means for rendering the audio independent of changes in positions of the more than one user.
  • Some examples include means for configuring one or more of a position, orientation, and/or size of a visual representation of any audio object to be based at least in part on the apparent position of the respective audio object.
  • the means for rendering further comprises: means for rendering, based on determining that a single first user is consuming the three-dimensional scene, audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user. In some examples, the means for rendering further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user.
  • the three-dimensional scene is one or more of a virtual reality scene, augmented reality scene, or a mixed reality scene.
  • the first loudspeaker is selected from the plurality of loudspeakers based on which of the plurality of loudspeakers have the least distance from the first audio object.
  • the distance relates to one or more of horizontal distance, vertical distance, or angular distance.
  • the means may comprise: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured, with the at least one processor, to cause the performance of the apparatus.
  • this specification describes a method comprising: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
  • Some examples include determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises means for configuring, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, an apparent position of the second audio object to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • configuring the apparent positions of the first and second audio objects to be mapped on to the positions of the first and second loudspeakers comprises performing one or more of shifting, rotating and/or scaling the three-dimensional scene.
  • the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped on to the first and second loudspeakers.
  • Some examples include: determining whether the three-dimensional scene comprises more than a threshold number of audio objects; wherein if the three-dimensional scene comprises more than the threshold number of audio objects, the rendering comprises means for configuring apparent position(s) of at least one of the audio objects to be mapped outside of the first loudspeaker arrangement.
  • the threshold number is two.
  • the apparent positions of each of at least two of the audio objects is configured to be mapped on positions of each of at least two loudspeakers respectively.
  • rendering audio corresponding to audio objects to be mapped outside of the first loudspeaker arrangement comprises rendering said audio with spatial extent.
  • the means for rendering comprises means for rendering the audio independent of changes in positions of the more than one user.
  • Some examples include means for configuring one or more of a position, orientation, and/or size of a visual representation of any audio object to be based at least in part on the apparent position of the respective audio object.
  • the means for rendering further comprises: means for rendering, based on determining that a single first user is consuming the three-dimensional scene, audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user. In some examples, the means for rendering further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user.
  • the three-dimensional scene is one or more of a virtual reality scene, augmented reality scene, or a mixed reality scene.
  • the first loudspeaker is selected from the plurality of loudspeakers based on which of the plurality of loudspeakers have the least distance from the first audio object.
  • the distance relates to one or more of horizontal distance, vertical distance, or angular distance.
  • this specification describes an apparatus configured to perform any method as described with reference to the second aspect.
  • this specification describes computer-readable instructions which, when executed by computing apparatus, cause the computing apparatus to perform any method as described with reference to the second aspect.
  • this specification describes a computer program comprising instructions for causing an apparatus to perform at least the following: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • this specification describes a computer-readable medium (such as a non-transitory computer-readable medium) comprising program instructions stored thereon for performing at least the following: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • this specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: render audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • this specification describes an apparatus comprising a first module configured to render audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining, by a second module, a number of users consuming the three-dimensional scene; and configuring, by a third module, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • FIG. 1 is a block diagram of an example system, indicated generally by the reference numeral 10.
  • the system 10 shows a three-dimensional scene, such as a virtual reality scene, augmented reality scene, and/or a mixed reality scene, where the system 10 comprises a plurality of loudspeakers 11a-11e.
  • the loudspeakers 11 may be arranged to provide six degrees of freedom (6DoF) audio to user(s) experiencing the scene, such as a virtual reality, augmented reality, and/or mixed reality scene (referred to as VR scene hereon). Audio may be processed (e.g. as processed in MPEG-I) for being rendered as 6DoF audio using a plurality of loudspeakers 11a-11e (e.g. instead of being rendered with headphones/other in-ear devices).
  • 6DoF six degrees of freedom
  • Audio may be processed (e.g. as processed in MPEG-I) for being rendered as 6DoF audio using a plurality of loudspeakers 11a-11e (e.g. instead of being rendered with headphones/other in
  • the system 10 further comprises a user 12 (user 12 shown to be in a position 12a) and an audio object 13a.
  • the VR scene may be defined such that the audio object 13a corresponds to a visual object shown to the user 12 in the position of the audio object 13a, where the audio object is positioned within the periphery of the loudspeaker arrangement of the loudspeakers 11a-11e.
  • An apparent position 13b of the audio object 13a may be configured based, at least in part, on the position 12a of the user 12 and the positioning of one or more of the plurality of loudspeakers 11a-11e.
  • line 15 represents direction of the audio object 13a relative to the user position 12a, and therefore the apparent position 13b is configured to be placed along the line 15 and between the loudspeakers 11a and 11b (denoted by the line 14).
  • loudspeakers 11a and 11b may render audio such that the apparent position 13b of the audio object 13a is configured to fall along the line 15.
  • the user 12 may perceive the audio to be generated from a direction of the audio object 13a.
  • a determination of which loudspeakers e.g. loudspeaker pair
  • the audio may be panned based on movement of the user 12 such that the apparent position 13b may move along the line 14 so that the perception of the audio being rendered from the direction of the audio object 13a.
  • FIG. 2 is a block diagram of an example system, indicated generally by the reference numeral 20.
  • the system 20 shows the VR scene of FIG. 1 , with the position of the user 12 being moved from position 12a to position 12b, and line 16 representing a direction of the audio object 13a relative to the user position 12b.
  • the audio object 13a has an updated apparent position 13c along the line 16 based on the updated position 12b of the user 12.
  • the apparent position 13c is configured to be placed along the line 16 and between the loudspeakers 11a and 11e (denoted by the line 17).
  • loudspeakers 11a and 11e may render audio such that the apparent position 13c of the audio object 13a is configured to fall along the line 16.
  • the user 12 may perceive the audio to be generated from a direction of the audio object 13a.
  • FIG. 3 is a block diagram of an example system, indicated generally by the reference numeral 30.
  • the system 30 shows the VR scene of FIG. 1 , with the position of the user 12 being moved from position 12a to position 12c, and line 15 representing a direction of the audio object 13a relative to the user position 12c.
  • the user 12 is shown to move closer to the audio object 13a along the line 15.
  • the gain of the loudspeakers 11a and 11b may be adjusted so as to simulate distance gain effects.
  • the audio may be rendered by the loudspeakers 11a and 11b at the position 13d at a higher volume (denoted by a larger circle representation) to that shown in FIG. 1 .
  • FIG. 4 is a block diagram of an example system, indicated generally by the reference numeral 40.
  • the system 40 shows a VR scene similar to that of FIG. 1 , with the loudspeakers 11a-11e arranged in a similar manner.
  • the system 40 further shows an audio object 41a, such that the direction, shown by the line 42, of the audio object 41a relative to the user 12 matches a direction of the loudspeaker 11a relative to the user 12.
  • the apparent position 41b is placed behind the audio object 41a along the line 42 such that the apparent position 41b may fall on the position of the loudspeaker 11a. Therefore, the audio for the audio object 41a may be rendered from a single loudspeaker 11a.
  • the apparent position of the audio object need not be changed based on the user's movement, as the audio from the audio object may appear to be generated from the direction of the respective loudspeaker regardless of the change in user position.
  • the gain of the loudspeakers may be adjusted so as to simulate distance gain effects.
  • FIG. 5 is a block diagram of an example system, indicated generally by the reference numeral 50.
  • the system 50 shows a VR scene similar to that of FIG. 1 , with another user 51 and an audio object 52a.
  • line 54 represents the direction of the audio object 52a relative to the user 12.
  • an apparent position 52b of the audio object 52a is configured to appear from behind the audio object 52a along the line 54, thus the audio being rendered by the loudspeakers 11a and 11b and the audio may be panned along the line 55 based on movement of user 12.
  • the apparent position of the audio object is at the apparent position 52b, the user 51 may perceive the audio object 52a to be placed in an incorrect position.
  • Line 56 represents the direction of the apparent position 52b relative to the user 51. For example, when audio of the audio object 52a appears to be generated from the apparent position 52b, the user 51 may perceive the audio object to be located along line 56, for example, at position 53.
  • the mismatch may be more noticeable if the user 51 is stationary while the user 12 moves, as the movement of user 12 may cause the apparent position 52b to be moved as well, thus causing the audio object position 53 to move around for the user 51 even though user 51 is not moving.
  • Such a mismatch may not be desirable and may interfere in the user 51 having an immersive experience.
  • the example embodiments described below may aim to address problems arising from a situation as described in FIG. 5 .
  • FIG. 6 is a flowchart of an algorithm, indicated generally by the reference numeral 60, in accordance with an example embodiment.
  • FIG. 6 may be viewed in conjunction with FIG. 7 for better understanding.
  • FIG. 7 is a block diagram of a system, indicated generally by the reference numeral 70, in accordance with an example embodiment.
  • the system 70 shows a scene 71 with a single user 12, and a scene 72 with a plurality of users (users 12 and 78).
  • the scene 71 may be similar to the scene as described with reference to FIG. 1 .
  • the algorithm 60 describes a method for rendering audio for a three-dimensional scene, such as a VR scene, at one or more of a plurality of loudspeakers 11a-11b.
  • the plurality of loudspeakers is arranged in a first loudspeaker arrangement (e.g. shown by the loudspeakers 11a-11e) suitable for providing spatial audio corresponding to the three-dimensional scene.
  • the three-dimensional scene may comprise a first audio object 73a.
  • the algorithm 60 may start at operation 601 for determining a number of users consuming the three-dimensional scene.
  • the algorithm moves to operation 604, and the apparent position 76 of the first audio object 73a is configured to be mapped, as shown by the arrow 77, to a position of a first loudspeaker, such as the loudspeaker 11a of the plurality of loudspeakers.
  • Configuring the apparent position may comprise rendering audio corresponding to the first audio object 73a from the loudspeaker 11a.
  • the rendering may comprise rendering the audio independent of changes in positions of the users 12 and/or 78.
  • the algorithm may move to the operation 603, and the apparent position (e.g. 73b) of the first audio object 73a is configured to be determined based on an initial position of the audio object (e.g. position of the audio object 73a) and/or relative position of the user (e.g. position of user 12).
  • audio may be rendered from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object 73a and a position of the first user 12, and optionally based on the positioning of one or more of the plurality of loudspeakers 11a-11e.
  • line 74 represents direction of the audio object 73a relative to the user position 12, and therefore the apparent position 73b is configured to be placed along the line 74 and between the loudspeakers 11a and 11b (denoted by the line 75).
  • loudspeakers 11a and 11b may render audio such that the apparent position 73b of the audio object 13a is configured to fall along the line 75.
  • the user 12 may perceive the audio to be generated from a direction of the audio object 73a.
  • rendering audio in this scenario may further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user 12 (e.g. as shown with reference to FIGs. 2 and 3 ).
  • both the users 12 and 78 may perceive the audio in a similar manner as being rendered from the position of the loudspeaker 11a, thus preventing a situation of audio and visual mismatch as described with reference to FIG. 5 . Further, the apparent position 76 may remain unchanged even if one of the users move, or the users move in different directions, as the apparent position 76 is not determined dependent on the position of the users.
  • one or more of a position, orientation, and/or size of a visual representation of the audio object 73a may be configured to be based at least in part on the apparent position 76 of the respective audio object 73a.
  • a visual representation of the audio object 73a is also configured to appear at the apparent position 76 in order to improve consistency between the positions of the audio and visual objects.
  • the apparent position 73b may gradually change to the apparent position 76 such that the user 12 may perceive the change to be gradual rather than an abrupt change.
  • the first loudspeaker such as the loudspeaker 11a may be selected based, at least in part, on which of the plurality of loudspeakers have the least distance from the first audio object.
  • the loudspeaker 11a is selected from the plurality of loudspeakers as the loudspeaker 11a may be closest to the audio object 73a.
  • the distance may relate to one or more of horizontal distance between the loudspeaker(s) and the audio object 73a, vertical distance between the loudspeaker(s) and the audio object 73a, or angular distance between the loudspeaker(s) and the audio object 73a.
  • a loudspeaker e.g. loudspeaker 11a
  • the floor of the VR scene may match the real-life room floor as well as possible.
  • the VR scene may be modified by adjusting the height of the audio object (and any associated visual object) such that when placed at the loudspeaker position, the VR scene floor and the real-life room floor positions match.
  • any changes in audio rendering based on the user's relative position may be paused and/or cancelled when more than one user is experiencing the VR scene.
  • one or more of a position, orientation, and/or size of a visual representation of any audio object may be configured to be based at least in part on the apparent position of the respective audio object.
  • FIG. 8 is a flowchart of an algorithm, indicated generally by the reference numeral 80, in accordance with an example embodiment.
  • FIG. 8 may be viewed in conjunction with FIG. 9 for better understanding.
  • FIG. 9 is a block diagram of a system, indicated generally by the reference numeral 90, in accordance with an example embodiment.
  • the system 90 shows a three-dimensional scene (e.g. scene 91) with the plurality of loudspeakers 11a-11e, and two audio objects 92a (similar to the first audio object 73a) and 93a.
  • the audio may be rendered such that the scene is being consumed by more than one user, such as users 12 and 78 (the users are not depicted here for simplicity).
  • Scenes 97, 98, and 99 show how the scene 91 may be modified according to an example embodiment.
  • the algorithm 80 may be performed after operation 604 of algorithm 60, where an apparent position of the first audio object (73a, 92a) is mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • the algorithm 80 may start at operation 801, where it is determined that the three-dimensional scene comprises more than one audio objects, such as a second audio object 93a, in addition to a first audio object 92a. If it is determined that the scene comprises more than one audio object, the algorithm may move to operation 802.
  • an apparent position of the second audio object 93a may be configured to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • the mapping may comprise performing one or more of shifting, rotating and/or scaling the three-dimensional scene (e.g. one or more audio objects of the three-dimensional scene).
  • the scene may be shifted as shown by the arrow 94, such that the audio object 92a is mapped on the position of the loudspeaker 11a (e.g. similar to the audio object 73a being mapped on the loudspeaker 11a).
  • the new positions of the audio objects are shown by the apparent positions 92b and 93b in scene 97.
  • scene may be rotated as shown by the arrow 95, such that the position of the audio objects may be updated to the apparent positions 92c and 93c as shown in the scene 98.
  • the scene may be panned as shown by the arrow 96, such that the position of the audio objects may be updated to the apparent positions 92d and 93d as shown in the scene 99.
  • the audio objects 92a and 93a are mapped on the positions of loudspeakers 11a and 11b respectively.
  • plurality of users may perceive the audio from audio object 92a in a similar manner as being rendered from the position of the loudspeaker 11a and the audio from the audio object 93a in a similar manner as being rendered from the position of the loudspeaker 11b, thus preventing a situation of audio and visual mismatch as described with reference to FIG. 5 .
  • the apparent positions 92d and 93d may remain unchanged even if one of the users move, or the users move in different directions or different distances, as the apparent positions 92d and/or 93d are not determined dependent on the position of the users.
  • one or more of a position, orientation, and/or size of a visual representation of the audio objects 92a and 93a may be configured to be based at least in part on the apparent position 92d and 93d respectively.
  • a visual representation of the audio object 92a is configured to appear at the apparent position 92d
  • a visual representation of the audio object 93a is configured to appear at the apparent position 93d in order to improve consistency between the positions of the audio and visual objects.
  • the loudspeaker 11b is selected for rendering the audio from the audio object 93a based on the loudspeaker 11b being a nearest one (e.g. least horizontal, vertical, and/or angular distance) of the plurality of loudspeakers in relation to the second audio object 93a.
  • FIG. 10 is a flowchart of an algorithm, indicated generally by the reference numeral 100, in accordance with an example embodiment.
  • FIG. 10 may be viewed in conjunction with FIG. 11 for better understanding.
  • FIG. 11 is a block diagram of a system, indicated generally by the reference numeral 110, in accordance with an example embodiment.
  • the system 110 shows a three-dimensional scene (e.g. scene 111a) with the plurality of loudspeakers 11a-11e, and a plurality of audio objects 112a, 113a 114a, 115a, 116a, and 117a.
  • the audio may be rendered such that the scene is being consumed by more than one user, such as users 12 and 78 (the users are not depicted here for simplicity).
  • Scenes 111b-111d show how the scene 111a may be modified according to an example embodiment.
  • the algorithm 100 may start at operation 1001, where it is determined that the three-dimensional scene comprises more than a threshold number of audio objects, such as audio objects 112a- 117a.
  • the threshold number may be two or three.
  • the threshold may be dependent on the number of loudspeakers.
  • the threshold number may be half the number of loudspeakers (e.g. half the number rounded up or down to the nearest integer), the threshold number may be equal to number of speakers, the threshold number may be twice the number of loudspeakers, the threshold number may be the number of loudspeakers plus or minus a predefined number, or the like.
  • the algorithm may move to operation 1002.
  • the threshold number may be two, such that when there are more than two audio objects, operation 1002 may be performed.
  • apparent positions of at least one of the audio objects is configured to be mapped outside of the first loudspeaker arrangement.
  • the threshold number e.g. two
  • the remaining audio objects in addition to the threshold number may be mapped outside of the first loudspeaker arrangement.
  • the scene 111b shows how the audio objects 112a to 117a may be initially positioned relative to each other (e.g. audio object constellation shown using dotted lines between the positions 112b to 117b).
  • the operation 1002 may be performed accordingly.
  • the operation 1002 may be performed, for example by shifting, panning, and/or zooming one or more audio objects of the scene.
  • the scene 111c shows audio objects 112a (at position 112c) and 116a (at position 116c) being mapped on the loudspeakers 11b and 11c respectively, and also shows audio objects 113a, 114a, 115a, and 117a (at positions 113c, 114c, 115c, and 117c respectively) to be mapped outside of the loudspeaker arrangement of loudspeakers 11a-11e.
  • the rendering audio corresponding to audio objects mapped outside the first loudspeaker arrangement may comprise rendering said audio with spatial extent.
  • the loudspeakers 11b and 11c may be selected for the audio objects 112a (at position 112d) and 116a (at position 116d) respectively based on distance (e.g. horizontal, vertical, and/or angular) of the respective loudspeaker from the audio object, as described with reference to FIG. 9 .
  • the audio objects 112a and 116a may be selected for being mapped on a loudspeaker (11b, 11c respectively) as the two audio objects being located next to each other at the edge of the audio object constellation (e.g. based on triangulation of the audio object positions.
  • the scene may then be shifted, rotated, and scaled so that these audio objects 112d and 116d align with two loudspeakers 11b and 11c.
  • audio processing may be modified.
  • spatial extent may be applied to audio objects 113c, 114c, 115c, and 117c such that the user may perceive the rendered audio to be arriving from a generic direction of the loudspeakers 11b and 11c (e.g. such that any mismatch between the audio and visual objects may be less noticeable.
  • spatially extended sources may be rendered via creating multiple uncorrelated N "helper" audio objects from the original audio object and placing them around the original audio object.
  • the example embodiments above provide a system for loudspeaker rendering of 6DoF VR content to multiple users by modifying the VR scene based on the loudspeaker positions, and independent of the positions of each individual user, such that any mismatch between user perception of audio objects and corresponding visual objects may be minimized. This may be achieved by configuring apparent positions of one or more audio objects to be mapped on positions of the loudspeaker(s) respectively.
  • FIG. 12 is a schematic diagram of components of one or more of the example embodiments described previously, which hereafter are referred to generically as processing systems 300.
  • a processing system 300 may have a processor 302, a memory 304 closely coupled to the processor and comprised of a RAM 314 and ROM 312, and, optionally, user input 310 and a display 318.
  • the processing system 300 may comprise one or more network/apparatus interfaces 308 for connection to a network/apparatus, e.g. a modem which may be wired or wireless. Interface 308 may also operate as a connection to other apparatus such as device/apparatus which is not network side apparatus. Thus, direct connection between devices/apparatus without network participation is possible.
  • the processor 302 is connected to each of the other components in order to control operation thereof.
  • the memory 304 may comprise a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD).
  • the ROM 312 of the memory 304 stores, amongst other things, an operating system 315 and may store software applications 316.
  • the RAM 314 of the memory 304 is used by the processor 302 for the temporary storage of data.
  • the operating system 315 may contain computer program code which, when executed by the processor implements aspects of the algorithms 60, 80, 90 described above. Note that in the case of small device/apparatus the memory can be most suitable for small size usage i.e. not always hard disk drive (HDD) or solid-state drive (SSD) is used.
  • the processor 302 may take any suitable form. For instance, it may be a microcontroller, a plurality of microcontrollers, a processor, or a plurality of processors.
  • the processing system 300 may be a standalone computer, a server, a console, or a network thereof.
  • the processing system 300 and needed structural parts may be all inside device/apparatus such as IoT device/apparatus i.e. embedded to very small size.
  • the processing system 300 may also be associated with external software applications. These may be applications stored on a remote server device/apparatus and may run partly or exclusively on the remote server device/apparatus. These applications may be termed cloud-hosted applications.
  • the processing system 300 may be in communication with the remote server device/apparatus in order to utilize the software application stored there.
  • FIG. 13 shows tangible media, specifically a removable memory unit 365, storing computer-readable code which when run by a computer may perform methods according to example embodiments described above.
  • the removable memory unit 365 may be a memory stick, e.g. a USB memory stick, having internal memory 366 for storing the computer-readable code.
  • the internal memory 366 may be accessed by a computer system via a connector 367.
  • Other forms of tangible storage media may be used.
  • Tangible media can be any device/apparatus capable of storing data/information which data/information can be exchanged between devices/apparatus/network.
  • Embodiments of the present invention may be implemented in software, hardware, application logic or a combination of software, hardware and application logic.
  • the software, application logic and/or hardware may reside on memory, or any computer media.
  • the application logic, software or an instruction set is maintained on any one of various conventional computer-readable media.
  • a "memory" or “computer-readable medium” may be any non-transitory media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer.
  • references to, where relevant, "computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc., or a “processor” or “processing circuitry” etc. should be understood to encompass not only computers having differing architectures such as single/multi-processor architectures and sequencers/parallel architectures, but also specialised circuits such as field programmable gate arrays FPGA, application specify circuits ASIC, signal processing devices/apparatus and other devices/apparatus. References to computer program, instructions, code etc.
  • programmable processor firmware such as the programmable content of a hardware device/apparatus as instructions for a processor or configured or configuration settings for a fixed function device/apparatus, gate array, programmable logic device/apparatus, etc.
  • circuitry refers to all of the following: (a) hardware-only circuit implementations (such as implementations in only analogue and/or digital circuitry) and (b) to combinations of circuits and software (and/or firmware), such as (as applicable): (i) to a combination of processor(s) or (ii) to portions of processor(s)/software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a server, to perform various functions) and (c) to circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Stereophonic System (AREA)

Abstract

This specification describes an apparatus comprising means for rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises: means for determining a number of users consuming the three-dimensional scene; and means for configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.

Description

    Field
  • The present specification relates to rendering audio, particularly in rendering audio at loudspeaker(s).
  • Background
  • Rendering audio at loudspeakers is known. There remains a need for improvement in rendering audio at loudspeakers at multi-user settings.
  • Summary
  • In a first aspect, this specification provides an apparatus comprising: means for rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises: means for determining a number of users consuming the three-dimensional scene; and means for configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • In some examples, configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
  • Some examples include means for determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises means for configuring, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, an apparent position of the second audio object to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • In some examples, the means for configuring the apparent positions of the first and second audio objects to be mapped on to the positions of the first and second loudspeakers comprises means for performing one or more of shifting, rotating and/or scaling the three-dimensional scene. In some examples, the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped on to the first and second loudspeakers.
  • Some examples include: means for determining whether the three-dimensional scene comprises more than a threshold number of audio objects; wherein if the three-dimensional scene comprises more than the threshold number of audio objects, the means for rendering comprises means for configuring apparent position(s) of at least one of the audio objects to be mapped outside of the first loudspeaker arrangement. In some examples, the threshold number is two. In some examples, the apparent positions of each of at least two of the audio objects is configured to be mapped on positions of each of at least two loudspeakers respectively. In some examples, rendering audio corresponding to audio objects to be mapped outside of the first loudspeaker arrangement comprises rendering said audio with spatial extent.
  • In some examples, the means for rendering comprises means for rendering the audio independent of changes in positions of the more than one user.
  • Some examples include means for configuring one or more of a position, orientation, and/or size of a visual representation of any audio object to be based at least in part on the apparent position of the respective audio object.
  • In some examples, the means for rendering further comprises: means for rendering, based on determining that a single first user is consuming the three-dimensional scene, audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user. In some examples, the means for rendering further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user.
  • In some examples, the three-dimensional scene is one or more of a virtual reality scene, augmented reality scene, or a mixed reality scene.
  • In some examples, the first loudspeaker is selected from the plurality of loudspeakers based on which of the plurality of loudspeakers have the least distance from the first audio object.
  • In some examples, the distance relates to one or more of horizontal distance, vertical distance, or angular distance.
  • The means may comprise: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured, with the at least one processor, to cause the performance of the apparatus.
  • In a second aspect, this specification describes a method comprising: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • In some examples, configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
  • Some examples include determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises means for configuring, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, an apparent position of the second audio object to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • In some examples, configuring the apparent positions of the first and second audio objects to be mapped on to the positions of the first and second loudspeakers comprises performing one or more of shifting, rotating and/or scaling the three-dimensional scene. In some examples, the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped on to the first and second loudspeakers.
  • Some examples include: determining whether the three-dimensional scene comprises more than a threshold number of audio objects; wherein if the three-dimensional scene comprises more than the threshold number of audio objects, the rendering comprises means for configuring apparent position(s) of at least one of the audio objects to be mapped outside of the first loudspeaker arrangement. In some examples, the threshold number is two. In some examples, the apparent positions of each of at least two of the audio objects is configured to be mapped on positions of each of at least two loudspeakers respectively. In some examples, rendering audio corresponding to audio objects to be mapped outside of the first loudspeaker arrangement comprises rendering said audio with spatial extent.
  • In some examples, the means for rendering comprises means for rendering the audio independent of changes in positions of the more than one user.
  • Some examples include means for configuring one or more of a position, orientation, and/or size of a visual representation of any audio object to be based at least in part on the apparent position of the respective audio object.
  • In some examples, the means for rendering further comprises: means for rendering, based on determining that a single first user is consuming the three-dimensional scene, audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user. In some examples, the means for rendering further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user.
  • In some examples, the three-dimensional scene is one or more of a virtual reality scene, augmented reality scene, or a mixed reality scene.
  • In some examples, the first loudspeaker is selected from the plurality of loudspeakers based on which of the plurality of loudspeakers have the least distance from the first audio object.
  • In some examples, the distance relates to one or more of horizontal distance, vertical distance, or angular distance.
  • In a third aspect, this specification describes an apparatus configured to perform any method as described with reference to the second aspect.
  • In a fourth aspect, this specification describes computer-readable instructions which, when executed by computing apparatus, cause the computing apparatus to perform any method as described with reference to the second aspect.
  • In a fifth aspect, this specification describes a computer program comprising instructions for causing an apparatus to perform at least the following: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • In a sixth aspect, this specification describes a computer-readable medium (such as a non-transitory computer-readable medium) comprising program instructions stored thereon for performing at least the following: rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • In a seventh aspect, this specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: render audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users consuming the three-dimensional scene; and configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • In an eighth aspect, this specification describes an apparatus comprising a first module configured to render audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining, by a second module, a number of users consuming the three-dimensional scene; and configuring, by a third module, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  • Brief description of the drawings
  • Example embodiments will now be described, by way of example only, with reference to the following schematic drawings, in which:
    • FIGs. 1 to 5 are block diagrams of example systems;
    • FIG. 6 is a flowchart of an algorithm in accordance with an example embodiment;
    • FIG. 7 is a block diagram of a system in accordance with an example embodiment;
    • FIG. 8 is a flowchart of an algorithm in accordance with an example embodiment;
    • FIG. 9 is a block diagram of a system in accordance with an example embodiment;
    • FIG. 10 is a flowchart of an algorithm in accordance with an example embodiment;
    • FIG. 11 is a block diagram of a system in accordance with an example embodiment;
    • FIG. 12 is a block diagram of components of a system in accordance with an example embodiment; and
    • FIG. 13 shows an example of tangible media for storing computer-readable code which when run by a computer may perform methods according to example embodiments described above.
    Detailed description
  • The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in the specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
  • In the description and drawings, like reference numerals refer to like elements throughout.
  • FIG. 1 is a block diagram of an example system, indicated generally by the reference numeral 10. The system 10 shows a three-dimensional scene, such as a virtual reality scene, augmented reality scene, and/or a mixed reality scene, where the system 10 comprises a plurality of loudspeakers 11a-11e. For example, the loudspeakers 11 may be arranged to provide six degrees of freedom (6DoF) audio to user(s) experiencing the scene, such as a virtual reality, augmented reality, and/or mixed reality scene (referred to as VR scene hereon). Audio may be processed (e.g. as processed in MPEG-I) for being rendered as 6DoF audio using a plurality of loudspeakers 11a-11e (e.g. instead of being rendered with headphones/other in-ear devices).
  • The system 10 further comprises a user 12 (user 12 shown to be in a position 12a) and an audio object 13a. The VR scene may be defined such that the audio object 13a corresponds to a visual object shown to the user 12 in the position of the audio object 13a, where the audio object is positioned within the periphery of the loudspeaker arrangement of the loudspeakers 11a-11e. An apparent position 13b of the audio object 13a may be configured based, at least in part, on the position 12a of the user 12 and the positioning of one or more of the plurality of loudspeakers 11a-11e. For example, line 15 represents direction of the audio object 13a relative to the user position 12a, and therefore the apparent position 13b is configured to be placed along the line 15 and between the loudspeakers 11a and 11b (denoted by the line 14). In other words, loudspeakers 11a and 11b may render audio such that the apparent position 13b of the audio object 13a is configured to fall along the line 15. As such, the user 12 may perceive the audio to be generated from a direction of the audio object 13a. In one example, a determination of which loudspeakers (e.g. loudspeaker pair) can be selected for allowing the apparent position 13b to be behind the position of the audio object 13a, and thus the loudspeakers 11a and 11b are selected for rendering the audio on that basis. The audio may be panned based on movement of the user 12 such that the apparent position 13b may move along the line 14 so that the perception of the audio being rendered from the direction of the audio object 13a.
  • FIG. 2 is a block diagram of an example system, indicated generally by the reference numeral 20. The system 20 shows the VR scene of FIG. 1, with the position of the user 12 being moved from position 12a to position 12b, and line 16 representing a direction of the audio object 13a relative to the user position 12b. As such, the audio object 13a has an updated apparent position 13c along the line 16 based on the updated position 12b of the user 12. The apparent position 13c is configured to be placed along the line 16 and between the loudspeakers 11a and 11e (denoted by the line 17). In other words, loudspeakers 11a and 11e may render audio such that the apparent position 13c of the audio object 13a is configured to fall along the line 16. As such, the user 12 may perceive the audio to be generated from a direction of the audio object 13a.
  • FIG. 3 is a block diagram of an example system, indicated generally by the reference numeral 30. The system 30 shows the VR scene of FIG. 1, with the position of the user 12 being moved from position 12a to position 12c, and line 15 representing a direction of the audio object 13a relative to the user position 12c. As such, the user 12 is shown to move closer to the audio object 13a along the line 15. As the user 12 moves closer to the audio object 13a, the gain of the loudspeakers 11a and 11b may be adjusted so as to simulate distance gain effects. For example, the audio may be rendered by the loudspeakers 11a and 11b at the position 13d at a higher volume (denoted by a larger circle representation) to that shown in FIG. 1.
  • FIG. 4 is a block diagram of an example system, indicated generally by the reference numeral 40. The system 40 shows a VR scene similar to that of FIG. 1, with the loudspeakers 11a-11e arranged in a similar manner. The system 40 further shows an audio object 41a, such that the direction, shown by the line 42, of the audio object 41a relative to the user 12 matches a direction of the loudspeaker 11a relative to the user 12. As such, the apparent position 41b is placed behind the audio object 41a along the line 42 such that the apparent position 41b may fall on the position of the loudspeaker 11a. Therefore, the audio for the audio object 41a may be rendered from a single loudspeaker 11a. In some examples, if an audio object is placed at a same position as a loudspeaker, the apparent position of the audio object need not be changed based on the user's movement, as the audio from the audio object may appear to be generated from the direction of the respective loudspeaker regardless of the change in user position. However, the gain of the loudspeakers may be adjusted so as to simulate distance gain effects.
  • FIG. 5 is a block diagram of an example system, indicated generally by the reference numeral 50. The system 50 shows a VR scene similar to that of FIG. 1, with another user 51 and an audio object 52a. In this scenario, it may not be feasible to render audio for the audio object 52a based on positions of both the user 12 and user 51, as their positions are different, and they would perceive the audio coming from different relative directions. For example, if the position of user 12 is taken into consideration, line 54 represents the direction of the audio object 52a relative to the user 12. As described with reference to FIG. 1, an apparent position 52b of the audio object 52a is configured to appear from behind the audio object 52a along the line 54, thus the audio being rendered by the loudspeakers 11a and 11b and the audio may be panned along the line 55 based on movement of user 12. However, if the apparent position of the audio object is at the apparent position 52b, the user 51 may perceive the audio object 52a to be placed in an incorrect position. Line 56 represents the direction of the apparent position 52b relative to the user 51. For example, when audio of the audio object 52a appears to be generated from the apparent position 52b, the user 51 may perceive the audio object to be located along line 56, for example, at position 53. This may cause the user 51 to perceive a mismatch between the audio direction and a position of a visual representation corresponding to the audio object 52a (e.g. a visual object placed at the same position as the audio object 52a). The mismatch may be more noticeable if the user 51 is stationary while the user 12 moves, as the movement of user 12 may cause the apparent position 52b to be moved as well, thus causing the audio object position 53 to move around for the user 51 even though user 51 is not moving. Such a mismatch may not be desirable and may interfere in the user 51 having an immersive experience.
  • The example embodiments described below may aim to address problems arising from a situation as described in FIG. 5.
  • FIG. 6 is a flowchart of an algorithm, indicated generally by the reference numeral 60, in accordance with an example embodiment. FIG. 6 may be viewed in conjunction with FIG. 7 for better understanding.
  • FIG. 7 is a block diagram of a system, indicated generally by the reference numeral 70, in accordance with an example embodiment. The system 70 shows a scene 71 with a single user 12, and a scene 72 with a plurality of users (users 12 and 78). The scene 71 may be similar to the scene as described with reference to FIG. 1.
  • The algorithm 60 describes a method for rendering audio for a three-dimensional scene, such as a VR scene, at one or more of a plurality of loudspeakers 11a-11b. The plurality of loudspeakers is arranged in a first loudspeaker arrangement (e.g. shown by the loudspeakers 11a-11e) suitable for providing spatial audio corresponding to the three-dimensional scene. The three-dimensional scene may comprise a first audio object 73a.
  • The algorithm 60 may start at operation 601 for determining a number of users consuming the three-dimensional scene.
  • If it is determined, at operation 602, that more than one user, such as user 12 and user 78, is consuming the scene (e.g. scene 72), the algorithm moves to operation 604, and the apparent position 76 of the first audio object 73a is configured to be mapped, as shown by the arrow 77, to a position of a first loudspeaker, such as the loudspeaker 11a of the plurality of loudspeakers. Configuring the apparent position may comprise rendering audio corresponding to the first audio object 73a from the loudspeaker 11a. The rendering may comprise rendering the audio independent of changes in positions of the users 12 and/or 78.
  • If it is determined at operation 602, that a single user, such as user 12, is consuming the scene, the algorithm may move to the operation 603, and the apparent position (e.g. 73b) of the first audio object 73a is configured to be determined based on an initial position of the audio object (e.g. position of the audio object 73a) and/or relative position of the user (e.g. position of user 12). For example, audio may be rendered from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object 73a and a position of the first user 12, and optionally based on the positioning of one or more of the plurality of loudspeakers 11a-11e. For example, line 74 represents direction of the audio object 73a relative to the user position 12, and therefore the apparent position 73b is configured to be placed along the line 74 and between the loudspeakers 11a and 11b (denoted by the line 75). In other words, loudspeakers 11a and 11b may render audio such that the apparent position 73b of the audio object 13a is configured to fall along the line 75. As such, the user 12 may perceive the audio to be generated from a direction of the audio object 73a. This may be similar to the audio rendering described with reference to FIG. 1. In one example, rendering audio in this scenario may further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user 12 (e.g. as shown with reference to FIGs. 2 and 3).
  • Moving to scene 72, as the apparent position 76 is configured to be mapped to a position of the loudspeaker 11a, both the users 12 and 78 may perceive the audio in a similar manner as being rendered from the position of the loudspeaker 11a, thus preventing a situation of audio and visual mismatch as described with reference to FIG. 5. Further, the apparent position 76 may remain unchanged even if one of the users move, or the users move in different directions, as the apparent position 76 is not determined dependent on the position of the users.
  • In one example, one or more of a position, orientation, and/or size of a visual representation of the audio object 73a may be configured to be based at least in part on the apparent position 76 of the respective audio object 73a. Thus, a visual representation of the audio object 73a is also configured to appear at the apparent position 76 in order to improve consistency between the positions of the audio and visual objects.
  • In one example, when a single user 12 is consuming a scene 71, and a second user 78 joins in, the apparent position 73b may gradually change to the apparent position 76 such that the user 12 may perceive the change to be gradual rather than an abrupt change.
  • In an example embodiment, the first loudspeaker, such as the loudspeaker 11a may be selected based, at least in part, on which of the plurality of loudspeakers have the least distance from the first audio object. For example, the loudspeaker 11a is selected from the plurality of loudspeakers as the loudspeaker 11a may be closest to the audio object 73a. The distance may relate to one or more of horizontal distance between the loudspeaker(s) and the audio object 73a, vertical distance between the loudspeaker(s) and the audio object 73a, or angular distance between the loudspeaker(s) and the audio object 73a.
  • For example, when the vertical distance is taken into consideration, a loudspeaker (e.g. loudspeaker 11a) may be selected whose distance from the floor matches most to the elevation of the audio object (e.g. y coordinate in MPEG-I Audio). This way, as the VR scene is moved, the floor of the VR scene may match the real-life room floor as well as possible. In another example, if matching of the loudspeaker elevation and audio object elevation is not feasible, the VR scene may be modified by adjusting the height of the audio object (and any associated visual object) such that when placed at the loudspeaker position, the VR scene floor and the real-life room floor positions match.
  • In an example embodiment, any changes in audio rendering based on the user's relative position (e.g. the changes in position/volume described with reference to FIGs. 1 to 3, such as loudspeaker directivity compensation) may be paused and/or cancelled when more than one user is experiencing the VR scene.
  • In one example, one or more of a position, orientation, and/or size of a visual representation of any audio object may be configured to be based at least in part on the apparent position of the respective audio object.
  • FIG. 8 is a flowchart of an algorithm, indicated generally by the reference numeral 80, in accordance with an example embodiment. FIG. 8 may be viewed in conjunction with FIG. 9 for better understanding.
  • FIG. 9 is a block diagram of a system, indicated generally by the reference numeral 90, in accordance with an example embodiment. The system 90 shows a three-dimensional scene (e.g. scene 91) with the plurality of loudspeakers 11a-11e, and two audio objects 92a (similar to the first audio object 73a) and 93a. The audio may be rendered such that the scene is being consumed by more than one user, such as users 12 and 78 (the users are not depicted here for simplicity). Scenes 97, 98, and 99 show how the scene 91 may be modified according to an example embodiment.
  • The algorithm 80 may be performed after operation 604 of algorithm 60, where an apparent position of the first audio object (73a, 92a) is mapped to a position of a first loudspeaker of the plurality of loudspeakers. The algorithm 80 may start at operation 801, where it is determined that the three-dimensional scene comprises more than one audio objects, such as a second audio object 93a, in addition to a first audio object 92a. If it is determined that the scene comprises more than one audio object, the algorithm may move to operation 802.
  • At operation 802, an apparent position of the second audio object 93a may be configured to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  • In one example, the mapping may comprise performing one or more of shifting, rotating and/or scaling the three-dimensional scene (e.g. one or more audio objects of the three-dimensional scene). For example, the scene may be shifted as shown by the arrow 94, such that the audio object 92a is mapped on the position of the loudspeaker 11a (e.g. similar to the audio object 73a being mapped on the loudspeaker 11a). The new positions of the audio objects are shown by the apparent positions 92b and 93b in scene 97. Next, as shown in scene 97, the scene may be rotated as shown by the arrow 95, such that the position of the audio objects may be updated to the apparent positions 92c and 93c as shown in the scene 98. Next, the scene may be panned as shown by the arrow 96, such that the position of the audio objects may be updated to the apparent positions 92d and 93d as shown in the scene 99. Thus, the audio objects 92a and 93a are mapped on the positions of loudspeakers 11a and 11b respectively. As such, plurality of users may perceive the audio from audio object 92a in a similar manner as being rendered from the position of the loudspeaker 11a and the audio from the audio object 93a in a similar manner as being rendered from the position of the loudspeaker 11b, thus preventing a situation of audio and visual mismatch as described with reference to FIG. 5. Further, the apparent positions 92d and 93d may remain unchanged even if one of the users move, or the users move in different directions or different distances, as the apparent positions 92d and/or 93d are not determined dependent on the position of the users. In one example, one or more of a position, orientation, and/or size of a visual representation of the audio objects 92a and 93a may be configured to be based at least in part on the apparent position 92d and 93d respectively. Thus, a visual representation of the audio object 92a is configured to appear at the apparent position 92d, and a visual representation of the audio object 93a is configured to appear at the apparent position 93d in order to improve consistency between the positions of the audio and visual objects.
  • In an example embodiment, the loudspeaker 11b is selected for rendering the audio from the audio object 93a based on the loudspeaker 11b being a nearest one (e.g. least horizontal, vertical, and/or angular distance) of the plurality of loudspeakers in relation to the second audio object 93a.
  • FIG. 10 is a flowchart of an algorithm, indicated generally by the reference numeral 100, in accordance with an example embodiment. FIG. 10 may be viewed in conjunction with FIG. 11 for better understanding.
  • FIG. 11 is a block diagram of a system, indicated generally by the reference numeral 110, in accordance with an example embodiment. The system 110 shows a three-dimensional scene (e.g. scene 111a) with the plurality of loudspeakers 11a-11e, and a plurality of audio objects 112a, 113a 114a, 115a, 116a, and 117a. The audio may be rendered such that the scene is being consumed by more than one user, such as users 12 and 78 (the users are not depicted here for simplicity). Scenes 111b-111d show how the scene 111a may be modified according to an example embodiment.
  • The algorithm 100 may start at operation 1001, where it is determined that the three-dimensional scene comprises more than a threshold number of audio objects, such as audio objects 112a- 117a. In one example, the threshold number may be two or three. In another example, the threshold may be dependent on the number of loudspeakers. For example, the threshold number may be half the number of loudspeakers (e.g. half the number rounded up or down to the nearest integer), the threshold number may be equal to number of speakers, the threshold number may be twice the number of loudspeakers, the threshold number may be the number of loudspeakers plus or minus a predefined number, or the like.
  • If it is determined at operation 1001 that the scene comprises more than a threshold number of audio object, the algorithm may move to operation 1002. In one example, the threshold number may be two, such that when there are more than two audio objects, operation 1002 may be performed. At operation 1002, apparent positions of at least one of the audio objects is configured to be mapped outside of the first loudspeaker arrangement. For example, the threshold number (e.g. two) of the audio objects may be mapped on respective loudspeakers, and the remaining audio objects in addition to the threshold number may be mapped outside of the first loudspeaker arrangement.
  • For example, the scene 111b shows how the audio objects 112a to 117a may be initially positioned relative to each other (e.g. audio object constellation shown using dotted lines between the positions 112b to 117b). As the number of audio objects 112a to 117a may be more than a threshold number (e.g. 2), the operation 1002 may be performed accordingly. The operation 1002 may be performed, for example by shifting, panning, and/or zooming one or more audio objects of the scene. The scene 111c shows audio objects 112a (at position 112c) and 116a (at position 116c) being mapped on the loudspeakers 11b and 11c respectively, and also shows audio objects 113a, 114a, 115a, and 117a (at positions 113c, 114c, 115c, and 117c respectively) to be mapped outside of the loudspeaker arrangement of loudspeakers 11a-11e. In an example embodiment, the rendering audio corresponding to audio objects mapped outside the first loudspeaker arrangement may comprise rendering said audio with spatial extent.
  • In one example, the loudspeakers 11b and 11c may be selected for the audio objects 112a (at position 112d) and 116a (at position 116d) respectively based on distance (e.g. horizontal, vertical, and/or angular) of the respective loudspeaker from the audio object, as described with reference to FIG. 9.
  • In one example, the audio objects 112a and 116a may be selected for being mapped on a loudspeaker (11b, 11c respectively) as the two audio objects being located next to each other at the edge of the audio object constellation (e.g. based on triangulation of the audio object positions. The scene may then be shifted, rotated, and scaled so that these audio objects 112d and 116d align with two loudspeakers 11b and 11c. For the audio objects, such as audio objects 113c, 114c, 115c, and 117c, that do not match the position of a loudspeaker, audio processing may be modified. For example, for spatial extent may be applied to audio objects 113c, 114c, 115c, and 117c such that the user may perceive the rendered audio to be arriving from a generic direction of the loudspeakers 11b and 11c (e.g. such that any mismatch between the audio and visual objects may be less noticeable. In one example, spatially extended sources may be rendered via creating multiple uncorrelated N "helper" audio objects from the original audio object and placing them around the original audio object.
  • The example embodiments above provide a system for loudspeaker rendering of 6DoF VR content to multiple users by modifying the VR scene based on the loudspeaker positions, and independent of the positions of each individual user, such that any mismatch between user perception of audio objects and corresponding visual objects may be minimized. This may be achieved by configuring apparent positions of one or more audio objects to be mapped on positions of the loudspeaker(s) respectively.
  • For completeness, FIG. 12 is a schematic diagram of components of one or more of the example embodiments described previously, which hereafter are referred to generically as processing systems 300. A processing system 300 may have a processor 302, a memory 304 closely coupled to the processor and comprised of a RAM 314 and ROM 312, and, optionally, user input 310 and a display 318. The processing system 300 may comprise one or more network/apparatus interfaces 308 for connection to a network/apparatus, e.g. a modem which may be wired or wireless. Interface 308 may also operate as a connection to other apparatus such as device/apparatus which is not network side apparatus. Thus, direct connection between devices/apparatus without network participation is possible.
  • The processor 302 is connected to each of the other components in order to control operation thereof.
  • The memory 304 may comprise a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD). The ROM 312 of the memory 304 stores, amongst other things, an operating system 315 and may store software applications 316. The RAM 314 of the memory 304 is used by the processor 302 for the temporary storage of data. The operating system 315 may contain computer program code which, when executed by the processor implements aspects of the algorithms 60, 80, 90 described above. Note that in the case of small device/apparatus the memory can be most suitable for small size usage i.e. not always hard disk drive (HDD) or solid-state drive (SSD) is used.
  • The processor 302 may take any suitable form. For instance, it may be a microcontroller, a plurality of microcontrollers, a processor, or a plurality of processors.
  • The processing system 300 may be a standalone computer, a server, a console, or a network thereof. The processing system 300 and needed structural parts may be all inside device/apparatus such as IoT device/apparatus i.e. embedded to very small size.
  • In some example embodiments, the processing system 300 may also be associated with external software applications. These may be applications stored on a remote server device/apparatus and may run partly or exclusively on the remote server device/apparatus. These applications may be termed cloud-hosted applications. The processing system 300 may be in communication with the remote server device/apparatus in order to utilize the software application stored there.
  • FIG. 13 shows tangible media, specifically a removable memory unit 365, storing computer-readable code which when run by a computer may perform methods according to example embodiments described above. The removable memory unit 365 may be a memory stick, e.g. a USB memory stick, having internal memory 366 for storing the computer-readable code. The internal memory 366 may be accessed by a computer system via a connector 367. Other forms of tangible storage media may be used. Tangible media can be any device/apparatus capable of storing data/information which data/information can be exchanged between devices/apparatus/network.
  • Embodiments of the present invention may be implemented in software, hardware, application logic or a combination of software, hardware and application logic. The software, application logic and/or hardware may reside on memory, or any computer media. In an example embodiment, the application logic, software or an instruction set is maintained on any one of various conventional computer-readable media. In the context of this document, a "memory" or "computer-readable medium" may be any non-transitory media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer.
  • Reference to, where relevant, "computer-readable storage medium", "computer program product", "tangibly embodied computer program" etc., or a "processor" or "processing circuitry" etc. should be understood to encompass not only computers having differing architectures such as single/multi-processor architectures and sequencers/parallel architectures, but also specialised circuits such as field programmable gate arrays FPGA, application specify circuits ASIC, signal processing devices/apparatus and other devices/apparatus. References to computer program, instructions, code etc. should be understood to express software for a programmable processor firmware such as the programmable content of a hardware device/apparatus as instructions for a processor or configured or configuration settings for a fixed function device/apparatus, gate array, programmable logic device/apparatus, etc.
  • As used in this application, the term "circuitry" refers to all of the following: (a) hardware-only circuit implementations (such as implementations in only analogue and/or digital circuitry) and (b) to combinations of circuits and software (and/or firmware), such as (as applicable): (i) to a combination of processor(s) or (ii) to portions of processor(s)/software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a server, to perform various functions) and (c) to circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present.
  • If desired, the different functions discussed herein may be performed in a different order and/or concurrently with each other. Furthermore, if desired, one or more of the above-described functions may be optional or may be combined. Similarly, it will also be appreciated that the flow charts of Figures 6, 8, and 10 are examples only and that various operations depicted therein may be omitted, reordered and/or combined.
  • It will be appreciated that the above described example embodiments are purely illustrative and are not limiting on the scope of the invention. Other variations and modifications will be apparent to persons skilled in the art upon reading the present specification.
  • Moreover, the disclosure of the present application should be understood to include any novel features or any novel combination of features either explicitly or implicitly disclosed herein or any generalization thereof and during the prosecution of the present application or of any application derived therefrom, new claims may be formulated to cover any such features and/or combination of such features.

Claims (15)

  1. An apparatus comprising:
    means for rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises:
    means for determining a number of users consuming the three-dimensional scene; and
    means for configuring, based on determining that more than one user is consuming the three-dimensional scene, an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  2. An apparatus as claimed in claim 1, wherein configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
  3. An apparatus as claimed in any one of the preceding claims, further comprising:
    means for determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object;
    wherein the means for rendering further comprises means for configuring, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, an apparent position of the second audio object to be mapped on to a position of a second loudspeaker of the plurality of loudspeaker.
  4. An apparatus as claimed in claim 3, wherein the means for configuring the apparent positions of the first and second audio objects to be mapped on to the positions of the first and second loudspeakers comprises means for performing one or more of shifting, rotating and/or scaling the three-dimensional scene.
  5. An apparatus as claimed in claim 4, wherein the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped on to the first and second loudspeakers.
  6. An apparatus as claimed in any one of the preceding claims, further comprising:
    means for determining whether the three-dimensional scene comprises more than a threshold number of audio objects;
    wherein if the three-dimensional scene comprises more than the threshold number of audio objects, the means for rendering comprises means for configuring apparent position(s) of at least one of the audio objects to be mapped outside of the first loudspeaker arrangement.
  7. An apparatus as claimed in claim 6, wherein the apparent positions of each of at least two of the audio objects is configured to be mapped on positions of each of at least two loudspeakers respectively.
  8. An apparatus as claimed in claim 6, wherein rendering audio corresponding to audio objects to be mapped outside of the first loudspeaker arrangement comprises rendering said audio with spatial extent.
  9. An apparatus as claimed in any one of the preceding claims, wherein the means for rendering comprises means for rendering the audio independent of changes in positions of the more than one user.
  10. An apparatus as claimed in any one of the preceding claims, further comprising means for configuring one or more of a position, orientation, and/or size of a visual representation of any audio object to be based at least in part on the apparent position of the respective audio object.
  11. An apparatus as claimed in any one of the preceding claims, wherein the means for rendering further comprises:
    means for rendering, based on determining that a single first user is consuming the three-dimensional scene, audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user.
  12. An apparatus as claimed in claim 11, wherein the means for rendering further comprises changing the direction and/or volume of the audio based at least in part on changes in the position of the first user.
  13. An apparatus as claimed in any one of the preceding claims, wherein the first loudspeaker is selected from the plurality of loudspeakers based on which of the plurality of loudspeakers have the least distance from the first audio object.
  14. A method comprising:
    rendering audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises:
    determining a number of users consuming the three-dimensional scene; and
    based on determining that more than one user is consuming the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
  15. A computer program comprising instructions, which, when executed by an apparatus, cause the apparatus to:
    render audio for a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises:
    determine a number of users consuming the three-dimensional scene; and
    based on determining that more than one user is consuming the three-dimensional scene, configure an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
EP25166950.3A 2024-04-23 2025-03-28 Rendering audio Pending EP4642055A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
GB2405692.1A GB2640535A (en) 2024-04-23 2024-04-23 Rendering audio

Publications (1)

Publication Number Publication Date
EP4642055A1 true EP4642055A1 (en) 2025-10-29

Family

ID=91275248

Family Applications (1)

Application Number Title Priority Date Filing Date
EP25166950.3A Pending EP4642055A1 (en) 2024-04-23 2025-03-28 Rendering audio

Country Status (3)

Country Link
EP (1) EP4642055A1 (en)
CN (1) CN120835264A (en)
GB (1) GB2640535A (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170245086A1 (en) * 2013-04-26 2017-08-24 Sony Corporation Sound processing apparatus and method, and program
US20190246229A1 (en) * 2018-02-06 2019-08-08 Sony Interactive Entertainment Inc Localization of sound in a speaker system
US20210168548A1 (en) * 2017-12-12 2021-06-03 Sony Corporation Signal processing device and method, and program
US20220167111A1 (en) * 2019-06-12 2022-05-26 Google Llc Three-dimensional audio source spatialization

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170245086A1 (en) * 2013-04-26 2017-08-24 Sony Corporation Sound processing apparatus and method, and program
US20210168548A1 (en) * 2017-12-12 2021-06-03 Sony Corporation Signal processing device and method, and program
US20190246229A1 (en) * 2018-02-06 2019-08-08 Sony Interactive Entertainment Inc Localization of sound in a speaker system
US20220167111A1 (en) * 2019-06-12 2022-05-26 Google Llc Three-dimensional audio source spatialization

Also Published As

Publication number Publication date
GB202405692D0 (en) 2024-06-05
CN120835264A (en) 2025-10-24
GB2640535A (en) 2025-10-29

Similar Documents

Publication Publication Date Title
CN111434126B (en) Signal processing device and method, and program
EP2804402B1 (en) Sound field control device, sound field control method and program
US11055057B2 (en) Apparatus and associated methods in the field of virtual reality
WO2020148120A2 (en) Processing audio signals
US10542366B1 (en) Speaker array behind a display screen
EP2992690A1 (en) Sound field adaptation based upon user tracking
CN111492342B (en) Audio scene processing
CN112071326B (en) Sound effect processing method and device
CN111512640B (en) multi-camera device
CN109462811B (en) Sound field reconstruction method, device, storage medium and device based on non-central point
GB2640535A (en) Rendering audio
US12126987B2 (en) Virtual scene
EP4380196A1 (en) Spatial sound improvement for seat audio using spatial sound zones
EP3595361A1 (en) Use of local link to support transmission of spatial audio in a virtual environment
CN117750270A (en) Spatial blending of audio
US12538091B2 (en) Methods, apparatus and systems for modelling audio objects with extent
CN118678286B (en) Audio data processing method, device and system, electronic equipment and storage medium
US9473871B1 (en) Systems and methods for audio management
US10200807B2 (en) Audio rendering in real time
HK40104983A (en) Methods, apparatus and systems for modelling audio objects with extent
HK40100828B (en) Methods, apparatus and systems for modelling audio objects with extent
HK40100828A (en) Methods, apparatus and systems for modelling audio objects with extent
CN117223299A (en) Methods, apparatus and systems for modeling audio objects with ranges

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN PUBLISHED

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE