EP4684538A1 - Rendering audio over multiple loudspeakers utilizing interaural cues for height virtualization - Google Patents
Rendering audio over multiple loudspeakers utilizing interaural cues for height virtualizationInfo
- Publication number
- EP4684538A1 EP4684538A1 EP24719406.1A EP24719406A EP4684538A1 EP 4684538 A1 EP4684538 A1 EP 4684538A1 EP 24719406 A EP24719406 A EP 24719406A EP 4684538 A1 EP4684538 A1 EP 4684538A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio
- loudspeakers
- loudspeaker
- processing method
- panning gains
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/01—Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]
Definitions
- the disclosure pertains to systems and methods for rendering audio for playback by a set of speakers.
- Audio devices including but not limited to smart audio devices, have been widely deployed and are becoming common features of many homes. Although existing systems and methods for controlling audio devices provide benefits, improved systems and methods would be desirable.
- loudspeaker and “loudspeaker” are used synonymously to denote any sound-emitting transducer (or set of transducers) driven by a single speaker feed.
- a typical set of headphones includes two speakers.
- performing an operation “on” a signal or data e.g., filtering, scaling, transforming, or applying gain to, the signal or data
- a signal or data e.g., filtering, scaling, transforming, or applying gain to, the signal or data
- performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
- system is used in a broad sense to denote a device, system, or subsystem.
- a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
- processor is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data).
- processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and/or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.
- the term “couples” or “coupled” is used to mean either a direct or indirect connection. Thus, if a first device couples to a second device, that connection may be through a direct connection, or through an indirect connection via other devices and connections.
- a single purpose audio device is a device (e.g., a TV or a mobile phone) including or coupled to at least one microphone (and optionally also including or coupled to at least one speaker) and which is designed largely or primarily to achieve a single purpose.
- a TV typically can play (and is thought of as being capable of playing) audio from program material, in most instances a modem TV runs some operating system on which applications run locally, including the application of watching television.
- the audio input and output in a mobile phone may do many things, but these are serviced by the applications running on the phone.
- a single purpose audio device having speaker(s) and microphone(s) is often configured to run a local application and/or service to use the speaker(s) and microphone(s) directly.
- Some single purpose audio devices may be configured to group together to achieve playing of audio over a zone or user configured area.
- a virtual assistant e.g., a connected virtual assistant
- a device e.g., a smart speaker or voice assistant integrated device
- at least one microphone and optionally also including or coupled to at least one speaker
- Virtual assistants may sometimes work together, e.g., in a discrete and conditionally defined way. For example, two or more virtual assistants may work together in the sense that one of them, for example, the one which is most confident that it has heard a wakeword, responds to the word.
- the connected devices may form a sort of constellation, which may be managed by one main application which may be (or implement) a virtual assistant.
- “wakeword” is used in a broad sense to denote any sound (e.g., a word uttered by a human, or some other sound), where a smart audio device is configured to awake in response to detection of (“hearing”) the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone).
- to “awake” denotes that the device enters a state in which it awaits (i.e., is listening for) a sound command.
- a “wakeword” may include more than one word, e.g., a phrase.
- wakeword detector denotes a device configured (or software that includes instructions for configuring a device) to search continuously for alignment between real-time sound (e.g., speech) features and a trained model.
- a wakeword event is triggered whenever it is determined by a wakeword detector that the probability that a wakeword has been detected exceeds a predefined threshold.
- the threshold may be a predetermined threshold which is tuned to give a good compromise between rates of false acceptance and false rejection.
- a device Following a wakeword event, a device might enter a state (which may be referred to as an “awakened” state or a state of “attentiveness”) in which it listens for a command and passes on a received command to a larger, more computationally-intensive recognizer.
- a wakeword event a state in which it listens for a command and passes on a received command to a larger, more computationally-intensive recognizer.
- At least some aspects of the present disclosure may be implemented via methods, such as audio processing methods.
- the methods may be implemented, at least in part, by a control system such as those disclosed herein.
- Some methods involve receiving, by a control system, audio data.
- the audio data may include a set of audio signals and associated spatial data.
- the set of audio signals may include one or more audio signals.
- the spatial data may indicate a desired perceived spatial position in three dimensions corresponding to an audio signal.
- Some methods may involve obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers with respect to a desired listening position or area.
- the set of loudspeakers may include two or more loudspeakers.
- Some methods may involve computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers.
- the relative activation of loudspeakers may be computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions.
- the relative activation of loudspeakers may be computed via combinations of two or more sets of panning gains constructed via amplitude panning. Each set of panning gains may approximate one or more interaural cues of the desired audio signal spatial position.
- the interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD).
- ITD interaural time difference
- ILD interaural level difference
- Some methods may involve rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers.
- Some methods may involve providing the set of audio signals to the set of loudspeakers.
- the relative activation of loudspeakers may vary continuously as a function of the desired perceived spatial position of the audio signals.
- the desired perceived spatial position corresponding to at least one audio signal may be at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located.
- computing the relative activation of loudspeakers may involve determining a cone of confusion corresponding to the audio signal elevation.
- computing the relative activation of loudspeakers may involve determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection.
- the panning gains may be power-preserving panning gains.
- the rendering may involve combining the panning gains for each of the points of intersection.
- the panning gains for each of the points of intersection may be determined via a flexible rendering process.
- the flexible rendering process may be a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof.
- the panning gains for each of the points of intersection may be determined by applying a cost function that is based on a sum of a spatial term and a proximity term.
- the panning gains may be combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual speakers located at each of the points of intersection. The additional panning gains may be determined using any of the disclosed methods.
- the intended perceived spatial position may correspond with a channel of a channel-based audio format. According to some examples, the intended perceived spatial position may be indicated by audio object metadata.
- Some or all of the operations, functions and/or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media.
- Such non-transitory media may include one or more memory devices such as those described herein, including but not limited to one or more random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon.
- an apparatus may include an interface system and a control system.
- the control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
- DSPs digital signal processors
- ASICs application specific integrated circuits
- FPGAs field programmable gate arrays
- Figure 1A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure.
- Figure IB shows examples of a listener and an array of loudspeakers in an audio environment.
- Figure 2 shows the audio environment of Figure IB from a different perspective.
- Figure 3 shows an example of a cone of confusion.
- Figure 4 shows panning gains generated for the loudspeaker layout shown in Figure 1.
- Figure 5 shows panning gains generated for the loudspeaker layout shown in Figure 1 according to some examples of the present disclosure.
- Figure 6 shows a floor plan of a listening environment, which is a living space in this example.
- Figure 7 is a graph of points indicating speaker activations, in an example embodiment.
- Figure 8 is a graph of tri-linear interpolation between points indicative of speaker activations according to one example.
- Figure 9 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein.
- Amplitude panning is a common method for rendering spatial audio to an array of loudspeakers.
- Amplitude panning involves varying the amplitudes of panning gains for a given audio signal as a function of the desired spatial position of the signal such that the perceived position of the sound energy aligns with the target signal position.
- the desired audio signal position lies between two speakers, this can be accomplished with pairwise panning, meaning only the two speakers on either side of the desired audio signal position participate in its rendering. (When the desired signal position aligns with the physical location of a loudspeaker, only that loudspeaker participates in its rendering.)
- Amplitude panning is an apt method for rendering to a myriad of speaker layouts.
- Playback of spatial audio in a consumer environment has typically been tied to a prescribed number of loudspeakers placed in prescribed positions, such as positions corresponding to Dolby 5.1 or 7.1 surround sound.
- content is authored specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby DigitalTM, Dolby Digital PlusTM, etc.)
- immersive, object-based spatial audio formats have been introduced (such as Dolby AtmosTM) which break this association between the content and specific loudspeaker locations.
- the content may be described as a collection of individual audio objects, each with possibly time varying metadata describing the desired perceived location of said audio objects in three- dimensional space and, in some examples, other properties of the audio object.
- the audio content is transformed into loudspeaker feeds by a Tenderer which adapts to the number and location of loudspeakers in the playback system.
- Such rendering methods may be improved upon by considering known perceptual cues that indicate elevated or depressed audio objects as above or below the head to the human ears.
- Interaural cues are critical for sound localization, including sound localization for elevated or depressed audio objects. These interaural cues are functions of the position of an audio object with respect to the head and refer to the resulting differences in some qualities of the sound when it reaches each ear.
- Two such cues are the interaural time difference (ITD) and the interaural level difference (ILD).
- ITD interaural time difference
- ILD interaural level difference
- the energy from a sound source directly in front of a listener produces interaural time and level differences of zero, because the source is equidistant from each ear.
- the energy from a sound source directly to the left of a listener produces an ITD equal to the time it takes the energy to travel past the listener’s left ear and around the listener’s head to the listener’s right ear, and an ILD equal to the air attenuation over the distance from left ear around the head to the right ear, plus any absorption or damping contributed by the head.
- ITD the time it takes the energy to travel past the listener’s left ear and around the listener’s head to the listener’s right ear
- ILD equal to the air attenuation over the distance from left ear around the head to the right ear, plus any absorption or damping contributed by the head.
- the maximum interaural differences a sound source would produce around a person’s head occurs for a source directly to the person’s left or right.
- Elevated sound sources may produce much smaller interaural differences than their head-height counterparts, so loudspeakers in such positions with high interaural differences are unsuitable for rendering elevated objects alone.
- amplitude panners endeavor to reproduce the spatial effect of a loudspeaker placed at the object location (in the loudspeaker plane), it is unsuitable to render height objects with traditional amplitude panning gains for such locations.
- the present disclosure presents solutions to these challenges.
- Figure 1A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 1A are merely provided by way of example. Other implementations may include more, fewer and/or different types and numbers of elements. According to some examples, the apparatus 100 may be, or may include, a smart audio device that is configured for performing at least some of the methods disclosed herein.
- the apparatus 100 may be, or may include, another device that is configured for performing at least some of the methods disclosed herein, such as a laptop computer, a cellular telephone, a tablet device, a smart home hub, etc. In some such implementations the apparatus 100 may be, or may include, a server.
- the apparatus 100 includes an interface system 105 and a control system 110.
- the interface system 105 may, in some implementations, be configured for receiving audio data.
- the audio data may include audio signals that are to be reproduced by at least some speakers of an environment.
- the audio data may include one or more audio signals and associated spatial data.
- the spatial data corresponding to an audio signal may indicate the intended perceived spatial position of that audio signal.
- the spatial data may be, or may include, audio object metadata.
- the spatial data may correspond with a channel of a channel-based audio format.
- the intended perceived spatial position may be derived from the audio format, such as with higher-order Ambisonics (HO A) or other spherical harmonic based sound field representations.
- HO A Ambisonics
- the interface system 105 may be configured for providing rendered audio signals to at least some loudspeakers of the set of loudspeakers of the environment.
- the interface system 105 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
- the control system 110 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and/or discrete hardware components.
- DSP digital signal processor
- ASIC application specific integrated circuit
- FPGA field programmable gate array
- control system 110 may reside in more than one device.
- a portion of the control system 110 may reside in a device within one of the environments depicted herein and another portion of the control system 110 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc.
- a portion of the control system 110 may reside in a device within one of the environments depicted herein and another portion of the control system 110 may reside in one or more other devices of the environment.
- control system functionality may be distributed across multiple smart audio devices of an environment, or may be shared by an orchestrating device (such as what may be referred to herein as a smart home hub) and one or more other devices of the environment.
- the interface system 105 also may, in some such examples, reside in more than one device.
- the control system 110 may be configured for performing, at least in part, the methods disclosed herein.
- the control system 110 may be configured for receiving audio data including a set of audio signals and associated spatial data.
- the set of audio signals may include one or more audio signals and the spatial data may indicate a desired perceived spatial position in three dimensions corresponding to an audio signal.
- the control system 110 may be configured for obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area.
- control system 110 may be configured for computing a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers.
- control system 110 may be configured for computing the relative activation of loudspeakers as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions.
- the control system 110 may be configured for computing the relative activation of loudspeakers via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial position.
- the interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD).
- Each set of panning gains may, for example, “approximate” one or more interaural cues of the desired audio signal spatial position by providing audio signals audio signals within a time difference threshold and/or level difference threshold — as compared to the exact ITD and/or ILD of the desired audio signal spatial position — for a particular perceived spatial position of the audio signals, such as within 1%, within 3%, within 5%, within 8%, within 10%, or within another threshold.
- the control system 110 may be configured for rendering the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers.
- the control system 110 may be configured for providing the set of audio signals to the set of loudspeakers.
- Non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc.
- RAM random access memory
- ROM read-only memory
- the one or more non-transitory media may, for example, reside in the optional memory system 115 shown in Figure 1A and/or in the control system 110. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon.
- the software may, for example, include instructions for controlling at least one device to process audio data.
- the software may, for example, be executable by one or more components of a control system such as the control system 110 of Figure 1A.
- the apparatus 100 may include the optional microphone system 120 shown in Figure 1A.
- the optional microphone system 120 may include one or more microphones.
- one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc.
- the apparatus 100 may include the optional loudspeaker system 125 shown in Figure 1A.
- the optional loudspeaker system 125 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as “speakers.”
- at least some loudspeakers of the optional loudspeaker system 125 may be arbitrarily located .
- at least some speakers of the optional loudspeaker system 125 may be placed in locations that do not correspond to any standard prescribed speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc.
- At least some loudspeakers of the optional loudspeaker system 125 may be placed in locations that are convenient to the space (e.g., in locations where there is space to accommodate the loudspeakers), but not in any standard prescribed loudspeaker layout.
- the apparatus 100 may include the optional sensor system 130 shown in Figure 1A.
- the optional sensor system 130 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc.
- the optional sensor system 130 may include one or more cameras.
- the cameras may be free-standing cameras.
- one or more cameras of the optional sensor system 130 may reside in a smart audio device, which may be a single purpose audio device or a virtual assistant.
- one or more cameras of the optional sensor system 130 may reside in a TV, a mobile phone or a smart speaker.
- the apparatus 100 may include the optional display system 135 shown in Figure 1A.
- the optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays.
- the optional display system 135 may include one or more organic light-emitting diode (OLED) displays.
- the sensor system 130 may include a touch sensor system and/or a gesture sensor system proximate one or more displays of the display system 135.
- the control system 110 may be configured for controlling the display system 135 to present a graphical user interface (GUI), such as one of the GUIs disclosed herein.
- GUI graphical user interface
- the apparatus 100 may be, or may include, a smart audio device.
- the apparatus 100 may be, or may include, a wakeword detector.
- the apparatus 100 may be, or may include, a virtual assistant.
- Cones of confusion are three-dimensional (3D) surfaces around a listener’ s head on which sound sources will impart equivalent interaural cues. (An example of a cone of confusion is shown in Figure 3 and will be described in detail below.) Cones of confusion extend to the left from a listener’s left ear and to the right from a listener’s right ear. The open end of the cone forms a circle that intersects the head-height plane at its diameter.
- Cones of confusion may be used to determine coordinates for head-height (or zero -elevation) sound sources that would produce the same interaural differences as a target elevated source.
- Each cone of confusion intersects a circle in the head-height plane having its center at the center of the listener’s head at two points.
- Some disclosed examples involve rendering a set of one or more audio signals, each with an associated desired perceived spatial position, over a set of two or more loudspeakers, wherein:
- the relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the one or more audio signals and the locations of the loudspeakers, wherein: o the spatial position of the audio signals is given in three dimensions; o speaker activation varies as a function of spatial position of the audio signals, particularly the elevation of the audio signal positions with respect to the loudspeaker elevations; and o speaker activation is determined via combinations of two or more sets of panning gains (constructed via amplitude panning), each set matching the interaural cues of the desired audio signal spatial position.
- Figure IB shows examples of a listener and an array of loudspeakers in an audio environment.
- the audio environment 150 includes an array of loudspeakers Si, S2, S3 and S4 at the height of a listener’s head 155. This height may be referred to herein as “head height.”
- the array of loudspeakers S1-S4 are being used to render the audio object o ( -.
- Figure IB shows a coordinate system having its center at the center of the listener’s head 155.
- Figure 2 shows the audio environment of Figure IB from a different perspective.
- the same set of loudspeakers may be used to render an audio object Oj with the same x and y coordinates, but now with nonzero height Zj.
- Figure 3 shows an example of a cone of confusion.
- cones of confusion are 3D surfaces around a listener’s head on which sound sources will impart equivalent interaural cues. Cones of confusion extend to the left from a listener’s left ear and to the right from a listener’s right ear.
- the origin of the coordinate system 305 is at the center of the listener’s head 155.
- the listener’s head 155 is facing in the positive direction of the y axis, the positive direction of the x axis extends from the left of the listener’s head 155 and the positive direction of the z axis extends through the top of the listener’s head 155.
- the open end of the cone of confusion 310 forms a circle 315 that intersects the circle 320 at height Zj, which includes position Oj.
- Many amplitude panners (including those operating according to ITU-R BS.2127) would produce the same panning gains gj for an object at position Oj as g t for an object at position o r
- the circle 315 also intersects the circle 325 in the x-y plane — which is also referred to herein as a “head-height plane” — at positions o' 7 y and o' j b .
- the subscripts /and b represent front and back, respectively. This means that an audio object positioned at either position o' jf or position o' j b would produce the same interaural differences as an audio object at position 6j.
- Panning gains g' jf(oj) for these two points using the amplitude panner, and new panning gains g ⁇ for Oj, which produce the correct interaural cues, may be constructed as a linear combination of the two sets of equivalent head-height gains, for example as follows:
- the two coefficients (o 7 ) may vary continuously as a function of the object position Oj and may also be related to each other for spatial continuity, for example as follows:
- Equation 7 ensures that the final set of panning gains gj is power preserving.
- CMAP center of mass amplitude panning
- the set ⁇ s denotes the positions of a set of M loudspeakers
- 6 denotes the desired perceived spatial position of the audio signal
- g denotes an M dimensional vector of speaker activations.
- each activation in the vector represents a gain per speaker.
- An optimal vector of activations may be found by minimizing the cost function across activations:
- C spatiai is derived from a model that places the perceived spatial position of an audio signal playing from a set of loudspeakers at the center of mass of those loudspeakers’ positions weighted by their associated activating gains g t (elements of the vector g):
- Equation (10) may then be manipulated into a spatial cost representing the squared error between the desired audio position and that produced by the activated loudspeakers:
- Equation (11) the spatial term of the cost function for CMAP defined in Equation (11) can be rearranged into a matrix quadratic as a function of speaker activations g:
- Equation (12) A represents an M x M square matrix, B represents a 1 x M vector, and C represents a scalar.
- the matrix A is of rank 2, and therefore when M > 2 there exist an infinite number of speaker activations g for which the spatial error term equals zero.
- C proximity removes this indeterminacy and results in a particular solution with perceptually beneficial properties in comparison to the other possible solutions.
- C proximity may be constructed such that activation of speakers whose position s t is distant from the desired audio signal position 6 is penalized more than activation of speakers whose position is close to the desired position. This construction yields an optimal set of speaker activations that is sparse, where only speakers in close proximity to the desired audio signal’s position are significantly activated, and practically results in a spatial reproduction of the audio signal that is perceptually more robust to listener movement around the set of speakers.
- C proximity may be defined as a distance- weighted sum of the absolute values squared of speaker activations. This may be represented compactly in matrix form as:
- the distance penalty function can take on many forms, but the following is a useful parameterization:
- Equation (13c) represents the Euclidean distance between the desired audio position and speaker position and a and ft represent tunable parameters.
- the parameter a indicates the global strength of the penalty;
- d Q corresponds to the spatial extent of the distance penalty (loudspeakers at a distance around d Q or futher away will be penalized), and P accounts for the abruptness of the onset of the penalty at distance d Q .
- Equation (15) may yield speaker activations that are negative in value.
- Equation (15) may be minimized subject to all activations remaining positive.
- Other flexible rendering examples may involve vector base amplitude panning (VBAP), the amplitude panner of ITU-R BS.2127, other flexible rendering techniques, or combinations thereof.
- FIG 4 shows panning gains generated for the loudspeaker layout shown in Figure 1 for an audio object with an elevation of 45 degrees above the head height and a varying azimuth angle indicated by the x-axis of the figure.
- the panning gains were generated via a CMAP-based process which essentially projects the position of the audio object into the head height plane when computing the panning gains.
- each speaker is soloed — meaning that there is zero gain for other loudspeakers — when the azimuth angle of the object matches that of the loudspeaker’s angular position, as a result of the enforced sparsity. This is not a desirable outcome because loudspeakers S2 and S4 are positioned directly to the right and left of the head, generating maximum interaural differences, impossible for an audio object at 45 degrees elevation.
- FIG 5 shows panning gains generated for the same loudspeaker layout shown in Figure 1 according to some examples of the present disclosure.
- loudspeakers S2 and S4 are never soloed because the overall panning gains are computed as interpolations between the two sets of panning gains corresponding to the two projection points of the object position along the cone of confusion into the head-height plane.
- neither of these two points corresponds to a case where loudspeakers S2 and S4 are soloed. Instead, each point generates interaural cues commensurate with an object position at that particular azimuth and 45 degrees elevation.
- Figure 6 shows a floor plan of a listening environment, which is a living space in this example.
- the types and numbers of elements shown in Figure 6 are merely provided by way of example. Other implementations may include more, fewer and/or different types and numbers of elements.
- the environment 600 includes a living room 610 at the upper left, a kitchen 615 at the lower center, and a bedroom 622 at the lower right.
- Boxes and circles distributed across the living space represent a set of loudspeakers 605a-605h, at least some of which may be smart speakers in some implementations, placed in locations convenient to the space, but not adhering to any standard prescribed layout (arbitrarily placed).
- flexible rendering of spatial audio may be rendered to the loudspeakers 605a-605h according to one or more disclosed embodiments.
- the environment 600 may include a smart home hub for implementing at least some of the disclosed methods.
- the smart home hub may include at least a portion of the above-described control system 110.
- a smart device (such as a smart speaker, a mobile phone, a smart television, a device used to implement a virtual assistant, etc.) may implement the smart home hub.
- the environment 600 includes cameras 61 la-61 le, which are distributed throughout the environment.
- one or more smart audio devices in the environment 600 also may include one or more cameras.
- the one or more smart audio devices may be single purpose audio devices or virtual assistants.
- one or more cameras of the optional sensor system 130 may reside in or on the television 630, in a mobile phone or in a smart speaker, such as one or more of the loudspeakers 605b, 605d, 605e or 605h.
- cameras 61 la-61 le are not shown in every depiction of the environment 600 presented in this disclosure, each of the environments 600 may nonetheless include one or more cameras in some implementations.
- Figure 7 is a graph of points indicating speaker activations, in an example embodiment.
- the x and y dimensions are sampled with 15 points and the z dimension is sampled with 5 points.
- Other implementations may include more samples or fewer samples.
- each point represents the M speaker activations for the CMAP solution.
- the process of successive linear interpolation includes interpolation of each pair of points in the top plane to determine first and second interpolated points 805a and 805b, interpolation of each pair of points in the bottom plane to determine third and fourth interpolated points 810a and 810b, interpolation of the first and second interpolated points 805a and 805b to determine a fifth interpolated point 815 in the top plane, interpolation of the third and fourth interpolated points 810a and 810b to determine a sixth interpolated point 820 in the bottom plane, and interpolation of the fifth and sixth interpolated points 815 and 820 to determine a seventh interpolated point 825 between the top and bottom planes.
- tri-linear interpolation is an effective interpolation method
- tri-linear interpolation is just one possible interpolation method that may be used in implementing aspects of the present disclosure, and that other examples may include other interpolation methods.
- Figure 9 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein.
- the blocks of method 900 like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of method 900 may be performed concurrently. Moreover, some implementations of method 900 may include more or fewer blocks than shown and/or described.
- the blocks of method 900 may be performed by one or more devices, which may be (or may include) a control system such as the control system 110 that is shown in Figure 1A and described above.
- block 905 involves receiving, by a control system, audio data.
- the audio data includes a set of one or more audio signals and associated spatial data.
- the spatial data indicates an intended perceived spatial position corresponding to an audio signal.
- the spatial data may be, or may include, spatial metadata of an object-based audio format such as Dolby AtmosTM.
- the intended perceived spatial position may be represented as Oj or as o ( -, as disclosed herein.
- the spatial data may be, or may correspond with, channels of a channelbased audio format such as a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4 or Dolby 9.1 format.
- block 910 involves obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area.
- block 910 may involve obtaining the loudspeaker location data and/or location data corresponding to the desired listening position or area from a data structure stored in a memory of, or accessible by, the control system.
- block 910 may involve determining the loudspeaker location data and/or location data corresponding to the desired listening position or area.
- the loudspeaker location data and/or location data corresponding to the desired listening position or area may be obtained through numerous mechanisms known in the art. According to some examples, loudspeaker location data and/or location data corresponding to the desired listening position or area may be specified according to a standard loudspeaker layout, such as a Dolby 5.1 loudspeaker layout. In some such examples, a user could, for example, provide input to a device — such as an audio/video receiver (AVR) indicating that that they have a set of loudspeakers in a Dolby 5.1 layout.
- AVR audio/video receiver
- the AVR may then assume a "canonical" Dolby 5.1 layout and apply one or more disclosed methods — such as the operations of blocks 915 and 920 — according to the corresponding loudspeaker layout.
- loudspeaker location data and location data corresponding to the desired listening position or area are fixed and can be physically measured, e.g. with a tape measure, or obtained from layout information such as computer assisted drafting CAD data.
- block 910 may involve a more adaptable approach that can automatically detect these loudspeaker and/or user locations and orientations through a one-time setup procedure or even dynamically across time.
- AES 133rd Convention September 2012
- numerous commercially available techniques for tracking both the position and orientation of a listener’s head in the context of spatial audio reproduction systems are presented.
- One particular example discussed is the Microsoft Kinect.
- U.S. Patent No. 10,779,084 entitled “Automatic Discovery and Localization of Speaker Locations in Surround Sound Systems,” which is hereby incorporated by reference, a system is described which can automatically locate the positions of loudspeakers and microphones in a listening environment by acoustically measuring the time-of-arrival (TOA) between each speaker and microphone.
- a listening position or area may be detected by placing and locating a microphone at a desired listening position (a microphone in a mobile phone held by the listener, for example), and an associated listening orientation may be defined by placing another microphone at a point in the viewing direction of the listener, e.g. at the TV.
- the listening orientation may be defined by locating a loudspeaker in the viewing direction, e.g. the loudspeakers on the TV.
- the listening orientation is inherently defined as the line connecting the detected listening position and the component of the reproduction system that includes the linear microphone array, such as a sound bar that is co-located with a television (placed directly above or below the television). Because the sound bar’s location is predictably placed directly above or below the video screen, the geometry of the measured distance and incident angle can be translated to an absolute position relative to any point in front of that reference sound bar location using simple trigonometric principles.
- the distance between a loudspeaker and a microphone of the linear microphone array can be estimated by playing a test signal and measuring the time of flight (TOF) between the emitting loudspeaker and the receiving microphone. The time delay of the direct component of a measured impulse response can be used for this purpose.
- TOF time of flight
- the impulse response between the loudspeaker and a microphone array element can be obtained by playing a test signal through the loudspeaker under analysis.
- a test signal For example, either a maximum length sequence (MLS) or a chirp signal (also known as logarithmic sine sweep) can be used as the test signal.
- the room impulse response can be obtained by calculating the circular cross-correlation between the captured signal and the MLS input.
- Fig. 2 of this reference shows an echoic impulse response obtained using a MLS input. This impulse response is said to be similar to a measurement taken in a typical office or living room.
- the delay of the direct component is used to estimate the distance between the loudspeaker and the microphone array element. For loudspeaker distance estimation, any loopback latency of the audio device used to playback the test signal should be computed and removed from the measured TOF estimate.
- Such acoustic mapping may sometimes be referred to as “continuous,” in the sense that the acoustic mapping may be continued after an initial set-up process and may be responsive to changing conditions in the audio environment, such as changing noise sources and/or levels, loudspeaker relocation, the deployment of additional loudspeakers, the relocation and/or re-orientation of one or more listeners, etc.
- Some disclosed methods involve generating calibration signals that are injected (e.g., mixed) into the audio content being rendered by audio devices in an audio environment.
- the calibration signals may be, or may include, acoustic direct sequence spread spectrum (DSSS) signals.
- DSSS acoustic direct sequence spread spectrum
- the calibration signals may be, or may include, other types of acoustic calibration signals, such as swept sinusoidal acoustic signals, white noise, “colored noise,” such as pink noise (a spectrum of frequencies that decreases in intensity at a rate of three decibels per octave), acoustic signals corresponding to music, etc.
- acoustic calibration signals such as swept sinusoidal acoustic signals, white noise, “colored noise,” such as pink noise (a spectrum of frequencies that decreases in intensity at a rate of three decibels per octave), acoustic signals corresponding to music, etc.
- Some disclosed methods for estimating an audio device location in an environment involve obtaining direction of arrival (DOA) data for each audio device of a plurality of audio devices in the environment and determining interior angles for each of a plurality of triangles based on the DOA data. Each triangle has vertices that correspond with audio device locations.
- the method involves determining a side length for each side of each of the triangles, performing a forward alignment process of aligning each of the plurality of triangles produce a forward alignment matrix and performing a reverse alignment process of aligning each of the plurality of triangles in a reverse sequence to produce a reverse alignment matrix.
- a final estimate of each audio device location is based, at least in part, on values of the forward alignment matrix and values of the reverse alignment matrix.
- some multi-speaker environments may include television (TV) speakers and a couch positioned for TV viewing. After locating the speakers in the environment, some methods may involve finding a vector pointing to the TV and locating the speech of a user sitting on the couch by triangulation. Some such methods may then involve having the TV emit a sound from its speakers and/or prompting the user to walk up to the TV and locating the user’s speech by triangulation. Some implementations may involve rendering an audio object that pans around the environment.
- a user may provide user input (e.g., saying “Stop”) indicating when the audio object is in one or more predetermined positions within the environment, such as the front of the environment, at a TV location of the environment, etc.
- user input e.g., saying “Stop”
- the user may be located by finding the intersection of directions of arrival of sounds emitted by multiple speakers.
- Some implementations involve determining an estimated distance between at least two audio devices and scaling the distances between other audio devices in the environment according to the estimated distance.
- loudspeaker location data and location data corresponding to the desired listening position or area may be obtained, and all such methods (as well as relevant future methods that may be developed) are meant to be applicable to the implementations of the present disclosure. Accordingly, the specific details disclosed herein should merely be regarded as examples.
- block 915 involves computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers.
- the relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions.
- the relative activation of loudspeakers is computed via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial position.
- the interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD).
- Each set of panning gains may, for example, “approximate” one or more interaural cues of the desired audio signal spatial position by providing audio signals audio signals within a time difference threshold and/or level difference threshold — as compared to the exact ITD and/or ILD of the desired audio signal spatial position — for a particular perceived spatial position of the audio signals, such as within 1%, within 3%, within 5%, within 8%, within 10%, or within another threshold.
- the relative activation of loudspeakers may vary continuously as a function of the desired perceived spatial position of the audio signals.
- block 920 involves rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers.
- method 900 may involve providing a set of rendered audio signals to the set of loudspeakers.
- the desired perceived spatial position corresponding to at least one audio signal may be at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located.
- computing the relative activation of loudspeakers may involve determining a cone of confusion corresponding to the audio signal elevation.
- the cone of confusion may be determined based on an assumed or measured position and orientation of a listener’s head.
- computing the relative activation of loudspeakers may involve determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection. Examples of such two points of intersection are positions o' 7 y and o'j b , which are shown in Figure 3 and described above.
- the rendering process of block 920 may involve combining the panning gains for each of the points of intersection.
- the panning gains may be power-preserving panning gains.
- the power-preserving panning gains may be determined according to Equation (7) or Equation (9b).
- the panning gains for each of the points of intersection may be determined via a flexible rendering process.
- the flexible rendering process may be a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof.
- the panning gains for each of the points of intersection may be determined by applying a cost function that is based on a sum of a spatial term and a proximity term. Equation (8) provides an example of one such cost function.
- the panning gains may be combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual loudspeakers located at each of the points of intersection.
- the control system may determine panning gains for positions o' 7 y and o'j b , which are shown in Figure 3 and described above, as if a virtual loudspeaker were located at each of these points of intersection.
- the additional panning gains may be determined according to a flexible rendering process such as a CMAP process, an FV process, or a combination thereof.
- Other examples may determine the additional panning gains according to a different rendering process, such as a VBAP process or one of the other disclosed methods.
- panning gains gj for Oj may be determined from the additional panning gains for positions o' 7 y and o'j b according to Equation 5.
- Some disclosed implementations include a system or device configured (e.g., programmed) to perform any embodiment of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing any embodiment of the disclosed methods or steps thereof.
- the disclosed system can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and/or otherwise configured to perform any of a variety of operations on data, including an embodiment of the disclosed method or steps thereof.
- a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and/or otherwise configured) to perform an embodiment of the disclosed method (or steps thereof) in response to data asserted thereto.
- Some embodiments of the disclosed system are implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of an embodiment of the disclosed method.
- DSP digital signal processor
- embodiments of the disclosed system (or elements thereof) are implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and/or otherwise configured to perform any of a variety of operations including an embodiment of the disclosed method.
- PC personal computer
- microprocessor which may include an input device and a memory
- elements of some embodiments of the disclosed system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform an embodiment of the disclosed method, and the system also includes other elements (e.g., one or more loudspeakers and/or one or more microphones).
- a general purpose processor configured to perform an embodiment of the disclosed method would typically be coupled to an input device (e.g., a mouse and/or a keyboard), a memory, and a display device.
- an input device e.g., a mouse and/or a keyboard
- a memory e.g., a display device.
- Another aspect of the present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) any disclosed method or steps thereof.
- EEEs Enumerated Example Embodiments
- An audio processing method comprising: receiving, by a control system, audio data, the audio data including a set of audio signals and associated spatial data, the set of audio signals including one or more audio signals and the spatial data indicating a desired perceived spatial position in three dimensions corresponding to an audio signal; obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area, computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers, wherein the relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions, and wherein the relative activation of loudspeakers is computed via combinations of two or more sets of panning gains constructed
- EEE3 The audio processing method of EEE 1 or EEE2, wherein the desired perceived spatial position corresponding to at least one audio signal is at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located.
- EEE5 The audio processing method of EEE4, wherein computing the relative activation of loudspeakers involves determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection.
- EEE6 The audio processing method of EEE5, wherein the panning gains are power-preserving panning gains.
- EEE7 The audio processing method of EEE5 or EEE6, wherein the rendering involves combining the panning gains for each of the points of intersection.
- EEE8 The audio processing method of any one of EEE5 to EEE7, wherein the panning gains for each of the points of intersection are determined via a flexible rendering process.
- EEE9. The audio processing method of EEE8, wherein the flexible rendering process is a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof.
- EEE10 The audio processing method of any one of EEE5 to EEE9, wherein the panning gains for each of the points of intersection are determined by applying a cost function that is based on a sum of a spatial term and a proximity term.
- EEE11 The audio processing method of any one of EEE7 to EEE 10, wherein the panning gains are combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual speakers located at each of the points of intersection.
- EEE12 The method of EEE11, wherein the additional panning gains are determined using any of the methods of EEE8 - EEE 10.
- EEE13 The audio processing method of any one of EEE1 to EEE12, wherein the intended perceived spatial position corresponds with a channel of a channel-based audio format.
- EEE14 The audio processing method of any one of EEE1 to EEE13, wherein the intended perceived spatial position is indicated by audio object metadata.
- EEE15 The audio processing method of any one of EEE1 to EEE14, further comprising providing the set of audio signals to the set of loudspeakers.
- EEE16 An apparatus configured to perform the method of any one of EEE1 to EEE15.
- EEE17 A system configured to perform the method of any one of EEE1 to EEE15.
- EEE18 One or more computer-readable and non-transitory media having instructions stored thereon for controlling one or more devices to perform the method of any one of EEE1 to EEE15.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Stereophonic System (AREA)
Abstract
Some methods may involve receiving audio data including a set of audio signals and associated 3D spatial data, computing a relative activation of each loudspeaker as a function of the desired perceived spatial positions of the audio signals and loudspeaker locations, and rendering the audio signals on the set of loudspeakers. The relative activations may be computed as a function of the desired perceived spatial positions (oj) of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions, and may be computed via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains corresponding to spatial positions (o'jf, o'jb) and approximating interaural cues of the desired audio signal spatial position. Said approximation of interaural cues may be achieved by considering the cone of confusion (310).
Description
RENDERING AUDIO OVER MULTIPLE LOUDSPEAKERS UTILIZING INTERAURAL CUES FOR HEIGHT VIRTUALIZATION
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63/597,644 filed November 9, 2023, and to U.S. Provisional Application No. 63/491,822 filed Mar 23, 2024, each of which is incorporated by reference in its entirety.
TECHNICAL FIELD
[0002] The disclosure pertains to systems and methods for rendering audio for playback by a set of speakers.
BACKGROUND
[0003] Audio devices, including but not limited to smart audio devices, have been widely deployed and are becoming common features of many homes. Although existing systems and methods for controlling audio devices provide benefits, improved systems and methods would be desirable.
NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, “speaker” and “loudspeaker” are used synonymously to denote any sound-emitting transducer (or set of transducers) driven by a single speaker feed. A typical set of headphones includes two speakers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and/or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set. [0008] Throughout this disclosure including in the claims, the term “couples” or “coupled” is used to mean either a direct or indirect connection. Thus, if a first device couples to a second device, that connection may be through a direct connection, or through an indirect connection via other devices and connections.
[0009] Herein, we use the expression “smart audio device” to denote a smart device which is either a single purpose audio device or a virtual assistant (e.g., a connected virtual assistant). A single purpose audio device is a device (e.g., a TV or a mobile phone) including or coupled to at least one microphone (and optionally also including or coupled to at least one speaker) and which is designed largely or primarily to achieve a single purpose. Although a TV typically can play (and is thought of as being capable of playing) audio from program material, in most instances a modem TV runs some operating system on which applications run locally, including the application of watching television. Similarly, the audio input and output in a mobile phone may do many things, but these are serviced by the applications running on the phone. In this sense, a single purpose audio device having speaker(s) and microphone(s) is often configured to run a local application and/or service to use the speaker(s) and microphone(s) directly. Some single purpose audio devices may be configured to group together to achieve playing of audio over a zone or user configured area.
[0010] A virtual assistant (e.g., a connected virtual assistant) is a device (e.g., a smart speaker or voice assistant integrated device) including or coupled to at least one microphone (and optionally also including or coupled to at least one speaker) and which may provide an ability to utilize multiple devices (distinct from the virtual assistant) for applications that are in a sense cloud enabled or otherwise not implemented in or on the virtual assistant itself. Virtual assistants may sometimes work together, e.g., in a discrete and conditionally defined way. For example, two or more virtual assistants may work together in the sense that one of them, for example, the one which is most confident that it has heard a wakeword, responds to the word. The connected devices may form a sort of constellation, which may be managed by one main application which may be (or implement) a virtual assistant.
[0011] Herein, “wakeword” is used in a broad sense to denote any sound (e.g., a word uttered by a human, or some other sound), where a smart audio device is configured to awake in response to detection of (“hearing”) the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, to “awake” denotes that the device enters a state in which it awaits (i.e., is listening for) a sound command. In some instances, what may be referred to herein as a “wakeword” may include more than one word, e.g., a phrase.
[0012] Herein, the expression “wakeword detector” denotes a device configured (or software that includes instructions for configuring a device) to search continuously for alignment between real-time sound (e.g., speech) features and a trained model. Typically, a wakeword event is triggered whenever it is determined by a wakeword detector that the probability that a wakeword has been detected exceeds a predefined threshold. For example, the threshold may be a predetermined threshold which is tuned to give a good compromise between rates of false acceptance and false rejection. Following a wakeword event, a device might enter a state (which may be referred to as an “awakened” state or a state of “attentiveness”) in which it listens for a command and passes on a received command to a larger, more computationally-intensive recognizer.
SUMMARY
[0013] At least some aspects of the present disclosure may be implemented via methods, such as audio processing methods. In some instances, the methods may be implemented, at least in part, by a control system such as those disclosed herein. Some methods involve receiving, by a control system, audio data. In some examples, the audio data may include a set of audio signals and associated spatial data. The set of audio signals may include one or more audio signals. The spatial data may indicate a desired perceived spatial position in three dimensions corresponding to an audio signal.
[0014] Some methods may involve obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers with respect to a desired listening position or area. The set of loudspeakers may include two or more loudspeakers.
[0015] Some methods may involve computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers. In some examples, the relative activation of loudspeakers may be computed as
a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions. According to some examples, the relative activation of loudspeakers may be computed via combinations of two or more sets of panning gains constructed via amplitude panning. Each set of panning gains may approximate one or more interaural cues of the desired audio signal spatial position. The interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD). Some methods may involve rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers. Some methods may involve providing the set of audio signals to the set of loudspeakers.
[0016] In some examples, the relative activation of loudspeakers may vary continuously as a function of the desired perceived spatial position of the audio signals.
[0017] According to some examples, the desired perceived spatial position corresponding to at least one audio signal may be at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located. In some examples, computing the relative activation of loudspeakers may involve determining a cone of confusion corresponding to the audio signal elevation. According to some examples, computing the relative activation of loudspeakers may involve determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection. In some examples, the panning gains may be power-preserving panning gains. According to some examples, the rendering may involve combining the panning gains for each of the points of intersection. In some examples, the panning gains for each of the points of intersection may be determined via a flexible rendering process. According to some such examples, the flexible rendering process may be a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof. In some examples, the panning gains for each of the points of intersection may be determined by applying a cost function that is based on a sum of a spatial term and a proximity term. According to some examples, the panning gains may be combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual speakers located at each of the points of intersection. The additional panning gains may be determined using any of the disclosed methods.
[0018] In some examples, the intended perceived spatial position may correspond with a channel of a channel-based audio format. According to some examples, the intended perceived spatial position may be indicated by audio object metadata.
[0019] Some or all of the operations, functions and/or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include one or more memory devices such as those described herein, including but not limited to one or more random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon.
[0020] At least some aspects of the present disclosure may be implemented via apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus may include an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0021] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.
BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0023] Figure IB shows examples of a listener and an array of loudspeakers in an audio environment.
[0024] Figure 2 shows the audio environment of Figure IB from a different perspective. [0025] Figure 3 shows an example of a cone of confusion.
[0026] Figure 4 shows panning gains generated for the loudspeaker layout shown in Figure 1.
[0027] Figure 5 shows panning gains generated for the loudspeaker layout shown in Figure 1 according to some examples of the present disclosure.
[0028] Figure 6 shows a floor plan of a listening environment, which is a living space in this example.
[0029] Figure 7 is a graph of points indicating speaker activations, in an example embodiment.
[0030] Figure 8 is a graph of tri-linear interpolation between points indicative of speaker activations according to one example.
[0031] Figure 9 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein.
DETAILED DESCRIPTION OF EMBODIMENTS
[0032] Amplitude panning is a common method for rendering spatial audio to an array of loudspeakers. Amplitude panning involves varying the amplitudes of panning gains for a given audio signal as a function of the desired spatial position of the signal such that the perceived position of the sound energy aligns with the target signal position. When the desired audio signal position lies between two speakers, this can be accomplished with pairwise panning, meaning only the two speakers on either side of the desired audio signal position participate in its rendering. (When the desired signal position aligns with the physical location of a loudspeaker, only that loudspeaker participates in its rendering.) Amplitude panning is an apt method for rendering to a myriad of speaker layouts.
[0033] Playback of spatial audio in a consumer environment has typically been tied to a prescribed number of loudspeakers placed in prescribed positions, such as positions corresponding to Dolby 5.1 or 7.1 surround sound. In these cases, content is authored specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby Digital™, Dolby Digital Plus™, etc.) More recently, immersive, object-based spatial audio formats have been introduced (such as Dolby Atmos™) which break this association between the content and specific loudspeaker locations. Instead, the content may be described as a collection of individual audio objects, each with possibly time varying metadata describing the desired perceived location of said audio objects in three- dimensional space and, in some examples, other properties of the audio object. At playback time, the audio content is transformed into loudspeaker feeds by a Tenderer which adapts to the number and location of loudspeakers in the playback system.
[0034] Many typical loudspeaker layouts place all full-range loudspeakers at head height. Layouts including elevated or overhead loudspeakers (for example, Dolby 7.1.4) are
relatively less common. Spatial audio formats - especially object-based formats such as Dolby Atmos - encode audio signals at multiple elevations, causing some of those objects to fall outside the head plane and therefore outside the likely plane of loudspeakers. Rendering this audio for such typical loudspeaker layouts then involves virtualizing elevated and depressed objects with head-height loudspeakers. The protocol for the panner described in International Telecommunication Union Recommendation BS.2127 (ITU-R BS.2127) for these spatial audio situations is to collapse elevated object positions to the loudspeaker plane and render them from there.
[0035] Such rendering methods may be improved upon by considering known perceptual cues that indicate elevated or depressed audio objects as above or below the head to the human ears. Interaural cues are critical for sound localization, including sound localization for elevated or depressed audio objects. These interaural cues are functions of the position of an audio object with respect to the head and refer to the resulting differences in some qualities of the sound when it reaches each ear. Two such cues are the interaural time difference (ITD) and the interaural level difference (ILD). For example, the energy from a sound source directly in front of a listener produces interaural time and level differences of zero, because the source is equidistant from each ear. On the other hand, the energy from a sound source directly to the left of a listener produces an ITD equal to the time it takes the energy to travel past the listener’s left ear and around the listener’s head to the listener’s right ear, and an ILD equal to the air attenuation over the distance from left ear around the head to the right ear, plus any absorption or damping contributed by the head. In fact, the maximum interaural differences a sound source would produce around a person’s head occurs for a source directly to the person’s left or right.
[0036] Elevated sound sources may produce much smaller interaural differences than their head-height counterparts, so loudspeakers in such positions with high interaural differences are unsuitable for rendering elevated objects alone. Likewise, because amplitude panners endeavor to reproduce the spatial effect of a loudspeaker placed at the object location (in the loudspeaker plane), it is unsuitable to render height objects with traditional amplitude panning gains for such locations. However, it would be imprudent to exclude such loudspeakers altogether, especially in a low-density array, for the sake of spatial continuity. [0037] The present disclosure presents solutions to these challenges. Some disclosed examples are capable of preserving interaural cues to impart an impression of object height when rendering audio objects whose elevations differ from those of the loudspeakers, while maintaining the continuity of traditional amplitude panning.
[0038] Figure 1A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 1A are merely provided by way of example. Other implementations may include more, fewer and/or different types and numbers of elements. According to some examples, the apparatus 100 may be, or may include, a smart audio device that is configured for performing at least some of the methods disclosed herein. In other implementations, the apparatus 100 may be, or may include, another device that is configured for performing at least some of the methods disclosed herein, such as a laptop computer, a cellular telephone, a tablet device, a smart home hub, etc. In some such implementations the apparatus 100 may be, or may include, a server.
[0039] In this example, the apparatus 100 includes an interface system 105 and a control system 110. The interface system 105 may, in some implementations, be configured for receiving audio data. The audio data may include audio signals that are to be reproduced by at least some speakers of an environment. The audio data may include one or more audio signals and associated spatial data. The spatial data corresponding to an audio signal may indicate the intended perceived spatial position of that audio signal. In some examples, the spatial data may be, or may include, audio object metadata. In some examples, the spatial data may correspond with a channel of a channel-based audio format. According to some examples, the intended perceived spatial position may be derived from the audio format, such as with higher-order Ambisonics (HO A) or other spherical harmonic based sound field representations.
[0040] The interface system 105 may be configured for providing rendered audio signals to at least some loudspeakers of the set of loudspeakers of the environment. The interface system 105 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0041] The interface system 105 may include one or more network interfaces and/or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and/or a gesture sensor system. In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as the optional memory system 115 shown in Figure 1A. However, the control system 110 may include a memory system in some instances.
[0042] The control system 110 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and/or discrete hardware components.
[0043] In some implementations, the control system 110 may reside in more than one device. For example, a portion of the control system 110 may reside in a device within one of the environments depicted herein and another portion of the control system 110 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 110 may reside in a device within one of the environments depicted herein and another portion of the control system 110 may reside in one or more other devices of the environment. For example, control system functionality may be distributed across multiple smart audio devices of an environment, or may be shared by an orchestrating device (such as what may be referred to herein as a smart home hub) and one or more other devices of the environment. The interface system 105 also may, in some such examples, reside in more than one device. [0044] In some implementations, the control system 110 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 110 may be configured for receiving audio data including a set of audio signals and associated spatial data. The set of audio signals may include one or more audio signals and the spatial data may indicate a desired perceived spatial position in three dimensions corresponding to an audio signal. In some examples, the control system 110 may be configured for obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area.
[0045] According to some examples, the control system 110 may be configured for computing a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers. In some examples, the control system 110 may be configured for computing the relative activation of loudspeakers as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions. According to some examples, the control system 110 may be configured for computing the relative activation of loudspeakers via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial
position. The interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD). Each set of panning gains may, for example, “approximate” one or more interaural cues of the desired audio signal spatial position by providing audio signals audio signals within a time difference threshold and/or level difference threshold — as compared to the exact ITD and/or ILD of the desired audio signal spatial position — for a particular perceived spatial position of the audio signals, such as within 1%, within 3%, within 5%, within 8%, within 10%, or within another threshold. In some examples, the control system 110 may be configured for rendering the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers. According to some examples, the control system 110 may be configured for providing the set of audio signals to the set of loudspeakers.
[0046] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 115 shown in Figure 1A and/or in the control system 110. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, include instructions for controlling at least one device to process audio data. The software may, for example, be executable by one or more components of a control system such as the control system 110 of Figure 1A.
[0047] In some examples, the apparatus 100 may include the optional microphone system 120 shown in Figure 1A. The optional microphone system 120 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc.
[0048] According to some implementations, the apparatus 100 may include the optional loudspeaker system 125 shown in Figure 1A. The optional loudspeaker system 125 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as “speakers.” In some examples, at least some loudspeakers of the optional loudspeaker system 125 may be arbitrarily located . For example, at least some speakers of the optional loudspeaker system 125 may be placed in locations that do not correspond to any standard prescribed speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby
9.1, Hamasaki 22.2, etc. In some such examples, at least some loudspeakers of the optional loudspeaker system 125 may be placed in locations that are convenient to the space (e.g., in locations where there is space to accommodate the loudspeakers), but not in any standard prescribed loudspeaker layout.
[0049] In some implementations, the apparatus 100 may include the optional sensor system 130 shown in Figure 1A. The optional sensor system 130 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some implementations, the optional sensor system 130 may include one or more cameras. In some implementations, the cameras may be free-standing cameras. In some examples, one or more cameras of the optional sensor system 130 may reside in a smart audio device, which may be a single purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 130 may reside in a TV, a mobile phone or a smart speaker.
[0050] In some implementations, the apparatus 100 may include the optional display system 135 shown in Figure 1A. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples wherein the apparatus 100 includes the display system 135, the sensor system 130 may include a touch sensor system and/or a gesture sensor system proximate one or more displays of the display system 135. According to some such implementations, the control system 110 may be configured for controlling the display system 135 to present a graphical user interface (GUI), such as one of the GUIs disclosed herein.
[0051] According to some examples the apparatus 100 may be, or may include, a smart audio device. In some such implementations the apparatus 100 may be, or may include, a wakeword detector. For example, the apparatus 100 may be, or may include, a virtual assistant.
[0052] Cones of confusion are three-dimensional (3D) surfaces around a listener’ s head on which sound sources will impart equivalent interaural cues. (An example of a cone of confusion is shown in Figure 3 and will be described in detail below.) Cones of confusion extend to the left from a listener’s left ear and to the right from a listener’s right ear. The open end of the cone forms a circle that intersects the head-height plane at its diameter.
Cones of confusion may be used to determine coordinates for head-height (or zero -elevation) sound sources that would produce the same interaural differences as a target elevated source. Each cone of confusion intersects a circle in the head-height plane having its center at the center of the listener’s head at two points. By interpolating between the panning gains for
these two points as a function of target audio object position, it is possible to mimic the interaural cues while still maintaining continuity and some degree of sparsity in the resulting set of panning gains.
[0053] Some disclosed examples involve rendering a set of one or more audio signals, each with an associated desired perceived spatial position, over a set of two or more loudspeakers, wherein:
• The locations of the set of loudspeakers with respect to a desired listening position or area are provided to the Tenderer; and
• The relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the one or more audio signals and the locations of the loudspeakers, wherein: o the spatial position of the audio signals is given in three dimensions; o speaker activation varies as a function of spatial position of the audio signals, particularly the elevation of the audio signal positions with respect to the loudspeaker elevations; and o speaker activation is determined via combinations of two or more sets of panning gains (constructed via amplitude panning), each set matching the interaural cues of the desired audio signal spatial position.
[0054] Figure IB shows examples of a listener and an array of loudspeakers in an audio environment. According to this example, the audio environment 150 includes an array of loudspeakers Si, S2, S3 and S4 at the height of a listener’s head 155. This height may be referred to herein as “head height.” In this example, the array of loudspeakers S1-S4 are being used to render the audio object o(-.
[0055] Figure IB shows a coordinate system having its center at the center of the listener’s head 155. In this example, the audio object O is also positioned at head height, which we may define as zt = 0. Accordingly, the coordinates of the audio object O may be expressed as follows:
Oi = [* y 0] (i)
[0056] Using an amplitude panner to render the audio object O will result in a set of gains gt for each loudspeaker.
9i = [9n - 9is\ (2)
[0057] Figure 2 shows the audio environment of Figure IB from a different perspective. The same set of loudspeakers may be used to render an audio object Oj with the same x and y coordinates, but now with nonzero height Zj.
Oj = [Xi,yi,Zj] (3)
[0058] Figure 3 shows an example of a cone of confusion. As noted elsewhere herein, cones of confusion are 3D surfaces around a listener’s head on which sound sources will impart equivalent interaural cues. Cones of confusion extend to the left from a listener’s left ear and to the right from a listener’s right ear. In the example shown in Figure 3, the origin of the coordinate system 305 is at the center of the listener’s head 155. Here, the listener’s head 155 is facing in the positive direction of the y axis, the positive direction of the x axis extends from the left of the listener’s head 155 and the positive direction of the z axis extends through the top of the listener’s head 155.
[0059] The open end of the cone of confusion 310 forms a circle 315 that intersects the circle 320 at height Zj, which includes position Oj. Many amplitude panners (including those operating according to ITU-R BS.2127) would produce the same panning gains gj for an object at position Oj as gt for an object at position or
9j = 9t (4)
[0060] Equivalent gains produce equivalent interaural cues. However, one may observe that position Of is not on the cone of confusion 310. This means that an audio object at position o(- would not produce the same interaural differences that an audio object at position Oj would produce.
[0061] The circle 315 also intersects the circle 325 in the x-y plane — which is also referred to herein as a “head-height plane” — at positions o'7y and o' jb. The subscripts /and b represent front and back, respectively. This means that an audio object positioned at either position o' jf or position o' jb would produce the same interaural differences as an audio object at position 6j. By interpolating between the panning gains for positions o'7y and o'jb as a function of target audio object position, it is possible to mimic the interaural cues for audio
objects positioned above or below the head-height plane, while still maintaining continuity and some degree of sparsity in the resulting set of panning gains.
[0062] Every set of panning gains on the circle created by the intersection of the cone surface xjc + z?c = xj + z? with the plane y = y7 produce the same interaural cues as Oj. As the speakers are all coplanar with the head at z = 0, ZjC = 0 may be substituted in to solve for the locations of the two head-height points o y (toward the front of the head) and Ojb (toward the back of the head) at which a sound object would produce equivalent interaural cues to one at Oj. Panning gains g' jf(oj) for these two points
using the amplitude panner, and new panning gains g^ for Oj, which produce the correct interaural cues, may be constructed as a linear combination of the two sets of equivalent head-height gains, for example as follows:
[0063] The two coefficients (o7 )
may vary continuously as a function of the object position Oj and may also be related to each other for spatial continuity, for example as follows:
[0064] Sample parameters for this relationship include n = 2 and y = 1, which makes the interpolation between the front and back set of panning gains power-preserving.
Additionally, in some examples gj may itself be normalized as well for the sake of continuity. For m = 2, Equation 7 ensures that the final set of panning gains gj is power preserving.
[0065] In flexible rendering, spatial audio may be rendered over an arbitrary number of arbitrarily placed speakers. Various methods have been developed to implement flexible rendering, including but not limited to center of mass amplitude panning (CMAP) . At a high
level, CMAP involves rendering a set of one or more audio signals, each with an associated desired perceived spatial position, for playback over a set of two or more speakers, where the relative activation of speakers of the set is a function of a model of perceived spatial position of said audio signals played back over the speakers and a proximity of the desired perceived spatial position of the audio signals to the positions of the speakers. The model may be designed to ensure that the audio signal is heard by the listener near its intended spatial position, and the proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term favors the activation of speakers that are near the desired perceived spatial position of the audio signal. For CMAP, this functional relationship may be conveniently derived from a cost function written as the sum of two terms, one for the spatial aspect and one for proximity:
[0066] Here, the set {s denotes the positions of a set of M loudspeakers, 6 denotes the desired perceived spatial position of the audio signal, and g denotes an M dimensional vector of speaker activations. For CMAP, each activation in the vector represents a gain per speaker. An optimal vector of activations may be found by minimizing the cost function across activations:
[0067] With certain definitions of the cost function, it is difficult to control the absolute level of the optimal activations resulting from the above minimization, though the relative level between the components of gopt is appropriate. To deal with this problem, a subsequent normalization of gopt may be performed so that the absolute level of the activations is controlled. For example, normalization of the vector to have unit length may be desirable, which is in line with constant power panning rules:
Equation (9b) parallels Equation (7).
[0068] The exact behavior of this type of flexible rendering algorithm is dictated by the particular construction of the two terms of the cost function, Cspatiai and Cproximity. For CMAP, Cspatiai is derived from a model that places the perceived spatial position of an audio signal playing from a set of loudspeakers at the center of mass of those loudspeakers’ positions weighted by their associated activating gains gt (elements of the vector g):
[0069] Equation (10) may then be manipulated into a spatial cost representing the squared error between the desired audio position and that produced by the activated loudspeakers:
[0070] Conveniently, the spatial term of the cost function for CMAP defined in Equation (11) can be rearranged into a matrix quadratic as a function of speaker activations g:
[0071] In Equation (12), A represents an M x M square matrix, B represents a 1 x M vector, and C represents a scalar. The matrix A is of rank 2, and therefore when M > 2 there exist an infinite number of speaker activations g for which the spatial error term equals zero.
Introducing the second term of the cost function, Cproximity. removes this indeterminacy and results in a particular solution with perceptually beneficial properties in comparison to the other possible solutions. For CMAP, Cproximity may be constructed such that activation of speakers whose position st is distant from the desired audio signal position 6 is penalized more than activation of speakers whose position is close to the desired position. This construction yields an optimal set of speaker activations that is sparse, where only speakers in close proximity to the desired audio signal’s position are significantly activated, and
practically results in a spatial reproduction of the audio signal that is perceptually more robust to listener movement around the set of speakers.
[0072] To this end, the second term of the cost function, Cproximity, may be defined as a distance- weighted sum of the absolute values squared of speaker activations. This may be represented compactly in matrix form as:
[0073] In Equation (13a), D represents a diagonal matrix of distance penalties between the desired audio position and each speaker, which may be represented as follows: di = distance o, Si) (13b)
[0074] The distance penalty function can take on many forms, but the following is a useful parameterization:
[0075] In Equation (13c), ||o — s | represents the Euclidean distance between the desired audio position and speaker position and a and ft represent tunable parameters. The parameter a indicates the global strength of the penalty; dQ corresponds to the spatial extent of the distance penalty (loudspeakers at a distance around dQ or futher away will be penalized), and P accounts for the abruptness of the onset of the penalty at distance dQ.
[0076] Combining the two terms of the cost function defined in Equations 8 and 9a yields the overall cost function:
C(g) = g*Ag + Bg + C + g*Dg = g*(A + D)g + Bg + C (14)
[0077] Setting the derivative of this cost function with respect to g equal to zero and solving for g yields the optimal speaker activation solution:
[0078] In general, the optimal solution in Equation (15) may yield speaker activations that are negative in value. For the CMAP construction of the flexible Tenderer, such negative activations may not be desirable, and thus Equation (15) may be minimized subject to all activations remaining positive.
[0079] In some examples, a flexible rendering process may be implemented using CMAP as the baseline amplitude panner from which to select sets of panning gains with equivalent interaural cues. Additionally, Equation 8 may be optimized for the object position o = o(- and speaker set
in order to determine the particular dependence of o and pb on 6j. Other flexible rendering examples may involve vector base amplitude panning (VBAP), the amplitude panner of ITU-R BS.2127, other flexible rendering techniques, or combinations thereof.
[0080] Figure 4 shows panning gains generated for the loudspeaker layout shown in Figure 1 for an audio object with an elevation of 45 degrees above the head height and a varying azimuth angle indicated by the x-axis of the figure. In these examples, the panning gains were generated via a CMAP-based process which essentially projects the position of the audio object into the head height plane when computing the panning gains. Note each speaker is soloed — meaning that there is zero gain for other loudspeakers — when the azimuth angle of the object matches that of the loudspeaker’s angular position, as a result of the enforced sparsity. This is not a desirable outcome because loudspeakers S2 and S4 are positioned directly to the right and left of the head, generating maximum interaural differences, impossible for an audio object at 45 degrees elevation.
[0081] Figure 5 shows panning gains generated for the same loudspeaker layout shown in Figure 1 according to some examples of the present disclosure. In this case loudspeakers S2 and S4 are never soloed because the overall panning gains are computed as interpolations between the two sets of panning gains corresponding to the two projection points of the object position along the cone of confusion into the head-height plane. For any azimuth angle, neither of these two points corresponds to a case where loudspeakers S2 and S4 are soloed. Instead, each point generates interaural cues commensurate with an object position at that particular azimuth and 45 degrees elevation. In comparison to Figure 4, interpolation between the panning gains from these two points causes a spreading of energy from loudspeakers S2 and S4 to loudspeakers Si and S3. The end result is maintenance of interaural
cues commensurate with the audio object elevation using panning gains that vary continuously as a function of object azimuth.
[0082] Although the previous discussion involved simple examples of loudspeaker placement, the same underlying principles apply to the flexible rendering of spatial audio over an arbitrary number of arbitrarily placed loudspeakers, such as the arbitrarily placed loudspeakers shown in Figure 6. Figure 6 shows a floor plan of a listening environment, which is a living space in this example. As with other figures provided herein, the types and numbers of elements shown in Figure 6 are merely provided by way of example. Other implementations may include more, fewer and/or different types and numbers of elements. According to this example, the environment 600 includes a living room 610 at the upper left, a kitchen 615 at the lower center, and a bedroom 622 at the lower right. Boxes and circles distributed across the living space represent a set of loudspeakers 605a-605h, at least some of which may be smart speakers in some implementations, placed in locations convenient to the space, but not adhering to any standard prescribed layout (arbitrarily placed). In some examples, flexible rendering of spatial audio may be rendered to the loudspeakers 605a-605h according to one or more disclosed embodiments.
[0083] According to some examples, the environment 600 may include a smart home hub for implementing at least some of the disclosed methods. According to some such implementations, the smart home hub may include at least a portion of the above-described control system 110. In some examples, a smart device (such as a smart speaker, a mobile phone, a smart television, a device used to implement a virtual assistant, etc.) may implement the smart home hub.
[0084] In this example, the environment 600 includes cameras 61 la-61 le, which are distributed throughout the environment. In some implementations, one or more smart audio devices in the environment 600 also may include one or more cameras. The one or more smart audio devices may be single purpose audio devices or virtual assistants. In some such examples, one or more cameras of the optional sensor system 130 may reside in or on the television 630, in a mobile phone or in a smart speaker, such as one or more of the loudspeakers 605b, 605d, 605e or 605h. Although cameras 61 la-61 le are not shown in every depiction of the environment 600 presented in this disclosure, each of the environments 600 may nonetheless include one or more cameras in some implementations.
[0085] Figure 7 is a graph of points indicating speaker activations, in an example embodiment. In this example, the x and y dimensions are sampled with 15 points and the z dimension is sampled with 5 points. Other implementations may include more samples or
fewer samples. According to this example, each point represents the M speaker activations for the CMAP solution.
[0086] At runtime, to determine the actual activations for each speaker, tri-linear interpolation between the speaker activations of the nearest 8 points may be used in some examples. Figure 8 is a graph of tri-linear interpolation between points indicative of speaker activations according to one example. In this example, the process of successive linear interpolation includes interpolation of each pair of points in the top plane to determine first and second interpolated points 805a and 805b, interpolation of each pair of points in the bottom plane to determine third and fourth interpolated points 810a and 810b, interpolation of the first and second interpolated points 805a and 805b to determine a fifth interpolated point 815 in the top plane, interpolation of the third and fourth interpolated points 810a and 810b to determine a sixth interpolated point 820 in the bottom plane, and interpolation of the fifth and sixth interpolated points 815 and 820 to determine a seventh interpolated point 825 between the top and bottom planes. Although tri-linear interpolation is an effective interpolation method, one of skill in the art will appreciate that tri-linear interpolation is just one possible interpolation method that may be used in implementing aspects of the present disclosure, and that other examples may include other interpolation methods.
[0087] Figure 9 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 900, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of method 900 may be performed concurrently. Moreover, some implementations of method 900 may include more or fewer blocks than shown and/or described. The blocks of method 900 may be performed by one or more devices, which may be (or may include) a control system such as the control system 110 that is shown in Figure 1A and described above.
[0088] According to this example, block 905 involves receiving, by a control system, audio data. In this example, the audio data includes a set of one or more audio signals and associated spatial data. Here, the spatial data indicates an intended perceived spatial position corresponding to an audio signal. In some examples the spatial data may be, or may include, spatial metadata of an object-based audio format such as Dolby Atmos™. In some examples, the intended perceived spatial position may be represented as Oj or as o(-, as disclosed herein. In some instances the spatial data may be, or may correspond with, channels of a channelbased audio format such as a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4 or Dolby 9.1
format. Accordingly, the intended perceived spatial position may correspond with a channel of a channel-based audio format, may correspond with metadata, or may correspond with both the channel and the metadata. In some examples, the intended perceived spatial position may be derived from the audio format, such as with higher-order Ambisonics (HOA) or other spherical harmonic based sound field representations.
[0089] In this example, block 910 involves obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area. According to some examples, block 910 may involve obtaining the loudspeaker location data and/or location data corresponding to the desired listening position or area from a data structure stored in a memory of, or accessible by, the control system. In other examples, block 910 may involve determining the loudspeaker location data and/or location data corresponding to the desired listening position or area.
[0090] The loudspeaker location data and/or location data corresponding to the desired listening position or area may be obtained through numerous mechanisms known in the art. According to some examples, loudspeaker location data and/or location data corresponding to the desired listening position or area may be specified according to a standard loudspeaker layout, such as a Dolby 5.1 loudspeaker layout. In some such examples, a user could, for example, provide input to a device — such as an audio/video receiver (AVR) indicating that that they have a set of loudspeakers in a Dolby 5.1 layout. This is one example of “obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area” in block 910. The AVR may then assume a "canonical" Dolby 5.1 layout and apply one or more disclosed methods — such as the operations of blocks 915 and 920 — according to the corresponding loudspeaker layout.
[0091] In some applications, such as an automobile cabin, loudspeaker location data and location data corresponding to the desired listening position or area are fixed and can be physically measured, e.g. with a tape measure, or obtained from layout information such as computer assisted drafting CAD data.
[0092] In Other examples, such as examples involving the home environment shown in Figure 6, block 910 may involve a more adaptable approach that can automatically detect these loudspeaker and/or user locations and orientations through a one-time setup procedure or even dynamically across time. In Hess, Wolfgang, Head-Tracking Techniques for Virtual Acoustic Applications, (AES 133rd Convention, October 2012), which is hereby incorporated
by reference, numerous commercially available techniques for tracking both the position and orientation of a listener’s head in the context of spatial audio reproduction systems are presented. One particular example discussed is the Microsoft Kinect. With its depth sensing and standard cameras along with a publicly available software (Windows Software Development Kit (SDK)), the positions and orientations of the heads of several listeners in a space can be simultaneously tracked using a combination of skeletal tracking and facial recognition. Although the Kinect for Windows has been discontinued, the Azure Kinect developer kit (DK), which implements the next generation of Microsoft’s depth sensor, is currently available.
[0093] In U.S. Patent No. 10,779,084, entitled “Automatic Discovery and Localization of Speaker Locations in Surround Sound Systems,” which is hereby incorporated by reference, a system is described which can automatically locate the positions of loudspeakers and microphones in a listening environment by acoustically measuring the time-of-arrival (TOA) between each speaker and microphone. A listening position or area may be detected by placing and locating a microphone at a desired listening position (a microphone in a mobile phone held by the listener, for example), and an associated listening orientation may be defined by placing another microphone at a point in the viewing direction of the listener, e.g. at the TV. Alternatively, the listening orientation may be defined by locating a loudspeaker in the viewing direction, e.g. the loudspeakers on the TV.
[0094] In Shi, Guangi et al, Spatial Calibration of Surround Sound Systems including Listener Position Estimation, (AES 137th Convention, October 2014), which is hereby incorporated by reference, a system is described in which a single linear microphone array associated with a component of the reproduction system whose location is predictable, such as a soundbar a front center speaker, measures the time-difference-of-arrival (TDOA) for both satellite loudspeakers and a listener to locate the positions of both the loudspeakers and listener. In this case, the listening orientation is inherently defined as the line connecting the detected listening position and the component of the reproduction system that includes the linear microphone array, such as a sound bar that is co-located with a television (placed directly above or below the television). Because the sound bar’s location is predictably placed directly above or below the video screen, the geometry of the measured distance and incident angle can be translated to an absolute position relative to any point in front of that reference sound bar location using simple trigonometric principles. The distance between a loudspeaker and a microphone of the linear microphone array can be estimated by playing a test signal and measuring the time of flight (TOF) between the emitting loudspeaker and the
receiving microphone. The time delay of the direct component of a measured impulse response can be used for this purpose. The impulse response between the loudspeaker and a microphone array element can be obtained by playing a test signal through the loudspeaker under analysis. For example, either a maximum length sequence (MLS) or a chirp signal (also known as logarithmic sine sweep) can be used as the test signal. The room impulse response can be obtained by calculating the circular cross-correlation between the captured signal and the MLS input. Fig. 2 of this reference shows an echoic impulse response obtained using a MLS input. This impulse response is said to be similar to a measurement taken in a typical office or living room. The delay of the direct component is used to estimate the distance between the loudspeaker and the microphone array element. For loudspeaker distance estimation, any loopback latency of the audio device used to playback the test signal should be computed and removed from the measured TOF estimate.
[0095] International Publication Number WO 2022/118072, entitled “Pervasive Acoustic Mapping,” which is hereby incorporated by reference, discloses additional methods for estimating loudspeaker location data and location data corresponding to the desired listening position or area. This disclosure describes multiple techniques that may be used in various combinations in order to provide automated acoustic mapping. The acoustic mapping may be pervasive and ongoing. Such acoustic mapping may sometimes be referred to as “continuous,” in the sense that the acoustic mapping may be continued after an initial set-up process and may be responsive to changing conditions in the audio environment, such as changing noise sources and/or levels, loudspeaker relocation, the deployment of additional loudspeakers, the relocation and/or re-orientation of one or more listeners, etc. Some disclosed methods involve generating calibration signals that are injected (e.g., mixed) into the audio content being rendered by audio devices in an audio environment. In some such examples, the calibration signals may be, or may include, acoustic direct sequence spread spectrum (DSSS) signals. In other examples, the calibration signals may be, or may include, other types of acoustic calibration signals, such as swept sinusoidal acoustic signals, white noise, “colored noise,” such as pink noise (a spectrum of frequencies that decreases in intensity at a rate of three decibels per octave), acoustic signals corresponding to music, etc. [0096] United States Patent Application Publication No. 2023/0040846 Al, entitled “Audio Device Auto-Location,” which is hereby incorporated by reference, discloses additional methods for estimating loudspeaker location data and location data corresponding to the desired listening position or area. Some disclosed methods for estimating an audio device location in an environment involve obtaining direction of arrival (DOA) data for each audio
device of a plurality of audio devices in the environment and determining interior angles for each of a plurality of triangles based on the DOA data. Each triangle has vertices that correspond with audio device locations. The method involves determining a side length for each side of each of the triangles, performing a forward alignment process of aligning each of the plurality of triangles produce a forward alignment matrix and performing a reverse alignment process of aligning each of the plurality of triangles in a reverse sequence to produce a reverse alignment matrix. A final estimate of each audio device location is based, at least in part, on values of the forward alignment matrix and values of the reverse alignment matrix. Some such methods may yield a result that is correct up to an unknown scale and rotation. In many applications, absolute scale is unnecessary, and rotations can be resolved by placing additional constraints on the solution. For example, some multi-speaker environments may include television (TV) speakers and a couch positioned for TV viewing. After locating the speakers in the environment, some methods may involve finding a vector pointing to the TV and locating the speech of a user sitting on the couch by triangulation. Some such methods may then involve having the TV emit a sound from its speakers and/or prompting the user to walk up to the TV and locating the user’s speech by triangulation. Some implementations may involve rendering an audio object that pans around the environment. A user may provide user input (e.g., saying “Stop”) indicating when the audio object is in one or more predetermined positions within the environment, such as the front of the environment, at a TV location of the environment, etc. According to some such examples, after locating the speakers within an environment and determining their orientation, the user may be located by finding the intersection of directions of arrival of sounds emitted by multiple speakers. Some implementations involve determining an estimated distance between at least two audio devices and scaling the distances between other audio devices in the environment according to the estimated distance.
[0097] As can be seen, there exist numerous mechanisms through which loudspeaker location data and location data corresponding to the desired listening position or area may be obtained, and all such methods (as well as relevant future methods that may be developed) are meant to be applicable to the implementations of the present disclosure. Accordingly, the specific details disclosed herein should merely be regarded as examples.
[0098] According to this example, block 915 involves computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers. In this example, the relative activation of
loudspeakers is computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions. According to this example, the relative activation of loudspeakers is computed via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial position. The interaural cues may be, or may include, interaural time difference (ITD) and interaural level difference (ILD). Each set of panning gains may, for example, “approximate” one or more interaural cues of the desired audio signal spatial position by providing audio signals audio signals within a time difference threshold and/or level difference threshold — as compared to the exact ITD and/or ILD of the desired audio signal spatial position — for a particular perceived spatial position of the audio signals, such as within 1%, within 3%, within 5%, within 8%, within 10%, or within another threshold. In some examples, the relative activation of loudspeakers may vary continuously as a function of the desired perceived spatial position of the audio signals.
[0099] In this example, block 920 involves rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers. According to some examples, method 900 may involve providing a set of rendered audio signals to the set of loudspeakers.
[0100] According to some examples, the desired perceived spatial position corresponding to at least one audio signal may be at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located. In some such examples, computing the relative activation of loudspeakers may involve determining a cone of confusion corresponding to the audio signal elevation. As noted elsewhere herein, in some examples, the cone of confusion may be determined based on an assumed or measured position and orientation of a listener’s head.
[0101] In some examples, computing the relative activation of loudspeakers may involve determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection. Examples of such two points of intersection are positions o'7y and o'jb, which are shown in Figure 3 and described above. In some examples, the rendering process of block 920 may involve combining the panning gains for each of the points of intersection. According to some examples, the panning gains may be power-preserving panning gains. In some such
examples, the power-preserving panning gains may be determined according to Equation (7) or Equation (9b).
[0102] In some examples, the panning gains for each of the points of intersection may be determined via a flexible rendering process. In some such examples, the flexible rendering process may be a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof. According to some examples, the panning gains for each of the points of intersection may be determined by applying a cost function that is based on a sum of a spatial term and a proximity term. Equation (8) provides an example of one such cost function.
[0103] According to some examples, the panning gains may be combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual loudspeakers located at each of the points of intersection. For example, the control system may determine panning gains for positions o'7y and o'jb, which are shown in Figure 3 and described above, as if a virtual loudspeaker were located at each of these points of intersection. In some such examples, the additional panning gains may be determined according to a flexible rendering process such as a CMAP process, an FV process, or a combination thereof. Other examples may determine the additional panning gains according to a different rendering process, such as a VBAP process or one of the other disclosed methods. In one such example, panning gains gj for Oj may be determined from the additional panning gains for positions o'7y and o'jb according to Equation 5.
[0104] Some disclosed implementations include a system or device configured (e.g., programmed) to perform any embodiment of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing any embodiment of the disclosed methods or steps thereof. For example, the disclosed system can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and/or otherwise configured to perform any of a variety of operations on data, including an embodiment of the disclosed method or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and/or otherwise configured) to perform an embodiment of the disclosed method (or steps thereof) in response to data asserted thereto.
[0105] Some embodiments of the disclosed system are implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and
otherwise configured) to perform required processing on audio signal(s), including performance of an embodiment of the disclosed method. Alternatively, embodiments of the disclosed system (or elements thereof) are implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and/or otherwise configured to perform any of a variety of operations including an embodiment of the disclosed method. Alternatively, elements of some embodiments of the disclosed system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform an embodiment of the disclosed method, and the system also includes other elements (e.g., one or more loudspeakers and/or one or more microphones). A general purpose processor configured to perform an embodiment of the disclosed method would typically be coupled to an input device (e.g., a mouse and/or a keyboard), a memory, and a display device. [0106] Another aspect of the present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) any disclosed method or steps thereof.
[0107] While specific embodiments and applications have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope described and claimed herein. It should be understood that while certain forms have been shown and described, the scope of the present disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.
[0108] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):
EEE1. An audio processing method, comprising: receiving, by a control system, audio data, the audio data including a set of audio signals and associated spatial data, the set of audio signals including one or more audio signals and the spatial data indicating a desired perceived spatial position in three dimensions corresponding to an audio signal; obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area,
computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers, wherein the relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions, and wherein the relative activation of loudspeakers is computed via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial position, and rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers.
EEE2. The audio processing method of EEE1, wherein the relative activation of loudspeakers varies continuously as a function of the desired perceived spatial position of the audio signals.
EEE3. The audio processing method of EEE 1 or EEE2, wherein the desired perceived spatial position corresponding to at least one audio signal is at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located.
EEE4. The audio processing method of EEE3, wherein computing the relative activation of loudspeakers involves determining a cone of confusion corresponding to the audio signal elevation.
EEE5. The audio processing method of EEE4, wherein computing the relative activation of loudspeakers involves determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection.
EEE6. The audio processing method of EEE5, wherein the panning gains are power-preserving panning gains.
EEE7. The audio processing method of EEE5 or EEE6, wherein the rendering involves combining the panning gains for each of the points of intersection.
EEE8. The audio processing method of any one of EEE5 to EEE7, wherein the panning gains for each of the points of intersection are determined via a flexible rendering process.
EEE9. The audio processing method of EEE8, wherein the flexible rendering process is a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof.
EEE10. The audio processing method of any one of EEE5 to EEE9, wherein the panning gains for each of the points of intersection are determined by applying a cost function that is based on a sum of a spatial term and a proximity term.
EEE11. The audio processing method of any one of EEE7 to EEE 10, wherein the panning gains are combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual speakers located at each of the points of intersection.
EEE12. The method of EEE11, wherein the additional panning gains are determined using any of the methods of EEE8 - EEE 10.
EEE13. The audio processing method of any one of EEE1 to EEE12, wherein the intended perceived spatial position corresponds with a channel of a channel-based audio format.
EEE14. The audio processing method of any one of EEE1 to EEE13, wherein the intended perceived spatial position is indicated by audio object metadata.
EEE15. The audio processing method of any one of EEE1 to EEE14, further comprising providing the set of audio signals to the set of loudspeakers.
EEE16. An apparatus configured to perform the method of any one of EEE1 to EEE15.
EEE17. A system configured to perform the method of any one of EEE1 to EEE15.
EEE18. One or more computer-readable and non-transitory media having instructions stored thereon for controlling one or more devices to perform the method of any one of EEE1 to EEE15.
Claims
1. An audio processing method, comprising: receiving, by a control system, audio data, the audio data including a set of audio signals and associated spatial data, the set of audio signals including one or more audio signals and the spatial data indicating a desired perceived spatial position in three dimensions corresponding to an audio signal; obtaining, by the control system, loudspeaker location data indicating locations of each loudspeaker of a set of loudspeakers including two or more loudspeakers with respect to a desired listening position or area, computing, by the control system, a relative activation of each loudspeaker of the set of loudspeakers as a function of the desired perceived spatial positions of the one or more audio signals and the location of each loudspeaker of the set of loudspeakers, wherein the relative activation of loudspeakers is computed as a function of the desired perceived spatial positions of the audio signals, including an elevation of the audio signal positions with respect to loudspeaker positions, and wherein the relative activation of loudspeakers is computed via combinations of two or more sets of panning gains constructed via amplitude panning, each set of panning gains approximating interaural cues of the desired audio signal spatial position, and rendering, by the control system, the set of audio signals for reproduction on the set of loudspeakers based on the relative activation of loudspeakers.
2. The audio processing method of claim 1 , wherein the relative activation of loudspeakers varies continuously as a function of the desired perceived spatial position of the audio signals.
3. The audio processing method of claim 1 or claim 2, wherein the desired perceived spatial position corresponding to at least one audio signal is at an audio signal elevation relative to a loudspeaker plane in which the set of loudspeakers are located.
4. The audio processing method of claim 3, wherein computing the relative activation of loudspeakers involves determining a cone of confusion corresponding to the audio signal elevation.
5. The audio processing method of claim 4, wherein computing the relative activation of loudspeakers involves determining two points of intersection between the cone of confusion and the loudspeaker plane and determining panning gains for each of the points of intersection.
6. The audio processing method of claim 5, wherein the panning gains are powerpreserving panning gains.
7. The audio processing method of claim 5 or claim 6, wherein the rendering involves combining the panning gains for each of the points of intersection.
8. The audio processing method of any one of claims 5-7, wherein the panning gains for each of the points of intersection are determined via a flexible rendering process.
9. The audio processing method of claim 8, wherein the flexible rendering process is a center of mass amplitude panning process, a flexible virtualization process, or a combination thereof.
10. The audio processing method of any one of claims 5-9, wherein the panning gains for each of the points of intersection are determined by applying a cost function that is based on a sum of a spatial term and a proximity term.
11. The audio processing method of any one of claims 7-10, wherein the panning gains are combined as a function of additional panning gains determined as a function of the desired audio object position with respect to virtual speakers located at each of the points of intersection.
12. The audio processing method of claim 11, wherein the additional panning gains are determined using any of the methods of claims 8-10.
13. The audio processing method of any one of claims 1-12, wherein the intended perceived spatial position corresponds with a channel of a channel-based audio format.
14. The audio processing method of any one of claims 1-13, wherein the intended perceived spatial position is indicated by audio object metadata.
15. The audio processing method of any one of claims 1-14, further comprising providing the set of audio signals to the set of loudspeakers.
16. An apparatus configured to perform the method of any one of claims 1-15.
17. A system configured to perform the method of any one of claims 1-15.
18. One or more computer-readable and non-transitory media having instructions stored thereon for controlling one or more devices to perform the method of any one of claims 1-15.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363491822P | 2023-03-23 | 2023-03-23 | |
| US202363597644P | 2023-11-09 | 2023-11-09 | |
| PCT/US2024/021013 WO2024197200A1 (en) | 2023-03-23 | 2024-03-21 | Rendering audio over multiple loudspeakers utilizing interaural cues for height virtualization |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4684538A1 true EP4684538A1 (en) | 2026-01-28 |
Family
ID=90731404
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24719406.1A Pending EP4684538A1 (en) | 2023-03-23 | 2024-03-21 | Rendering audio over multiple loudspeakers utilizing interaural cues for height virtualization |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4684538A1 (en) |
| WO (1) | WO2024197200A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9949052B2 (en) * | 2016-03-22 | 2018-04-17 | Dolby Laboratories Licensing Corporation | Adaptive panner of audio objects |
| EP3519846B1 (en) | 2016-09-29 | 2023-03-22 | Dolby Laboratories Licensing Corporation | Automatic discovery and localization of speaker locations in surround sound systems |
| WO2021127286A1 (en) | 2019-12-18 | 2021-06-24 | Dolby Laboratories Licensing Corporation | Audio device auto-location |
| JP7815249B2 (en) | 2020-12-03 | 2026-02-17 | ドルビー・インターナショナル・アーベー | Pervasive Acoustic Mapping |
-
2024
- 2024-03-21 WO PCT/US2024/021013 patent/WO2024197200A1/en not_active Ceased
- 2024-03-21 EP EP24719406.1A patent/EP4684538A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024197200A1 (en) | 2024-09-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12003946B2 (en) | Adaptable spatial audio playback | |
| CN114846821B (en) | Automatic positioning of audio devices | |
| US12170875B2 (en) | Managing playback of multiple streams of audio over multiple speakers | |
| CN114788304B (en) | Method for reducing errors in an ambient noise compensation system | |
| EP3721187B1 (en) | An apparatus and method for processing volumetric audio | |
| US20240422503A1 (en) | Rendering based on loudspeaker orientation | |
| JP7789915B2 (en) | Distributed Audio Device Ducking | |
| WO2024197200A1 (en) | Rendering audio over multiple loudspeakers utilizing interaural cues for height virtualization | |
| US12003948B1 (en) | Multi-device localization | |
| US20240284136A1 (en) | Adaptable spatial audio playback | |
| EP4714129A1 (en) | Virtual sound sources and rendering techniques | |
| Galindo et al. | Microphone array design for spatial audio object early reflection parametrisation from room impulse responses | |
| KR102958371B1 (en) | Audio Device Auto-Location | |
| CN118216163A (en) | Rendering based on loudspeaker orientation | |
| Blanco Galindo et al. | Microphone Array Design for Spatial Audio Object Early Reflection Parametrisation from Room Impulse Responses | |
| HK40069549A (en) | Audio device auto-location | |
| CN118235435A (en) | Distributed Audio Device Ducking | |
| Vesa | Studies on Binaural and Monaural Signal Analysis |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250929 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |