EP2954703B1 - Bestimmung von renderern für kugelflächenfunktionskoeffizienten - Google Patents
Bestimmung von renderern für kugelflächenfunktionskoeffizienten Download PDFInfo
- Publication number
- EP2954703B1 EP2954703B1 EP14707870.3A EP14707870A EP2954703B1 EP 2954703 B1 EP2954703 B1 EP 2954703B1 EP 14707870 A EP14707870 A EP 14707870A EP 2954703 B1 EP2954703 B1 EP 2954703B1
- Authority
- EP
- European Patent Office
- Prior art keywords
- renderer
- speaker geometry
- geometry
- channel
- speaker
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
- 238000000034 method Methods 0.000 claims description 119
- 238000009877 rendering Methods 0.000 claims description 84
- 230000006870 function Effects 0.000 claims description 45
- 230000001788 irregular Effects 0.000 claims description 45
- 239000011159 matrix material Substances 0.000 description 117
- 238000004091 panning Methods 0.000 description 75
- 238000010586 diagram Methods 0.000 description 26
- 238000000605 extraction Methods 0.000 description 19
- 238000012545 processing Methods 0.000 description 15
- 239000013598 vector Substances 0.000 description 12
- 230000008569 process Effects 0.000 description 11
- 230000011664 signaling Effects 0.000 description 6
- 235000009508 confectionery Nutrition 0.000 description 5
- 238000013461 design Methods 0.000 description 5
- 238000013459 approach Methods 0.000 description 4
- 230000004807 localization Effects 0.000 description 4
- 238000013507 mapping Methods 0.000 description 4
- 238000004321 preservation Methods 0.000 description 4
- 230000009466 transformation Effects 0.000 description 4
- 230000008901 benefit Effects 0.000 description 3
- 238000004891 communication Methods 0.000 description 3
- 238000013500 data storage Methods 0.000 description 3
- 230000000694 effects Effects 0.000 description 3
- 238000001914 filtration Methods 0.000 description 3
- 239000000463 material Substances 0.000 description 3
- 230000004044 response Effects 0.000 description 3
- 230000005236 sound signal Effects 0.000 description 3
- 238000004458 analytical method Methods 0.000 description 2
- 238000003491 array Methods 0.000 description 2
- 230000015572 biosynthetic process Effects 0.000 description 2
- 238000006243 chemical reaction Methods 0.000 description 2
- 238000004590 computer program Methods 0.000 description 2
- 230000001419 dependent effect Effects 0.000 description 2
- 210000005069 ears Anatomy 0.000 description 2
- 238000005516 engineering process Methods 0.000 description 2
- 239000000835 fiber Substances 0.000 description 2
- 238000007667 floating Methods 0.000 description 2
- 230000007246 mechanism Effects 0.000 description 2
- 238000012986 modification Methods 0.000 description 2
- 230000004048 modification Effects 0.000 description 2
- 230000003287 optical effect Effects 0.000 description 2
- 230000008447 perception Effects 0.000 description 2
- 238000003786 synthesis reaction Methods 0.000 description 2
- 230000001052 transient effect Effects 0.000 description 2
- QNRATNLHPGXHMA-XZHTYLCXSA-N (r)-(6-ethoxyquinolin-4-yl)-[(2s,4s,5r)-5-ethyl-1-azabicyclo[2.2.2]octan-2-yl]methanol;hydrochloride Chemical compound Cl.C([C@H]([C@H](C1)CC)C2)CN1[C@@H]2[C@H](O)C1=CC=NC2=CC=C(OCC)C=C21 QNRATNLHPGXHMA-XZHTYLCXSA-N 0.000 description 1
- 230000009471 action Effects 0.000 description 1
- 239000000654 additive Substances 0.000 description 1
- 230000000996 additive effect Effects 0.000 description 1
- 230000003190 augmentative effect Effects 0.000 description 1
- 230000009286 beneficial effect Effects 0.000 description 1
- 230000005540 biological transmission Effects 0.000 description 1
- 238000000354 decomposition reaction Methods 0.000 description 1
- 230000007812 deficiency Effects 0.000 description 1
- 238000011156 evaluation Methods 0.000 description 1
- 238000012544 monitoring process Methods 0.000 description 1
- 238000005457 optimization Methods 0.000 description 1
- 238000011160 research Methods 0.000 description 1
- 238000012360 testing method Methods 0.000 description 1
- 238000012546 transfer Methods 0.000 description 1
- 238000000844 transformation Methods 0.000 description 1
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S5/00—Pseudo-stereo systems, e.g. in which additional channel signals are derived from monophonic signals by means of phase shifting, time delay or reverberation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/301—Automatic calibration of stereophonic sound system, e.g. with test microphone
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
Definitions
- This disclosure relates to audio rendering and, more specifically, rendering of spherical harmonic coefficients.
- a higher order ambisonics (HOA) signal (often represented by a plurality of spherical harmonic coefficients (SHC) or other hierarchical elements) is a three-dimensional representation of a sound field.
- This HOA or SHC representation may represent this sound field in a manner that is independent of the local speaker geometry used to playback a multi-channel audio signal rendered from this SHC signal.
- This SHC signal may also facilitate backwards compatibility as this SHC signal may be rendered to well-known and highly adopted multi-channel formats, such as a 5.1 audio channel format or a 7.1 audio channel format.
- the SHC representation therefore enables a better representation of a sound field that also accommodates backward compatibility.
- WO 0182651 A1 discloses techniques of making a recording or transmitting a sound filed from either multiple monaural or directional sound signals that reproduce through multiple discrete loud speakers a sound filed with spatial harmonics that substantially exactly match those of the original sound field.
- US 2011 305344 A1 discloses a method and apparatus to encode audio with spatial information in a manner that does not depend on exhibition set up.
- an audio renderer that suits a particular local speaker geometry. While the SHC may accommodate well-known multi-channel speaker formats, commonly the end-user listener does not properly place or locate the speakers in the manner required by these multi-channel formats, resulting in irregular speaker geometries.
- the techniques described in this disclosure may determine the local speaker geometry and then determine a renderer for rendering the SHC signals based on this local speaker geometry.
- the rendering device may select from among a number of different renderers, e.g., a mono renderer, a stereo renderer, a horizontal only renderer or a three-dimensional renderer, and generate this renderer based on the local speaker geometry. This renderer may account for irregular speaker geometries and thereby facilitate better reproduction of the sound field despite irregular speaker geometries in comparison to a regular renderer designed for regular speaker geometries.
- the techniques may render to a uniform speaker geometry, which may be referred to as a virtual speaker geometry, so as to maintain invertibility and recover the SHC.
- the techniques may then perform various operations to project these virtual speakers to different horizontal planes (which may a different elevation than the horizontal plane on which the virtual speaker was originally located).
- the techniques may enable a device to generate a renderer that maps these projected virtual speakers to different physical speakers arranged in an irregular speaker geometry. Projecting these virtual speakers in this manner may facilitate better reproduction of the sound field.
- surround sound formats include the popular 5.1 format (which includes the following six channels: front left (FL), front right (FR), center or front center, back left or surround left, back right or surround right, and low frequency effects (LFE)), the growing 7.1 format, and the upcoming 22.2 format (e.g., for use with the Ultra High Definition Television standard). Further examples include formats for a spherical harmonic array.
- the input to a future MPEG encoder (which may generally be developed in response to a ISO/IEC JTC1/SC29/WG11/N13411 document, entitled “Call for Proposals for 3D Audio,” dated January 2013 and published at the convention in Geneva, Switzerl and) is optionally one of three possible formats: (i) traditional channel-based audio, which is meant to be played through loudspeakers at pre-specified positions; (ii) object-based audio, which involves discrete pulse-code-modulation (PCM) data for single audio objects with associated metadata containing their location coordinates (amongst other information); and (iii) scene-based audio, which involves representing the sound field using coefficients of spherical harmonic basis functions (also called “spherical harmonic coefficients" or SHC).
- PCM pulse-code-modulation
- a hierarchical set of elements may be used to represent a sound field.
- the hierarchical set of elements may refer to a set of elements in which the elements are ordered such that a basic set of lower-ordered elements provides a full representation of the modeled sound field. As the set is extended to include higher-order elements, the representation becomes more detailed.
- SHC spherical harmonic coefficients
- k ⁇ c , c is the speed of sound ( ⁇ 343 m/s), ⁇ r r , ⁇ r , ⁇ r ⁇ is a point of reference (or observation point), j n ( ⁇ ) is the spherical Bessel function of order n, and Y n m ⁇ r ⁇ r are the spherical harmonic basis functions of order n and suborder m.
- the term in square brackets is a frequency-domain representation of the signal (i.e., S ( ⁇ , r r , ⁇ r , ⁇ r )) which can be approximated by various time-frequency transformations, such as the discrete Fourier transform (DFT), the discrete cosine transform (DCT), or a wavelet transform.
- DFT discrete Fourier transform
- DCT discrete cosine transform
- wavelet transform a frequency-domain representation of the signal
- hierarchical sets include sets of wavelet transform coefficients and other sets of coefficients of multiresolution basis functions.
- the spherical harmonic basis functions are shown in three-dimensional coordinate space with both the order and the suborder shown.
- the SHC A n m k can either be physically acquired (e.g., recorded) by various microphone array configurations or, alternatively, they can be derived from channel-based or object-based descriptions of the sound field.
- the former represents scene-based audio input to an encoder.
- a fourth-order representation involving 1+2 4 (25, and hence fourth order) coefficients may be used.
- a n m k g ⁇ ⁇ 4 ⁇ ik h n 2 kr s Y n m ⁇ ⁇ s ⁇ s , where i is ⁇ 1 , h n 2 ⁇ is the spherical Hankel function (of the second kind) of order n, and ⁇ r s , ⁇ s , ⁇ s ⁇ is the location of the object.
- Knowing the source energy g ( ⁇ ) as a function of frequency allows us to convert each PCM object and its location into the SHC A n m k . Further, it can be shown (since the above is a linear and orthogonal decomposition) that the A n m k coefficients for each object are additive. In this manner, a multitude of PCM objects can be represented by the A n m k coefficients (e.g., as a sum of the coefficient vectors for the individual objects).
- these coefficients contain information about the sound field (the pressure as a function of 3D coordinates), and the above represents the transformation from individual objects to a representation of the overall sound field, in the vicinity of the observation point ⁇ r r , ⁇ r , ⁇ r ⁇ .
- the remaining figures are described below in the context of object-based and SHC-based audio coding.
- FIG. 3 is a diagram illustrating a system 20 that may perform various aspects of the techniques described in this disclosure.
- the system 20 includes a content creator 22 and a content consumer 24.
- the content creator 22 may represent a movie studio or other entity that may generate multi-channel audio content for consumption by content consumers, such as the content consumers 24. Often, this content creator generates audio content in conjunction with video content.
- the content consumer 24 represents an individual that owns or has access to an audio playback system 32, which may refer to any form of audio playback system capable of playing back multi-channel audio content.
- the content consumer 24 includes an audio playback system 32.
- the content creator 22 includes an audio renderer 28 and an audio editing system 30.
- the audio renderer 26 may represent an audio processing unit that renders or otherwise generates speaker feeds (which may also be referred to as "loudspeaker feeds," “speaker signals,” or “loudspeaker signals”). Each speaker feed may correspond to a speaker feed that reproduces sound for a particular channel of a multi-channel audio system.
- the renderer 28 may render speaker feeds for conventional 5.1, 7.1 or 22.2 surround sound formats, generating a speaker feed for each of the 5, 7 or 22 speakers in the 5.1, 7.1 or 22.2 surround sound speaker systems.
- the renderer 28 may be configured to render speaker feeds from source spherical harmonic coefficients for any speaker configuration having any number of speakers, given the properties of source spherical harmonic coefficients discussed above.
- the renderer 28 may, in this manner, generate a number of speaker feeds, which are denoted in FIG. 3 as speaker feeds 29.
- the content creator may, during the editing process, render spherical harmonic coefficients 27 ("SHC 27"), listening to the rendered speaker feeds in an attempt to identify aspects of the sound field that do not have high fidelity or that do not provide a convincing surround sound experience.
- the content creator 22 may then edit source spherical harmonic coefficients (often indirectly through manipulation of different objects from which the source spherical harmonic coefficients may be derived in the manner described above).
- the content creator 22 may employ an audio editing system 30 to edit the spherical harmonic coefficients 27.
- the audio editing system 30 represents any system capable of editing audio data and outputting this audio data as one or more source spherical harmonic coefficients.
- the content creator 22 may generate the bitstream 31 based on the spherical harmonic coefficients 27. That is, the content creator 22 includes a bitstream generation device 36, which may represent any device capable of generating the bitstream 31. In some instances, the bitstream generation device 36 may represent an encoder that bandwidth compresses (by, as one example, entropy encoding) the spherical harmonic coefficients 27 and that arranges the bandwidth compressed version of the spherical harmonic coefficients 27 in an accepted format to form the bitstream 31.
- the bitstream generation device 36 may represent an audio encoder (possibly, one that complies with a known audio coding standard, such as MPEG surround, or a derivative thereof) that encodes the multi-channel audio content 29 using, as one example, processes similar to those of conventional audio surround sound encoding processes to compress the multi-channel audio content or derivatives thereof.
- the compressed multi-channel audio content 29 may then be entropy encoded or coded in some other way to bandwidth compress content 29 and arranged in accordance with an agreed upon format to form the bitstream 31.
- the content creator 22 may transmit the bitstream 31 to the content consumer 24.
- the content creator 22 may output the bitstream 31 to an intermediate device positioned between the content creator 22 and the content consumer 24.
- This intermediate device may store the bitstream 31 for later delivery to the content consumer 24, which may request this bitstream.
- the intermediate device may comprise a file server, a web server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smart phone, or any other device capable of storing the bitstream 31 for later retrieval by an audio decoder.
- the content creator 22 may store the bitstream 31 to a storage medium, such as a compact disc, a digital video disc, a high definition video disc or other storage mediums, most of which are capable of being read by a computer and therefore may be referred to as computer-readable storage mediums.
- a storage medium such as a compact disc, a digital video disc, a high definition video disc or other storage mediums, most of which are capable of being read by a computer and therefore may be referred to as computer-readable storage mediums.
- the transmission channel may refer to those channels by which content stored to these mediums are transmitted (and may include retail stores and other store-based delivery mechanism). In any event, the techniques of this disclosure should not therefore be limited in this respect to the example of FIG. 3 .
- the content consumer 24 includes an audio playback system 32.
- the audio playback system 32 may represent any audio playback system capable of playing back multi-channel audio data.
- the audio playback system 32 may include a number of different renderers.
- the audio playback system 32 may also include a renderer determination unit 40 that may represent a unit configured to determine or otherwise select an audio renderer 34 from among a plurality of audio renderers. In some instances, the renderer determination unit 40 may select the renderer 34 from a number of pre-defined renderers. In other instances, the renderer determination unit 40 may dynamically determine the audio renderer 34 based on local speaker geometry information 41.
- the local speaker geometry information 41 may specify a location of each speaker coupled to the audio playback system 32 relative to the audio playback system 32, a listener, or any other identifiable region or location. Often, a listener may interface with the audio playback system 32 via a graphical user interface (GUI) or other form of interface to input the local speaker geometry information 41. In some instances, the audio playback system 32 may automatically (meaning, in this example, without requiring any listener intervention) determine the local speaker geometry information 41 often by emitting certain tones and measuring the tones via a microphone coupled to the audio playback system 32.
- GUI graphical user interface
- the audio playback system 32 may further include an extraction device 38.
- the extraction device 38 may represent any device capable of extracting spherical harmonic coefficients 27' ("SHC 27'," which may represent a modified form of or a duplicate of spherical harmonic coefficients 27) through a process that may generally be reciprocal to that of the bitstream generation device 36.
- the audio playback system 32 may receive the spherical harmonic coefficients 27' and invoke the extraction device 38 to extract the SHC 27' and, if specified or available, the audio rendering information 39.
- each of the above renderers 34 may provide for a different form of rendering, where the different forms of rendering may include one or more of the various ways of performing vector-base amplitude panning (VBAP), one or more of the various ways of performing distance based amplitude panning (DBAP), one or more of the various ways of performing simple panning, one or more of the various ways of performing near field compensation (NFC) filtering and/or one or more of the various ways of performing wave field synthesis.
- the selected renderer 34 may then render spherical harmonic coefficients 27' to generate a number of speaker feeds 35 (corresponding to the number of loudspeakers electrically or possibly wirelessly coupled to audio playback system 32, which are not shown in the example of FIG. 3 for ease of illustration purposes).
- the audio playback system 32 may select any one of a plurality of audio renderers and may be configured to select one or more of audio renderers depending on the source from which the bitstream 31 is received (such as a DVD player, a Blu-ray player, a smartphone, a tablet computer, a gaming system, and a television to provide a few examples). While any one of the audio renderers may be selected, often the audio renderer used when creating the content provides for a better (and possibly the best) form of rendering due to the fact that the content was created by the content creator 22 using this one of audio renderers, i.e., the audio renderer 28 in the example of FIG. 3 . Selecting the one of the audio renderers 34 having a rendering form that is the same as or at least close to the rendering form of the local speaker geometry may provide for a better representation of the sound field that may result in a better surround sound experience for the content consumer 24.
- the source from which the bitstream 31 is received such as a DVD player, a Blu-ray player, a smartphone, a
- the bitstream generation device may generate the bitstream 31 to include the audio rendering information 39 ("audio rendering info 39").
- the audio rendering information 39 may include a signal value identifying an audio renderer used when generating the multi-channel audio content, i.e., the audio renderer 28 in the example of FIG. 3 .
- the signal value includes a matrix used to render spherical harmonic coefficients to a plurality of speaker feeds.
- the signal value includes two or more bits that define an index that indicates that the bitstream includes a matrix used to render spherical harmonic coefficients to a plurality of speaker feeds.
- the signal value when an index is used, further includes two or more bits that define a number of rows of the matrix included in the bitstream and two or more bits that define a number of columns of the matrix included in the bitstream.
- the signal value specifies a rendering algorithm used to render spherical harmonic coefficients to a plurality of speaker feeds.
- the rendering algorithm may include a matrix that is known to both the bitstream generation device 36 and the extraction device 38. That is, the rendering algorithm may include application of a matrix in addition to other rendering steps, such as panning (e.g., VBAP, DBAP or simple panning) or NFC filtering.
- the signal value includes two or more bits that define an index associated with one of a plurality of matrices used to render spherical harmonic coefficients to a plurality of speaker feeds.
- both the bitstream generation device 36 and the extraction device 38 may be configured with information indicating the plurality of matrices and the order of the plurality of matrices such that the index may uniquely identify a particular one of the plurality of matrices.
- the bitstream generation device 36 may specify data in the bitstream 31 defining the plurality of matrices and/or the order of the plurality of matrices such that the index may uniquely identify a particular one of the plurality of matrices.
- the signal value includes two or more bits that define an index associated with one of a plurality of rendering algorithms used to render spherical harmonic coefficients to a plurality of speaker feeds.
- both the bitstream generation device 36 and the extraction device 38 may be configured with information indicating the plurality of rendering algorithms and the order of the plurality of rendering algorithms such that the index may uniquely identify a particular one of the plurality of matrices.
- the bitstream generation device 36 may specify data in the bitstream 31 defining the plurality of matrices and/or the order of the plurality of matrices such that the index may uniquely identify a particular one of the plurality of matrices.
- bitstream generation device 36 specifies audio rendering information 39 on a per audio frame basis in the bitstream. In other instances, bitstream generation device 36 specifies the audio rendering information 39 a single time in the bitstream.
- the extraction device 38 may then determine audio rendering information 39 specified in the bitstream. Based on the signal value included in the audio rendering information 39, the audio playback system 32 may render a plurality of speaker feeds 35 based on the audio rendering information 39. As noted above, the signal value may in some instances include a matrix used to render spherical harmonic coefficients to a plurality of speaker feeds. In this case, the audio playback system 32 may configure one of the audio renderers 34 with the matrix, using this one of the audio renderers 34 to render the speaker feeds 35 based on the matrix.
- the signal value includes two or more bits that define an index that indicates that the bitstream includes a matrix used to render the spherical harmonic coefficients 27' to the speaker feeds 35.
- the extraction device 38 may parse the matrix from the bitstream in response to the index, whereupon the audio playback system 32 may configure one of the audio renderers 34 with the parsed matrix and invoke this one of the renderers 34 to render the speaker feeds 35.
- the extraction device 38 may parse the matrix from the bitstream in response to the index and based on the two or more bits that define a number of rows and the two or more bits that define the number of columns in the manner described above.
- the signal value specifies a rendering algorithm used to render the spherical harmonic coefficients 27' to the speaker feeds 35.
- some or all of the audio renderers 34 may perform these rendering algorithms.
- the audio playback device 32 may then utilize the specified rendering algorithm, e.g., one of the audio renderers 34, to render the speaker feeds 35 from the spherical harmonic coefficients 27'.
- the audio playback system 32 may render the speaker feeds 35 from the spherical harmonic coefficients 27' using the one of the audio renderers 34 associated with the index.
- the audio playback system 32 may render the speaker feeds 35 from the spherical harmonic coefficients 27' using one of the audio renderers 34 associated with the index.
- the extraction device 38 may determine the audio rendering information 39 on a per audio frame basis or a single time.
- the techniques may potentially result in better reproduction of the multi-channel audio content 35 and according to the manner in which the content creator 22 intended the multi-channel audio content 35 to be reproduced. As a result, the techniques may provide for a more immersive surround sound or multi-channel audio experience.
- the audio rendering information 39 may be specified as metadata separate from the bitstream or, in other words, as side information separate from the bitstream.
- the bitstream generation device 36 may generate this audio rendering information 39 separate from the bitstream 31 so as to maintain bitstream compatibility with (and thereby enable successful parsing by) those extraction devices that do not support the techniques described in this disclosure. Accordingly, while described as being specified in the bitstream, the techniques may allow for other ways by which to specify the audio rendering information 39 separate from the bitstream 31.
- the techniques may enable the bitstream generation device 36 to specify a portion of the audio rendering information 39 in the bitstream 31 and a portion of the audio rendering information 39 as metadata separate from the bitstream 31.
- the bitstream generation device 36 may specify the index identifying the matrix in the bitstream 31, where a table specifying a plurality of matrixes that includes the identified matrix may be specified as metadata separate from the bitstream.
- the audio playback system 32 may then determine the audio rendering information 39 from the bitstream 31 in the form of the index and from the metadata specified separately from the bitstream 31.
- the audio playback system 32 may, in some instances, be configured to download or otherwise retrieve the table and any other metadata from a pre-configured or configured server (most likely hosted by the manufacturer of the audio playback system 32 or a standards body).
- the content consumer 24 does not properly configure the speakers according to a specified (typically by the surround sound audio format body) geometry. Often, the content consumer 24 does not place the speakers at a fixed height and in precisely the specified location relative to the listener. The content consumer 24 may be unable to place speakers in these location or be unaware that there are even specified locations at which to place speakers to achieve a suitable surround sound experience.
- SHC enables a more flexible arrangement of speakers given that the SHC represent the sound field in two or three dimensions, meaning that from the SHC, an acceptable (or at least better sounding, in comparison to that of non-SHC audio systems) reproduction of sound field may be provided by speakers configured in most any speaker geometry.
- the techniques described in this disclosure may enable the renderer determination unit 40 not only to select a standard renderer using the audio rendering information 39 in the manner described above but to dynamically generate a renderer based on the local speaker geometry information 41.
- the techniques may provide for at least four exemplary ways by which to generate a renderer 34 tailored to a specific local speaker geometry specified by the local speaker geometry information 41.
- These four ways may include a way by which to generate a mono renderer 34, a stereo renderer 34, a horizontal multi-channel renderer 34 (where, for example, "horizontal multi-channel” refers to a multi-channel speaker configuration having more than two speakers in which all of the speakers are generally on or near the same horizontal plane), and a three-dimensional (3D) renderer 34 (where a three-dimensional renderer may render for multiple horizontal planes of speakers).
- the audio determination unit 40 may select renderer 34 based on the audio rendering information 39 or the local speaker geometry information 41.
- the content consumer 24 may specify a preference that the renderer determination unit 40 select the renderer 34 based on the audio rendering information 39 (when present, as this may not be present in all bitstreams) and, when not present, determine (or select if previously determined) the renderer 34 based on the local speaker geometry information 41.
- the content consumer 24 may specify a preference that the renderer determination unit 40 determine (or select if previously determined) the renderer 34 based on the local speaker geometry information 41 without ever considering the audio rendering information 39 during the selection of the renderer 34.
- any number of preferences may be specified for configuring how the renderer determination unit 40 selects the renderer 34 based on the audio rendering information 39 and/or the local speaker geometry 41. Accordingly, the techniques should not be limited in this respect to the two exemplary alternatives discussed above.
- the renderer determination unit 40 may first categorize the local speaker geometry into one of the four categories briefly mentioned above. That is, the renderer determination unit 40 may first determine whether the local speaker geometry information 41 indicates that the local speaker geometry generally conforms to a mono speaker geometry, a stereo speaker geometry, a horizontal multi-channel speaker geometry having three or more speakers on the same horizontal plane or a three-dimensional multi-channel speaker geometry having three or more speakers, two of which are on different horizontal planes (often separated by some threshold height).
- the renderer determination unit 40 may generate one of a mono renderer, a stereo renderer, a horizontal multi-channel renderer and a three-dimensional multi-channel renderer. The renderer determination unit 40 may then provide this renderer 34 for use by the audio playback system 32, whereupon the audio playback system 32 may render the SHC 27' in the manner described above to generate the multi-channel audio data 35.
- the techniques may enable audio playback system 32 to determine a local speaker geometry of one or more speakers used for playback of spherical harmonic coefficients representative of a sound field, and determine a two dimensional or three dimensional renderer based on the local speaker geometry.
- the audio playback system 32 may render the spherical harmonic coefficients using the determined renderer to generate multi-channel audio data.
- the audio playback system 32 may, when determining the renderer based on the local speaker geometry, determine a stereo renderer when the local speaker geometry conforms to a stereo speaker geometry.
- the audio playback system 32 may, when determining the renderer based on the local speaker geometry, determine a horizontal multi-channel renderer when the local speaker geometry conforms to horizontal multi-channel speaker geometry having more than two speakers.
- the audio playback system 32 may, when determining the renderer based on the local speaker geometry, determine a three-dimensional multi-channel renderer when the local speaker geometry conforms a three-dimensional multi-channel speaker geometry having more than two speakers on more than one horizontal plane.
- the audio playback system 32 may, when determining the local speaker geometry of the one or more speakers, receive input from a listener specifying local speaker geometry information describing the local speaker geometry.
- the audio playback system 32 may, when determining the local speaker geometry of the one or more speakers, receive input via a graphical user interface from a listener specifying local speaker geometry information describing the local speaker geometry.
- the audio playback system 32 may, when determining the local speaker geometry of the one or more speakers, automatically determine local speaker geometry information describing the local speaker geometry.
- a Higher Order Ambisonics signal such as SHC 27, is a representation of a three-dimensional sound field using spherical harmonic basis functions, where at least one of the spherical harmonic basis functions are associated with a spherical basis function having an order greater than one.
- This representation may provide an ideal sound format as it is independent of end user speaker geometry and, as a result, the representation may be rendered to any geometry at the content consumer without prior knowledge on the encoding side.
- the final speaker signals may then be derived by linear combination of the spherical harmonic coefficients, which generally represent a polar pattern pointing in the direction of that particular speaker.
- HOA or SHC renderer generation system/algorithm detects what type of speaker geometry is in use: mono, stereo, horizontal, three-dimensional or flagged as a known geometry/renderer matrix.
- FIG. 4 is a block diagram illustrating the renderer determination unit 40 of FIG. 3 in more detail.
- the renderer determination unit 40 may include a renderer selection unit 42, a layout determination unit 44, and a renderer generation unit 46.
- the renderer selection unit 42 may represent a unit configured to select a pre-defined renderer based on the rendering information 39 or select the renderer specified in the rendering information 39, outputting this selected or specified renderer as the renderer 34.
- the layout determination unit 44 may represent a unit configured to categorize a local speaker geometry based on local speaker geometry information 41.
- the layout determination unit 44 may categorize the local speaker geometry to one of the four categories described above: 1) mono speaker geometry, 2) stereo speaker geometry, 3) a horizontal multi-channel speaker geometry, and 4) a three-dimensional multi-channel speaker geometry.
- the layout determination unit 44 may pass categorization information 45 to the renderer generation unit 46 that indicates to which of the four categories the local speaker geometry most conforms.
- the renderer generation unit 46 may represent a unit configured to generate a renderer 34 based on the categorization information 45 and the local speaker geometry information 41.
- the renderer generation unit 46 may include a mono renderer generation unit 48D, a stereo renderer generation unit 48A, a horizontal renderer generation unit 48B, and a three-dimensional (3D) renderer generation unit 48C.
- the mono renderer generation unit 48A may represent a unit configured to generate a mono renderer based on the local speaker geometry information 41.
- the stereo renderer generation unit 48A may represent a unit configured to generate a stereo renderer based on the local speaker geometry information 41. The process employed by the stereo renderer generation unit 48A is described in more detail below with respect to the example of FIG. 6 .
- the horizontal renderer generation unit 48B may represent a unit configured to generate a horizontal multi-channel renderer based on the local speaker geometry information 41. The process employed by the horizontal renderer generation unit 48B is described in more detail below with respect to the example of FIG. 7 .
- the 3D renderer generation unit 48C may represent a unit configured to generate a 3D multi-channel renderer based on the local speaker geometry information 41. The process employed by the horizontal renderer generation unit 48B is described in more detail below with respect to the example of FIGS. 8 and 9 .
- FIG. 5 is a flow diagram illustrating exemplary operation of the renderer determination unit 40 shown in the example of FIG. 4 in performing various aspects of the techniques described in this disclosure.
- the flow diagram of FIG. 5 generally outlines the operations performed by the renderer determination unit 40 described above with respect to FIG. 4 , except for some minor notation changes.
- the renderer flag refers to a specific example of the audio rendering information 39.
- the "SHC order” refers to the maximum order of the SHC.
- the "stereo renderer” may refer to the stereo renderer generation unit 48A.
- the “horizontal renderer” may refer to the horizontal renderer generation unit 48B.
- the “3D renderer” may refer to the 3D renderer generation unit 48C.
- the “Renderer Matrix” may refer to the renderer selection unit 42.
- the renderer selection unit 42 may determine whether the render flag, which may be denoted as the render flag 39', is present in the bitstream 31 (or other side channel information associated with the bitstream 31) (60). When the renderer flag 39' is present in the bitstream 31 ("YES" 60), the renderer selection unit 42 may select the renderer from a potential plurality of renderers based on the renderer flag 39' and output the selected renderer as the renderer 34 (62, 64).
- the renderer selection unit 42 may select the renderer from a potential plurality of renderers based on the renderer flag 39' and output the selected renderer as the renderer 34 (62, 64).
- the renderer selection unit 42 may invoke the renderer determination unit 40, which may determine the local speaker geometry information 41. Based on the local speaker geometry information 41, the renderer determination unit 40 may invoke one of the mono renderer determination unit 48D, the stereo renderer determination unit 48A, the horizontal renderer determination unit 48B or the 3D renderer determination unit 48C.
- the renderer determination unit 40 may invoke the mono renderer determination unit 48D, which may determine a mono renderer (based potentially on the SHC order) and output the mono renderer as the renderer 34 (66, 64).
- the render determination unit 40 may invoke the stereo renderer determination unit 48A, which may determine a stereo render (based potentially on the SHC order) and output the stereo renderer as the renderer 34 (68, 64).
- the render determination unit 40 may invoke the horizontal renderer determination unit 48B, which may determine a horizontal render (based potentially on the SHC order) and output the horizontal renderer as the renderer 34 (70, 64).
- the render determination unit 40 may invoke the 3D renderer determination unit 48C, which may determine a 3D render (based potentially on the SHC order) and output the 3D renderer as the renderer 34 (72, 64).
- the techniques may enable the renderer determination unit 40 to determine a local speaker geometry of one or more speakers used for playback of spherical harmonic coefficients representative of a sound field, and determining a two-dimensional or three-dimensional renderer based on the local speaker geometry.
- FIG. 6 is a flow diagram illustrating exemplary operation of the stereo renderer generation unit 48A shown in the example of FIG. 4 .
- the stereo renderer generation unit 48A may receive the local speaker geometry information 41 (100) and then determine angular distances between the speakers relative to a listener position in what may be considered as the "sweet spot" for a given speaker geometry (102).
- the stereo renderer generation unit 48A may then calculate a highest allowed order, limited by the HOA/SHC order of the spherical harmonic coefficients (104).
- the stereo renderer generation unit 48A may next generate equal spaced azimuths based on the determined allowed order (106).
- the stereo renderer generation unit 48A may then sample the spherical basis functions at the locations of virtual or real speakers forming the two dimensional (2D) renderer.
- the stereo renderer generation unit 48A may then perform the pseudo-inverse (understood in the context of matrix mathematics) of this 2D renderer (108).
- this 2D renderer may be represented by the following matrix : h 0 2 kr 1 Y 0 0 ⁇ ⁇ 1 ⁇ 1 h 0 2 kr 2 Y 0 0 ⁇ ⁇ 2 ⁇ 2 . . . h 1 2 kr 1 Y 1 1 ⁇ ⁇ 1 ⁇ 1 . . . . . . . . . .
- this matrix may be V rows by (n+1) 2 , wherein V denotes the number of virtual speakers and n denotes the SHC order.
- h n 2 ⁇ is the spherical Hankel function (of the second kind) of order n.
- Y n m ⁇ r ⁇ r are the spherical harmonic basis functions of order n and suborder m.
- ⁇ ⁇ r , ⁇ r ⁇ is a point of reference (or observation point) in terms of spherical coordinates.
- the stereo renderer generation unit 48A may then rotate the azimuth to the right position and to the left position generating two different 2D renderers (110, 112) and then combines them into a 2D renderer matrix (114).
- the stereo renderer generation unit 48A may then convert this 2D renderer matrix to a 3D renderer matrix (116) and zero pad the difference between the allowed order (denoted as order' in the example of FIG. 6 ) and the order, n (120).
- the stereo renderer generation unit 48A may then perform energy preservation with respect to the 3D renderer matrix (122), outputting this 3D renderer matrix (124).
- the techniques may enable the stereo renderer generation unit 48A to generate a stereo rendering matrix based on the SHC order and angular distance between the left and right speaker positions.
- the stereo renderer generation unit 48A may then rotate the front position of the rendering matrix to match the left and then right speaker positions and then combine these left and right matrixes to form the final rendering matrix.
- FIG. 7 is a flow diagram illustrating exemplary operation of the horizontal renderer generation unit 48B shown in the example of FIG. 4 .
- the horizontal renderer generation unit 48B may receive the local speaker geometry information 41 (130) and then find angular distances between the speakers relative to a listener position in what may be considered as the "sweet spot" for a given speaker geometry (132).
- the horizontal renderer generation unit 48B may then compute the minimum angular distance and the maximum angular distance, comparing the minimum angular distance to the maximum angular distance (134).
- the horizontal renderer generation unit 48B determines that the local speaker geometry is regular.
- the horizontal renderer generation unit 48B may determine that the local speaker geometry is irregular.
- the horizontal renderer generation unit 48B may calculate a highest allowed order, limited by the HOA/SHC order of the spherical harmonic coefficients, as describe above (136).
- the horizontal renderer generation unit 48B may next generate the pseudo-inverse of the 2D renderer (138), convert this pseudo-inverse of the 2D renderer to a 3D renderer (140), and zero pad the 3D renderer (142).
- the horizontal renderer generation unit 48B may calculate the highest allowed order, limited by the HOA/SHC order of the spherical harmonic coefficients, as describe above (144). The horizontal renderer generation unit 48B may then generate equal spaced azimuths based on the allowed order (146) to generate a 2D renderer. The horizontal renderer generation unit 48B may perform a pseudo inverse of the 2D renderer (148), and perform an optional windowing operation (150). In some instances, the horizontal renderer generation unit 48B may not perform the windowing operation.
- the horizontal renderer generation unit 48B may also pan gains placing equal azimuth to real azimuths (of the irregular speaker geometry, 152) and perform a matrix multiplication of the pseudo-inverse 2D renderer by the panned gains (154).
- the panning gain matrix may represent a vector base amplitude panning (VBAP) matrix of size RxV that perform VBAP, where V again represents the number of virtual speakers and the R represents the number of real speakers.
- the VBAP matrix may be specified as follows: VBAP MATRIX ⁇ 1 RxV .
- the multiplication may be expressed as follows: VBAP MATRIX ⁇ 1 RxV D ⁇ 1 Vx n + 1 2 .
- the horizontal renderer generation unit 48B may then convert the output of the matrix multiplication, which is a 2D renderer, to a 3D renderer (156) and then zero pad the 3D renderer, again as described above (158).
- the techniques may be performed with respect to any way by which to map virtual speakers to real speakers.
- the matrix may be denoted as a "virtual-to-real speaker mapping matrix" having a size of RxV.
- the multiplication may therefore be expressed more generally as: Virtual _ to _ Real _ Speaker _ Mapping _ Matrix ⁇ 1 R ⁇ V D ⁇ 1 V ⁇ n + 1 2 .
- This Virtual_to_Real_Speaker_Mapping_Matrix may represent any panning or other matrix that may map virtual speakers to real speakers, including one or more of the matrices for performing vector-base amplitude panning (VBAP), one or more of the matrices for performing distance based amplitude panning (DBAP), one or more of the matrices for performing simple panning, one or more of the matrices for performing near field compensation (NFC) filtering and/or one or more of the matrices for performing wave field synthesis.
- VBAP vector-base amplitude panning
- DBAP distance based amplitude panning
- NFC near field compensation
- the horizontal renderer generation unit 48B may perform energy preservation with respect to the regular 3D renderer or the irregular 3D renderer (160). In some examples but not all, the horizontal renderer generation unit 48B may perform an optimization based on spatial properties of the 3D renderer (162), outputting this optimized 3D or non-optimized 3D renderer (164).
- the system may therefore generally detect whether the geometry of speakers is regularly spaced or irregular and then creates a rendering matrix based on the pseudo-inverse or the AllRAD approach.
- the AllRAD approach is discussed in more detail in a paper by Franz Zotter et al., entitled “Comparison of energy-preserving and all-round Ambisonic decoders," presented during the AIA-DAGA in Merano, 18-21 March 2013 .
- a rendering matrix is generated by creating a renderer matrix for a regular horizontal based on the HOA order and angular distance between the left and right speaker positions. The front position of the rendering matrix is then rotated to match the left and then right speaker positions and then combined to form the final rendering matrix.
- FIG. 8A-8B are flow diagrams illustrating exemplary operation of the 3D renderer generation unit 48C shown in the example of FIG. 4 .
- the 3D renderer generation unit 48C may receive the local speaker geometry information 41 (170) and then determine spherical harmonics basis functions using the geometry of the first order and the geometry of the HOA/SHC order, n (172, 174).
- the 3D renderer generation unit 48C may then determine condition numbers for both the first order and less basis functions and those basis function associated with spherical basis functions greater than an order of one but less than or equal to n (176, 178).
- the 3D renderer generation unit 48C compares both of the condition values to a so-called "regular value" (180), which may represent a threshold having a value of, in some examples, 1.05.
- the 3D renderer generation unit 48C may determine that the local speaker geometry is regular (symmetrical in some sense from left to right and front to back with equally spaced speakers).
- the 3D renderer generation unit 48C may compare the condition value computed from the first order and less spherical basis functions to the regular value (182).
- this first order or less condition number is less than the regular value ("YES" 182)
- the 3D renderer generation unit 48C determines that the local speaker geometry is nearly regular (or, as shown in the example of FIG. 8 , "near regular").
- this first order or less condition number is not below the regular value (“NO" 182), the 3D renderer generation unit 48C determines that the local geometry is irregular.
- the 3D renderer generation unit 48C determines the 3D rendering matrix in a manner similar to that described above with respect to the regular 3D matrix determination set forth with respect to the example of FIG. 7 , except that the 3D renderer generation unit 48C generates this matrix for multiple horizontal planes of speakers (184).
- the 3D renderer generation unit 48C determines the 3D rendering matrix in a manner similar to that described above with respect to the irregular 2D matrix determination set forth with respect to the example of FIG. 7 , except that the 3D renderer generation unit 48C generates this matrix for multiple horizontal planes of speakers (186).
- the 3D renderer generation unit 48C determines the 3D rendering matrix in a manner similar to that described in U.S. Provisional Application U.S. 61/762,302 , entitled "PERFORMING 2D AND/OR 3D PANNING WITH RESPECT TO HIERARCHICAL SETS OF ELEMENTS,” except for slight modification to accommodate the more general nature of this determination (in that the techniques of this disclosure not limited to 22.2 speaker geometries as provided by way of example in this provisional application, 188).
- the 3D renderer generation unit 48C performs energy preservation with respect to the generated matrix (190) followed by, in some instances, optimizing this 3D rendering matrix based on spatial properties of the 3D rendering matrix (192).
- the 3D renderer generation unit 48C may then output this renderer as renderer 34 (194).
- the system may detect a regular (using pseudo-inverse), a near regular (that is regular at first order, but not the HOA order, and uses the AllRAD method) or finally an irregular (this is based on the above referenced U.S. Provisional Application U.S. 61/762,302 , but implemented as a potentially more general approach).
- the three-dimensional irregular process 188 may generate, where appropriate, 3D-VBAP triangulation for areas covered by speakers, high and low panning rings at the top bottom, a horizontal band, stretch factors etc. to create an enveloping renderer for irregular three-dimensional listening. All of the foregoing options may use energy preservation so that on the fly switching between geometries have the same perceived energy. Most irregular or near irregular options use an optional spherical harmonic windowing.
- FIG. 8B is a flow diagram illustrating operation of the 3D renderer determination unit 48C in determining a 3D renderer for playback of audio content via an irregular 3D local speaker geometry.
- the 3D renderer determination unit 48C may calculate the highest allowed order, limited by the HOA/SHC order of the spherical harmonic coefficients, as describe above (196).
- the 3D renderer generation unit 48C may then generate equal spaced azimuths based on the allowed order (198) to generate a 3D renderer.
- the 3D renderer generation unit 48C may perform a pseudo inverse of the 3D renderer (200), and perform an optional windowing operation (202). In some instances, the 3D renderer generation unit 48C may not perform the windowing operation
- the 3D renderer determination unit 48C may also perform lower hemisphere processing and upper hemisphere processing as described in more detail below with respect to FIG. 9 (204, 206).
- the 3D renderer determination unit 48C may, when performing the lower and upper hemisphere processing, generate hemisphere data (which is described below in more detail) indicating an amount to "stretch" the angular distances between real speakers, a 2D pan limit that may specify a panning limit to limit panning to certain threshold heights, and a horizontal band amount that may specify a horizontal height band in which speakers are considered in the same horizontal plane.
- the 3D renderer determination unit 48C may, in some instances, perform a 3D VBAP operation to construct 3D VBAP triangles while possibly "stretching" the local speaker geometry based on the hemisphere data from one or more of the lower hemisphere processing and the upper hemisphere processing (208).
- the 3D renderer determination unit 48C may stretch the real speaker angular distances within a given hemisphere to cover more space.
- the 3D renderer determination unit 48C may also identify 2D panning duplets for the lower hemisphere and upper hemisphere (210, 212), where these duplets identify two real speakers for each virtual speaker in the lower and upper hemisphere, respectively.
- the 3D renderer determination unit 48C may then loop through each regular geometry position identified when generating the equally spaced geometry, and based on the 2D panning duplets of the lower and upper hemisphere virtual speakers and the 3D VBAP triangles perform the following analysis (214).
- the 3D renderer determination unit 48C may determine whether the virtual speakers are within the upper and lower horizontal band values specified in the hemisphere data for the lower and upper hemispheres (216). When the virtual speakers are within these band values ("YES" 216), the 3D renderer determination unit 48C sets the elevation for these virtual speakers to zero (218). In other words, the 3D renderer determination unit 48C may identify virtual speakers in the lower hemisphere and upper hemisphere close to the middle horizontal plane bisecting the sphere around the so-called "sweet spot" and set the location of these virtual speakers to be on this horizontal plane.
- the 3D renderer determination unit 48C may perform 3D VBAP panning (or any other form or way by which to map virtual speakers to real speakers) (220) to generate the horizontal plane portion of the 3D renderer used to map the virtual speakers to the real speakers along the middle horizontal plane.
- the 3D renderer determination unit 48C may, when looping through each regular geometry position of the virtual speakers, may evaluate those virtual speakers in the lower hemisphere to determine whether these lower hemisphere virtual speakers are below a lower hemisphere elevation limit specified in the lower hemisphere data (222). The 3D renderer determination unit 48C may perform a similar evaluation with respect to the upper hemisphere virtual speakers to determine whether these upper hemisphere virtual speakers are above an upper hemisphere elevation limit specified in the upper hemisphere data (224).
- the 3D renderer determination unit 48C may perform panning with the identified lower duplets and the upper duplets, respectively (230, 232), effectively creating what may be referred to as a panning ring that clips the elevation of the virtual speaker and pans it between the real speakers above the horizontal band of the given hemisphere.
- the 3D renderer determination unit 48C may then combine the 3D VBAP panning matrix with the lower duplets panning matrix and the upper duplets panning matrix (234) and perform a matrix multiplication to matrix multiple the 3D renderer by the combined panning matrix (236).
- the 3D renderer determation unit 48C may then zero pad the difference between the allowed order (denoted as order' in the example of FIG. 6 ) and the order, n (238), outputting the irregular 3D renderer.
- the techniques may enable the renderer determination unit 40 to determine an allowed order of spherical basis functions to which spherical harmonic coefficients are associated, the allowed order identifying those of the spherical harmonic coefficients that are required to be rendered, and determine the renderer based on the determined allowed order.
- the renderer determination unit 40 the allowed order identifies those of the spherical harmonic coefficients that are required to be rendered given a determined local speaker geometry of speakers used for playback of the spherical harmonic coefficients.
- the renderer determination unit 40 may, when determining the renderer, determine the renderer such that the renderer only renders those of the spherical harmonic coefficients associated with spherical basis functions having an order less than or equal to the determined allowed order.
- the renderer determination unit 40 may the allowed order is less than the maximum order N of the spherical basis functions to which the spherical harmonic coefficients are associated.
- the renderer determination unit 40 may render the spherical harmonic coefficients using the determined renderer to generate multi-channel audio data.
- the renderer determination unit 40 may determine a local speaker geometry of one or more speakers used for playback of the spherical harmonic coefficients. When determining the renderer, the renderer determination unit 40 may determine the renderer based on the determined allowed order and the local speaker geometry.
- the renderer determination unit 40 may, when determining the renderer based on the local speaker geometry, determine a stereo renderer to render those of the spherical harmonic coefficients of the allowed order when the local speaker geometry conforms to a stereo speaker geometry.
- the renderer determination unit 40 may, when determining the renderer based on the local speaker geometry, determine a horizontal multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the local speaker geometry conforms to horizontal multi-channel speaker geometry having more than two speakers.
- the renderer determination unit 40 may, when determining the horizontal multi-channel renderer, determine an irregular horizontal multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the determined local speaker geometry indicates an irregular speaker geometry.
- the renderer determination unit 40 may, when determining the horizontal multi-channel renderer, determine a regular horizontal multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the determined local speaker geometry indicates a regular speaker geometry.
- the renderer determination unit 40 may, when determining the renderer based on the local speaker geometry, determine a three-dimensional multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the local speaker geometry conforms a three-dimensional multi-channel speaker geometry having more than two speakers on more than one horizontal plane.
- the renderer determination unit 40 may, when determining the three-dimensional multi-channel renderer, determine an irregular three-dimensional multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the determined local speaker geometry indicates an irregular speaker geometry.
- the renderer determination unit 40 may, when determining the three-dimensional multi-channel renderer, determine a near regular three-dimensional multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the determined local speaker geometry indicates a near regular speaker geometry.
- the renderer determination unit 40 may, when determining the three-dimensional multi-channel renderer, determine a regular three-dimensional multi-channel renderer to render those of the spherical harmonic coefficients of the allowed order when the determined local speaker geometry indicates a regular speaker geometry.
- the renderer determination unit 40 may, when determining the local speaker geometry of the one or more speakers, receive input from a listener specifying local speaker geometry information describing the local speaker geometry.
- the renderer determination unit 40 may, when determining the local speaker geometry of the one or more speakers, receive input via a graphical user interface from a listener specifying local speaker geometry information describing the local speaker geometry.
- the renderer determination unit 40 may, when determining the local speaker geometry of the one or more speakers, automatically determine local speaker geometry information describing the local speaker geometry.
- FIG. 9 is flow diagram illustrating exemplary operation of the 3D renderer generation unit 48C shown in the example of FIG. 4 in performing lower hemisphere processing and upper hemisphere processing when determining the irregular 3D renderer. More information regarding the process shown in the example of FIG. 9 can be found in the above referenced U.S. Provisional Application U.S. 61/762,302 . The process shown in the example of FIG. 9 may represent the lower or upper hemisphere processing described above with respect to FIG. 8B .
- the 3D renderer determination unit 48C may receive the local speaker geometry information 41 and determine first hemisphere real speaker locations (250, 252). The 3D renderer determination unit 48C may then duplicate the first hemisphere onto the opposite hemisphere and generate spherical harmonics using the geometry for HOA order (254, 256). The 3D renderer determination unit 48C may determine the condition number (258), which may indicate the regularity (or uniformity) of the local speaker geometry.
- the 3D renderer determination unit 48C may determine hemisphere data that includes a stretch value of zero, a 2D pan limit value of sign (90) and a horizontal band value of zero (262).
- the stretch value indicates an amount to "stretch" the angular distances between real speakers, the 2D pan limit that may specify a panning limit to limit panning to certain threshold heights, and a horizontal band amount that may specify a horizontal height band in which speakers are considered in the same horizontal plane.
- the 3D renderer determination unit 48C may also determine the angular distance of azimuths of the highest/lowest (depending on whether upper or lower hemisphere processing is performed) speakers (264). When the condition number is greater than a threshold number or the maximum absolute value elevation difference between the real speakers is not equal to 90 degrees ("NO" 260), the 3D renderer determination unit 48C may determine whether the maximum absolute value elevation difference is greater than zero and whether the maximum angular distance is less than a threshold angular distance (266). When the maximum absolute value elevation difference is greater than zero and the maximum angular distance is less than a threshold angular distance ("YES" 266), the 3D renderer determination unit 48C may then determine whether the maximum absolute value of the elevation is greater than 70 (268).
- the 3D renderer determination unit 48C determines hemisphere data that includes a stretch value equal to zero, a 2D pan limit equal to the sign of the maximum of the absolute value of the elevation, and a horizontal band value equal to zero (270).
- the 3D renderer determination unit 48C may determine hemisphere data that includes a stretch value equal to 10 minus the maximum absolute value of the elevations times 70 multiplied by 10, a 2D pan limit equal to the signed form of the maximum of the absolute value of the elevation minus the stretch value, and a horizontal band value equal to the signed form of the maximum absolute value of the elevations multiplied by 0.1 (272).
- the 3D renderer determination unit 48C may then determine whether the minimum of the absolute value of the elevations is equal to zero (274). When the minimum of the absolute value of the elevations is equal to zero ("YES" 274), the 3D renderer determination unit 48C may determine hemisphere data that includes a stretch value equal to zero, a 2D pan limit equal to zero, a horizontal band value equal to zero and a bound hemisphere value identifying indices of real speakers whose elevation are equal to zero (276).
- the 3D renderer determination unit 48C may determine the bound hemisphere value to be equal to the indices of lowest elevation speakers (278). The 3D renderer determination unit 48C may then determine whether the maximum absolute value of the elevations is greater than 70 (280).
- the 3D renderer determination unit 48C may determine hemisphere data that includes a stretch value equal to zero, a 2D pan limit equal to the signed form of the maximum of the absolute value of the elevations, and a horizontal band value equal to zero.
- the 3D renderer determination unit 48C may determine hemisphere data that includes a stretch value equal to 10 minus the maximum absolute value of the elevations times 70 multiplied by 10, a 2D pan limit equal to the signed form of the maximum of the absolute value of the elevation minus the stretch value, and a horizontal band value equal to the signed form of the maximum absolute value of the elevations multiplied by 0.1.
- FIG. 10 is a diagram illustrating a graph 299 in unit space showing how a stereo renderer may be generated in accordance with the techniques set forth in this disclosure.
- virtual speakers 300A-300H are arranged in a uniform geometry around the circumference of the horizontal plane bisecting the unit sphere (centered around the so-called "sweet spot").
- Physical speaker 302A and 302B are positioned at angular distances of 30 degrees and -30 degrees (respectively) as measured from the virtual speaker 300A.
- the stereo renderer determination unit 48A may determine a stereo renderer 34 that maps the virtual speaker 300A to the physical speakers 302A and 302B in the manner described above in more detail.
- FIG. 11 is a diagram illustrating a graph 304 in unit space showing how an irregular horizontal renderer may be generated in accordance with the techniques set forth in this disclosure.
- virtual speakers 300A-300H are arranged in a uniform geometry around the circumference of the horizontal plane bisecting the unit sphere (centered around the so-called "sweet spot").
- Physical speaker 302A-302D (“physical speakers 302") are positioned irregularly around the circumference of the horizontal plane.
- the horizontal renderer determination unit 48B may determine a irregular horizontal renderer 34 that maps the virtual speakers 300A-300H ("virtual speakers 300") to the physical speakers 302 in the manner described above in more detail.
- the horizontal renderer determination unit 48B may map the virtual speakers 300 to the two of the real speakers 302 closest to each one of the virtual speakers (in terms of having the smallest angular distance).
- the mapping is set forth in the following table: VIRTUAL SPEAKER REAL SPEAKER 300A 302A and 302B 300B 302B and 302C 300C 302B and 302C 300D 302C and 302D 300E 302C and 302D 300F 302C and 302D 300G 302D and 302A 300H 302D and 302A
- FIGS. 12A and 12B are diagrams illustrating graphs 306A and 306B showing how an irregular 3D renderer may be generated in accordance with the techniques described in this disclosure.
- the graph 306A includes stretched speaker locations 308A-308H ("stretched speaker locations 308").
- the 3D renderer determination unit 48C may identify hemisphere data having stretched real speaker locations 308 in the manner described above with respect to the example of FIG. 9 .
- the graph 306A also shows real speakers locations 302A-302H ("real speaker locations 302") relative to the stretched speaker locations 308, where in some instances the real speaker locations 302 are the same as the stretched speaker locations 308 and, in other instances, the real speaker locations 302 are not the same as the stretched speaker locations 308.
- Graph 306A also includes upper 2D pan interpolated line 310A representative of upper 2D panning duplets and lower 2D pan interpolated line 310B representative of the lower 2D panning duplets, each of which is described above in more detail with respect to the example of FIG. 8 .
- the 3D renderer determination unit 48C may determine the upper 2D pan interpolated line 310A based on the upper 2D pan duplets and the lower 2D pan interpolated line 310B based on the lower 2D pan duplets.
- the upper 2D pan interpolated line 310A may represent the upper 2D pan matrix
- the lower 2D pan interpolated line 310B may represent the lower 2D pan matrix.
- the graph adds virtual speakers 300 to the graph 306A, where the virtual speakers 300 are not formally denoted in the example of FIG. 12B to avoid unnecessary confusion with the lines demonstrating the mapping of the virtual speakers 300 to the stretched speaker locations 308.
- the 3D renderer determination unit 48C maps each one of the virtual speakers 300 to two or more of the stretched speaker locations 308 that have the closest angular distance to the virtual speaker, similar to that shown in the horizontal examples of FIGS. 11 and 12 .
- the irregular 3D renderer may therefore map the virtual speakers to the stretched speaker locations in the manner shown in the example of FIG. 12B .
- the techniques may therefore provide for, in a first example, a device, such as the audio playback system 32, comprising means for determining a local speaker geometry, e.g., the renderer determination unit 40, of one or more speakers used for playback of spherical harmonic coefficients representative of a sound field, and means for determining, e.g., the renderer determination unit 40, a two-dimensional or three-dimensional renderer based on the local speaker geometry.
- a device such as the audio playback system 32, comprising means for determining a local speaker geometry, e.g., the renderer determination unit 40, of one or more speakers used for playback of spherical harmonic coefficients representative of a sound field, and means for determining, e.g., the renderer determination unit 40, a two-dimensional or three-dimensional renderer based on the local speaker geometry.
- the device of the first example may further comprise means for rendering, e.g., the audio renderer 34, the spherical harmonic coefficients using the determined two-dimensional or three-dimensional renderer to generate multi-channel audio data.
- means for rendering e.g., the audio renderer 34, the spherical harmonic coefficients using the determined two-dimensional or three-dimensional renderer to generate multi-channel audio data.
- the device of the first example wherein the means for determining the two-dimensional or three-dimensional renderer based on the local speaker geometry may comprise means for, when the local speaker geometry conforms to a stereo speaker geometry, determining a two-dimensional stereo renderer, e.g., the stereo renderer generation unit 48A.
- the device of the first example wherein the means for determining the two-dimensional or three-dimensional renderer based on the local speaker geometry comprises means for, when the local speaker geometry conforms to horizontal multi-channel speaker geometry having more than two speakers, determining, a horizontal two-dimensional multi-channel renderer, e.g., the horizontal renderer generation unit 48B.
- the device of the fourth example wherein the means for determining the horizontal two-dimensional multi-channel renderer comprises means for determining an irregular horizontal two-dimensional multi-channel renderer when the determined local speaker geometry indicates an irregular speaker geometry, as described with respect to the example of FIG. 7 .
- the device of the fourth example wherein the means for determining the horizontal two-dimensional multi-channel renderer comprises means for determining a regular horizontal two-dimensional multi-channel renderer when the determined local speaker geometry indicates a regular speaker geometry, as described with respect to the example of FIG. 7 .
- the device of the first example wherein the means for determining the two-dimensional or three-dimensional renderer based on the local speaker geometry comprises means for, when the local speaker geometry conforms a three-dimensional multi-channel speaker geometry having more than two speakers on more than one horizontal plane, determining a three-dimensional multi-channel renderer, e.g., the 3D renderer generation unit 48C.
- the device of the seventh example wherein the means for determining the three-dimensional multi-channel renderer comprises means for determining an irregular three-dimensional multi-channel renderer when the determined local speaker geometry indicates an irregular speaker geometry, as described above with respect to the examples of FIGS. 8A and 8B .
- the device of the seventh example wherein the means for determining the three-dimensional multi-channel renderer comprises means for determining a near regular three-dimensional multi-channel renderer when the determined local speaker geometry indicates a near regular speaker geometry, as described above with respect to the example of FIGS. 8A .
- the device of the seventh example wherein the means for determining the three-dimensional multi-channel renderer comprises means for determining a regular three-dimensional multi-channel renderer when the determined local speaker geometry indicates a regular speaker geometry, as described above with respect to the example of FIGS. 8A .
- the device of the first example wherein the means for determining the renderer comprises means for determining an allowed order of spherical basis functions to which the spherical harmonic coefficients are associated, the allowed order identifying those of the spherical harmonic coefficients that are required to be rendered given the determined local speaker geometry, and means for determining the renderer based on the determined allowed order, as described above with respect to the examples of FIGS. 5-8B .
- the device of the first example wherein the means for determining the two-dimensional or three-dimensional renderer comprises means for determining an allowed order of spherical basis functions to which the spherical harmonic coefficients are associated, the allowed order identifying those of the spherical harmonic coefficients that are required to be rendered given the determined local speaker geometry; and means for determining the two-dimensional or three-dimensional renderer such that the two-dimensional or three-dimensional renderer only renders those of the spherical harmonic coefficients associated with spherical basis functions having an order less than or equal to the determined allowed order, as described above with respect to the examples of FIGS. 5-8B .
- the device of the first example wherein the means for determining the local speaker geometry of the one or more speakers comprises means for receiving input from a listener specifying local speaker geometry information describing the local speaker geometry.
- the device of the first example, wherein determining the two-dimensional or three-dimensional renderer based on the local speaker geometry comprises, when the local speaker geometry conforms to a mono speaker geometry, determining a mono renderer, e.g., the mono renderer determination unit 48D.
- FIGS. 13A-13D are diagram illustrating bitstreams 31A-31D formed in accordance with the techniques described in this disclosure.
- bitstream 31A may represent one example of bitstream 31 shown in the example of FIG. 3 .
- the bitstream 31A includes audio rendering information 39A that includes one or more bits defining a signal value 54. This signal value 54 may represent any combination of the below described types of information.
- the bitstream 31A also includes audio content 58, which may represent one example of the audio content.
- the bitstream 31B may be similar to the bitstream 31A where the signal value 54 comprises an index 54A, one or more bits defining a row size 54B of the signaled matrix, one or more bits defining a column size 54C of the signaled matrix, and matrix coefficients 54D.
- the index 54A may be defined using two to five bits, while each of row size 54B and column size 54C may be defined using two to sixteen bits.
- the extraction device 38 may extract the index 54A and determine whether the index signals that the matrix is included in the bitstream 31B (where certain index values, such as 0000 or 1111, may signal that the matrix is explicitly specified in bitstream 31B).
- the bitstream 31B includes an index 54A signaling that the matrix is explicitly specified in the bitstream 31B.
- the extraction device 38 may extract the row size 54B and the column size 54C.
- the extraction device 38 may be configured to compute the number of bits to parse that represent matrix coefficients as a function of the row size 54B, the column size 54C and a signaled (not shown in FIG. 13A ) or implicit bit size of each matrix coefficient.
- the extraction device 38 may extract the matrix coefficients 54D, which the audio playback device 24 may use to configure one of the audio renderers 34 as described above. While shown as signaling the audio rendering information 39B a single time in the bitstream 31B, the audio rendering information 39B may be signaled multiple times in bitstream 31B or at least partially or fully in a separate out-of-band channel (as optional data in some instances).
- the bitstream 31C may represent one example of bitstream 31 shown in the example of FIG. 3 above.
- the bitstream 31C includes the audio rendering information 39C that includes a signal value 54, which in this example specifies an algorithm index 54E.
- the bitstream 31C also includes audio content 58.
- the algorithm index 54E may be defined using two to five bits, as noted above, where this algorithm index 54E may identify a rendering algorithm to be used when rendering the audio content 58.
- the extraction device 38 may extract the algorithm index and determine whether the algorithm index 54E signals that the matrix are included in the bitstream 31C (where certain index values, such as 0000 or 1111, may signal that the matrix is explicitly specified in bitstream 31C).
- the bitstream 31C includes the algorithm index 54E signaling that the matrix is not explicitly specified in bitstream 31C.
- the extraction device 38 forwards the algorithm index 54E to audio playback device, which selects the corresponding one (if available) the rendering algorithms (which are denoted as renderers 34 in the example of FIGS. 3 and 4 ). While shown as signaling audio rendering information 39C a single time in the bitstream 31C, in the example of FIG. 13C , audio rendering information 39C may be signaled multiple times in the bitstream 31C or at least partially or fully in a separate out-of-band channel (as optional data in some instances).
- the bitstream 31C may represent one example of bitstream 31 shown in FIGS. 4 , 5 and 8 above.
- the bitstream 31D includes the audio rendering information 39D that includes a signal value 54, which in this example specifies a matrix index 54F.
- the bitstream 31D also includes audio content 58.
- the matrix index 54F may be defined using two to five bits, as noted above, where this matrix index 54F may identify a rendering algorithm to be used when rendering the audio content 58.
- the extraction device 38 may extract the matrix index 50F and determine whether the matrix index 54F signals that the matrix are included in the bitstream 31D (where certain index values, such as 0000 or 1111, may signal that the matrix is explicitly specified in bitstream 31C).
- the bitstream 31D includes the matrix index 54F signaling that the matrix is not explicitly specified in bitstream 31D.
- the extraction device 38 forwards the matrix index 54F to audio playback device, which selects the corresponding one (if available) the renderers 34. While shown as signaling audio rendering information 39D a single time in the bitstream 31D, in the example of FIG. 13D , audio rendering information 39D may be signaled multiple times in the bitstream 31D or at least partially or fully in a separate out-of-band channel (as optional data in some instances).
- FIG. 14A and 14B are another example of a 3D renderer determination unit 48C that may perform various aspects of the techniques described in this disclosure. That is, 3D renderer determination unit 48C may represent a unit configured to, when a virtual speaker is arranged in a sphere geometry lower than a horizontal plane bisecting the sphere geometry, project the virtual speaker to a location on the horizontal plane, and perform two dimensional panning on a hierarchical set of elements that describe a sound field when generating a first plurality of loudspeaker channel signals that reproduce the sound field such that the reproduced sound field includes at least one sound that appears to originate from the projected location of the virtual speaker.
- the 3D renderer determination unit 48C may receive the SHC 27' and invoke virtual speaker renderer 350, which may represent a unit configured to perform virtual loudspeaker t-design rendering.
- the virtual speaker renderer 350 may render the SCH 27' and generate loudspeaker channel signals for a given number of virtual speakers (e.g., 22 or 32).
- the 3D renderer determination unit 48C further includes a spherical weighting unit 352, an upper hemisphere 3D panning unit 354, an ear-level 2D panning unit 356 and a lower hemisphere 2D panning unit 358.
- the spherical weighting unit 352 may represent a unit configured to weight certain channels.
- the upper hemisphere 3D panning unit 354 represents a unit configured to perform 3D panning on the spherically weighted virtual loudspeaker channel signals to pan these signals among the various upper hemisphere physical or, in other words, real speakers.
- the ear-level hemisphere 2D panning unit 356 represents a unit configured to perform 2D panning on the spherically weighted virtual loudspeaker channel signals to pan these signals among the various ear-level physical or, in other words, real speakers.
- the lower hemisphere 2D panning unit 358 represents a unit configured to perform 2D panning on the spherically weighted virtual loudspeaker channel signals to pan these signals among the various lower hemisphere physical or, in other words, real speakers.
- the 3D rendering determination unit 48C' may be similar to that shown in FIG. 14B except the 3D rendering determination unit 48C' may not perform spherical weighting or otherwise include the spherical weighting unit 352.
- the loudspeaker feeds are computed by assuming that each loudspeaker produces a spherical wave.
- the transform matrix may vary depending on, for example, which SHC were used in the subset (e.g., the basic set) and which definition of SH basis function is used.
- a transform matrix to convert from a selected basic set to a different channel format (e.g., 7.1, 22.2) may be constructed
- the way outlined above to manipulate the above framework to ensure invertibility may result in less-than-desirable audio-image quality. That is, the sound reproduction may not always result in a correct localization of sounds when compared to the audio being captured.
- the techniques may be further augmented to introduce a concept that may be referred to as "virtual speakers.”
- a concept that may be referred to as "virtual speakers.”
- the above framework may be modified to include some form of panning, such as vector base amplitude panning (VBAP), distance based amplitude panning, or other forms of panning.
- VBAP vector base amplitude panning
- VBAP distance based amplitude panning
- VBAP may effectively introduce what may be characterized as "virtual speakers.”
- VBAP may generally modify a feed to one or more loudspeakers so that these one or more loudspeakers effectively output sound that appears to originate from a virtual speaker at one or more of a location and angle different than at least one of the location and/or angle of the one or more loudspeakers that supports the virtual speaker.
- the VBAP matrix is of size M rows by N columns, where M denotes the number of speakers (and would be equal to five in the equation above) and N denotes the number of virtual speakers.
- the VBAP matrix may be computed as a function of the vectors from the defined location of the listener to each of the positions of the speakers and the vectors from the defined location of the listener to each of the positions of the virtual speakers.
- the D matrix in the above equation may be of size N rows by (order+1) 2 columns, where the order may refer to the order of the SH functions.
- the D matrix may represent the following matrix : h 0 2 kr 1 Y 0 0 ⁇ ⁇ 1 ⁇ 1 h 0 2 kr 2 Y 0 0 ⁇ ⁇ 2 ⁇ 2 . . . h 1 2 kr 1 Y 1 1 ⁇ ⁇ 1 ⁇ 1 . . . . . . . . . . . . . . .
- the VBAP matrix is an MxN matrix providing what may be referred to as a "gain adjustment" that factors in the location of the speakers and the position of the virtual speakers.
- Introducing panning in this manner may result in better reproduction of the multi-channel audio that results in a better quality image when reproduced by the local speaker geometry.
- the techniques may overcome poor speaker geometries that do not align with those specified in various standards.
- the equation may be inverted and employed to transform SHC back to a multi-channel feed for a particular geometry or configuration of loudspeakers, which may be referred to as geometry B below. That is, the equation may be inverted to solve for the g matrix.
- the g matrix may represent speaker gain for, in this example, each of the five loudspeakers in a 5.1 speaker configuration.
- the virtual speakers locations used in this configuration may correspond to the locations defined in a 5.1 multichannel format specification or standard.
- the location of the loudspeakers that may support each of these virtual speakers may be determined using any number of known audio localization techniques, many of which involve playing a tone having a particular frequency to determine a location of each loudspeaker with respect to a headend unit (such as an audio/video receiver (A/V receiver), television, gaming system, digital video disc system, or other types of headend systems).
- a user of the headend unit may manually specify the location of each of the loudspeakers.
- the headend unit may solve for the gains, assuming an ideal configuration of virtual loudspeakers by way of VBAP.
- the techniques may enable a device or apparatus to perform a vector base amplitude panning or other form of panning on the first plurality of loudspeaker channel signals to produce a first plurality of virtual loudspeaker channel signals.
- These virtual loudspeaker channel signals may represent signals provided to the loudspeakers that enable these loudspeakers to produce sounds that appear to originate from the virtual loudspeakers.
- the techniques may enable a device or apparatus to perform the first transform on the first plurality of virtual loudspeaker channel signals to produce the hierarchical set of elements that describes the sound field.
- the techniques may enable an apparatus to perform a second transform on the hierarchical set of elements to produce a second plurality of loudspeaker channel signals, where each of the second plurality of loudspeaker channel signals is associated with a corresponding different region of space, where the second plurality of loudspeaker channel signals comprise a second plurality of virtual loudspeaker channels and where the second plurality of virtual loudspeaker channel signals is associated with the corresponding different region of space.
- the techniques may, in some instances, enable a device to perform a vector base amplitude panning on the second plurality of virtual loudspeaker channel signals to produce a second plurality of loudspeaker channel signals.
- transformation matrix was derived from a 'mode matching' criteria
- alternative transform matrices can be derived from other criteria as well, such as pressure matching, energy matching, etc. It is sufficient that a matrix can be derived that allows the transformation between the basic set (e.g., SHC subset) and traditional multichannel audio and also that after manipulation (that does not reduce the fidelity of the multichannel audio), a slightly modified matrix can also be formulated that is also invertible.
- the above described 3D panning may introduce artifacts or otherwise result in lower quality playback of the speaker feeds.
- the 3D panning described above may be employed with respect to a 22.2 speaker geometry, which is shown in FIG. 15A and FIG. 15B .
- FIG. 15A and 15B illustrate the same 22.2 speaker geometry, where the black dots in the graph shown in FIG. 15A shows the location of all loudspeakers 22 speakers (and excluding the low frequency speakers) and FIG. 15B shows the location of these same speakers but additionally defines the semi-sphere positional nature of these speakers (which blocks those speakers located behind the shaded semi-sphere).
- M the number of which are denoted as M above
- the listener's head is positioned somewhere in the semi-sphere around the (x, y, z) point of (0, 0, 0) in the graphs of FIGS. 15A and 15B .
- the 3D renderer determination unit 48C shown in the example of FIG 14A may represent a unit to, when a virtual speaker is arranged in a sphere geometry lower than a horizontal plane bisecting the sphere geometry, project the virtual speaker to a location on the horizontal plane, and perform two dimensional panning on a hierarchical set of elements that describe a sound field when generating a first plurality of loudspeaker channel signals that reproduce the sound field such that the reproduced sound field includes at least one sound that appears to originate from the projected location of the virtual speaker.
- FIG. 16A shows a sphere 400 bisected by a horizontal plane 402 on to which virtual speakers are projected upwards in accordance with the technique described in this disclosure.
- the techniques may project the virtual speakers to any horizontal plane (e.g. elevation) within the sphere 400.
- FIG. 16B shows the sphere 400 bisected by a horizontal plane 402 on to which virtual speakers are projected downward in accordance with the techniques described in this disclosure.
- the 3D renderer determination unit 48C may project the virtual speakers 300A-300C down to the horizontal plane 402. While described as being projected onto a horizontal plane 402 that equally bisects the sphere 400, the techniques may project the virtual speakers to any horizontal plane (e.g. elevation) within the sphere 400.
- the techniques may enable the 3D renderer determination unit 48C to determine a position of one of a plurality of physical speakers relative to a position of one of a plurality of virtual speakers arranged in a geometry, and adjust the position of the one of the plurality of virtual speakers within the geometry based on the determined position.
- the 3D renderer determination unit 48C may be further configured to perform a first transform in addition to the two dimensional panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals, wherein each of the first plurality of loudspeaker channel signals is associated with a corresponding different region of space.
- This first transform may be reflected in the equations above as D -1 .
- the 3D renderer determination unit 48C may be further configured to, when performing two dimensional panning on the hierarchical set of elements, perform two dimensional vector base amplitude panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals.
- each of the first plurality of loudspeaker channel signals is associated with a corresponding different defined region of space.
- the different defined regions of space are defined in one or more of an audio format specification and an audio format standard.
- the 3D renderer determination unit 48C may also or alternatively be configured to, when a virtual speaker is arranged in a sphere geometry near a horizontal plane at or near ear level in the sphere geometry, perform two dimensional panning on a hierarchical set of elements that describe a sound field when generating a first plurality of loudspeaker channel signals that reproduce the sound field such that the reproduced sound field includes at least one sound that appears to originate from a location of the virtual speaker.
- the 3D renderer determination unit 48C may be further configured to perform a first transform (which again may refer to the D -1 transform noted above) in addition to the two dimensional panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals, where each of the first plurality of loudspeaker channel signals is associated with a corresponding different region of space.
- a first transform which again may refer to the D -1 transform noted above
- the 3D renderer determination unit 48C may be further configured to, when performing two dimensional panning on the hierarchical set of elements, perform two dimensional vector base amplitude panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals.
- each of the first plurality of loudspeaker channel signals is associated with a corresponding different defined region of space.
- the different defined regions of space may be defined in one or more of an audio format specification and an audio format standard.
- the one or more processors of device 10 may be further configured to, when a virtual speaker is arranged in a sphere geometry above a horizontal plane bisecting the sphere geometry, perform three dimensional panning on a hierarchical set of elements when generating a first plurality of loudspeaker channel signals that describe a sound field such that the sound field includes at least one sound that appears to originate from a location of the virtual speaker.
- the 3D renderer determination unit 48C may be further configured to perform a first transform in addition to the three dimensional panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals, wherein each of the first plurality of loudspeaker channel signals is associated with a corresponding different region of space.
- the 3D renderer determination unit 48C may be further configured to, when performing three dimensional panning on the hierarchical set of elements the first plurality of loudspeaker channel signals, perform three dimensional vector base amplitude panning on the hierarchical set of elements when generating the first plurality of loudspeaker channel signals.
- each of the first plurality of loudspeaker channel signals is associated with a corresponding different defined region of space.
- the different defined regions of space may be defined in one or more of an audio format specification and an audio format standard.
- the 3D renderer determination unit 48C may further be configure to, when performing both a three dimensional panning and a two dimensional panning in the generation of a plurality of loudspeaker channel signals from a hierarchical set of elements, perform a weighting with respect to the hierarchical set of elements based on an order of each of the hierarchical set of elements.
- the 3D renderer determination unit 48C may be further configure to, when performing the weighting, perform a window function with respect to the hierarchical set of elements based on the order of each of the hierarchical set of elements.
- This windowing function may be shown in the example of FIG. 17 , where the y-axis reflects decibels and the x-axis denotes the order of the SHC.
- the one or more processors of device 10 may further be configure to, when performing the weighting, perform a Kaiser Bessle window function, as one example, with respect to the hierarchical set of elements based on the order of each of the hierarchical set of elements.
- processors may each represent means for perform the various functions attributed to the one or more processors.
- Other means may include dedicated application specific hardware, field programmable gate arrays, application specific integrated circuits or any other form of hardware dedicated or capable of executing software that may perform the various aspects either alone or in combination of the techniques described in this disclosure.
- the problem identified and potentially solved by the techniques may be summarized as follows.
- the arrangement of the loudspeakers may be crucial.
- a three-dimensional sphere of equidistant loudspeakers may be desired.
- current loudspeaker setups are typically 1) not equally distributed, 2) exist only in the upper hemisphere around and above a listener, not in the lower hemisphere below and 3) for legacy support (e.g., 5.1 speaker setup) have usually a ring of loudspeakers at the height of the ears.
- the techniques may provide for a different treatment of the virtual loudspeaker signals:
- the first aspects of the techniques may enable device 10 to orthogonally map the virtual loudspeakers coming from the lower hemisphere onto the horizontal plane and projected onto the two closest real loudspeakers using a two-dimensional panning method.
- the first aspect of the techniques may minimize, reduce or remove localization errors caused by wrongly projected virtual loudspeakers.
- the virtual loudspeakers in the upper hemisphere that are at (or about) the height of the ears may also be projected to the two closest loudspeakers using a two-dimensional panning method in accordance with the second aspects of the techniques described in this disclosure.
- the reason behind this second modification may be that humans may not be as accurate in the perception of elevated sound sources, compared to the perception of the azimuthal direction.
- VBAP is generally known to be accurate in the creation of azimuthal direction of a virtual sound source, it is relatively inaccurate in the creation of elevated sounds - often the perceived virtual sounds sources are perceived with a higher elevation than intended.
- the second aspect of the techniques avoids using 3D-VBAP in spatial area which would not benefit from it, and may even cause a degraded quality.
- the third aspect of the techniques is that all the remaining virtual loudspeakers of the upper hemisphere above ear level are projected using a conventional three-dimensional panning method.
- a fourth aspect of the techniques may be performed where all the Higher Order Ambisonics / Spherical Harmonic Coefficients surround-sound material are weighted using a weighting function as a function of the spherical harmonics order to increase a smoother spatial reproduction of the material. This has been shown to be potentially beneficial for matching the energy of the 2D and the 3D panned virtual loudspeakers.
- the 3D renderer determination unit 48C may perform any combination of the aspects described in this disclosure, performing one or more of the four aspects.
- a different device that generates spherical harmonic coefficients may perform various aspects of the techniques in a reciprocal manner. While not described in detail to avoid redundancy, the techniques of this disclosure should not be strictly limited to the example of FIG. 14A .
- the above thus represents a lossless mechanism to convert between a hierarchical set of elements (e.g., a set of SHC) and multiple audio channels.
- a hierarchical set of elements e.g., a set of SHC
- No errors are incurred as long as the multichannel audio signals are not subjected to further coding noise.
- the conversion to SHC may incur errors.
- it is possible to account for these errors by monitoring the values of the coefficients and taking appropriate action to reduce their effect.
- These methods may take into account characteristics of the SHC, including the inherent redundancy in the SHC representation.
- the approach described herein provides a solution to a potential disadvantage in the use of SHC-based representation of sound fields. Without this solution, the SHC-based representation may not be deployed, due to the significant disadvantage imposed by not being able to have functionality in the millions of legacy playback systems.
- the techniques may therefore provide for, in a first example, a device comprising means for determining a difference in position between one of a plurality of physical speakers and one of a plurality of virtual speakers arranged in a geometry, e.g., the renderer determination unit 40, and means for adjusting a position of the one of the plurality of virtual speakers within the geometry based on the determined difference in position, e.g., the renderer determination unit 40.
- the device of the first example wherein the means for determining the difference in position comprises means for determining a difference in elevation between the one of the plurality of physical speakers and the one of the plurality of virtual speakers, e.g., the 3D renderer determination unit 48C.
- the device of the first example wherein the means for determining the difference in position comprises means for determining a difference in elevation between the one of the plurality of physical speakers and the one of the plurality of virtual speakers, and wherein the means for adjusting the position of the one of the plurality of virtual speakers comprises means for projecting the one of the plurality of virtual speakers to an elevation lower than an original elevation of the plurality of virtual speakers when the determined difference in elevation exceeds a threshold value, as described above in more detail with respect to the examples of FIGS. 8A-9 and 14A-16B .
- the device of the first examples wherein the means for determining the difference in position comprises means for determining a difference in elevation between the one of the plurality of physical speakers and the one of the plurality of virtual speakers, and wherein the means for adjusting the position of the one of the plurality of virtual speakers comprises means for projecting the one of the plurality of virtual speakers to an elevation higher than an original elevation of the one of the plurality of virtual speakers when the determined difference in elevation exceeds a threshold value, as described above in more detail with respect to the examples of FIGS. 8A-9 and 14A-16B .
- the device of the first example further comprising means for performing two dimensional panning on a hierarchical set of elements that describe a sound field when generating a plurality of loudspeaker channel signals to drive the plurality of physical speakers so as to reproduce the sound field such that the reproduced sound field includes at least one sound that appears to originate from the adjusted location of the virtual speaker, as described above in more detail with respect to the examples of FIGS. 8A and 8B .
- the device of the fifth example wherein the hierarchical set of elements comprise a plurality of spherical harmonic coefficients.
- the device of the fifth example wherein the means for performing two dimensional panning on the hierarchical set of elements comprises means for performing two dimensional vector based amplitude panning on the hierarchical set of elements when generating the plurality of loudspeaker channel signals, as described above in more detail with respect to the examples of FIGS. 8A and 8B .
- the device of the first example further comprising means for determining one or more stretched physical speaker positions that are different from positions of the corresponding one or more of the plurality of physical speakers, as described above in more detail with respect to the examples of FIGS. 8A-12B .
- the device of the first example further comprising means for determining one or more stretched physical speaker positions that are different from positions of the corresponding one or more of the plurality of physical speakers, wherein the means for determining the difference in position comprises means for determining a difference between at least one of the stretched physical speaker positions relative to the position of the one of the plurality of virtual speakers, as described above in more detail with respect to the examples of FIGS. 8A-12B .
- the device of the first example further comprising means for determining one or more stretched physical speaker positions that are different from positions of the corresponding one or more of the plurality of physical speakers, wherein the means for determining the difference in position comprises means for determining a difference in elevation between at least one of the stretched physical speaker positions and the position of the one of the plurality of virtual speakers, and wherein the means for adjusting the position of the one of the plurality of virtual speakers comprises means for projecting the one of the plurality of virtual speakers to an elevation lower than an original elevation of the plurality of virtual speakers when the determined difference in elevation exceeds a threshold value, as described above in more detail with respect to the examples of FIGS. 8A-12B and 14A-16B .
- the device of the first example further comprising means for determining one or more stretched physical speaker positions that are different from positions of the corresponding one or more of the plurality of physical speakers, wherein the means for determining the difference in position comprises means for determining a difference in elevation between at least one of the stretched physical speaker positions and the position of the one of the plurality of virtual speakers, and wherein the means for adjusting the position of the one of the plurality of virtual speakers comprises means for projecting the one of the plurality of virtual speakers to an elevation higher than an original elevation of the plurality of virtual speakers when the determined difference in elevation exceeds a threshold value, , as described above in more detail with respect to the examples of FIGS. 8A-12B and 14A-16B .
- the device of the first example wherein the plurality of virtual speakers are arranged in a spherical geometry, as described above in more detail with respect to the examples of FIGS. 8A-12B and 14A-16B .
- the device of the first example wherein the plurality of virtual speakers are arranged in a polyhedron geometry.
- the techniques may be performed with respect to any virtual speaker geometry, including any form of polyhedron geometry, such as a cubic geometry, a dodecahedron geometry, an icosidodecahedron geometry, a rhombic triacontahedron geometry, a prism geometry, and a pyramid geometry to provide a few examples.
- the device of the first example wherein the plurality of physical speakers are arranged in an irregular speaker geometry.
- the device of the first example wherein the plurality of physical speakers are arranged in an irregular speaker geometry on multiple different horizontal planes.
- the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit.
- Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol.
- computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave.
- Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and/or data structures for implementation of the techniques described in this disclosure.
- a computer program product may include a computer-readable medium.
- such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer.
- any connection is properly termed a computer-readable medium.
- a computer-readable medium For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium.
- DSL digital subscriber line
- Disk and disc includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
- processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry.
- DSPs digital signal processors
- ASICs application specific integrated circuits
- FPGAs field programmable logic arrays
- processors may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein.
- the functionality described herein may be provided within dedicated hardware and/or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
- the techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set).
- IC integrated circuit
- a set of ICs e.g., a chip set.
- Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and/or firmware
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Stereophonic System (AREA)
- Circuit For Audible Band Transducer (AREA)
Claims (15)
- Verfahren zum Erzeugen eines Audio-Renderers zur Wiedergabe von Kugelflächenfunktionskoeffizienten, wobei das Verfahren Folgendes beinhaltet:Bestimmen, durch einen oder mehrere Prozessoren eines Geräts, von lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) durch Kategorisieren einer lokalen Lautsprechergeometrie als eine aus einer Mehrzahl von Lautsprechergeometrien, wobei die Mehrzahl von Lautsprechergeometrien wenigstens eine oder mehrere aus einer Mono-Lautsprechergeometrie, einer Stereo-Lautsprechergeometrie, einer horizontalen Mehrkanal-Lautsprechergeometrie oder einer dreidimensionalen Mehrkanal-Lautsprechergeometrie beinhalten, auf der Basis von lokalen Lautsprechergeometrieinformationen (41) von einem oder mehreren Lautsprechern, die zur Wiedergabe von Kugelflächenfunktionskoeffizienten benutzt werden, die für ein Schallfeld repräsentativ sind; undErzeugen eines zweidimensionalen oder eines dreidimensionalen Renderers (34) auf der Basis der Kategorisierungsinformationen (45) und der lokalen Lautsprechergeometrieinformationen (41).
- Verfahren nach Anspruch 1, das ferner das Rendern der Kugelflächenfunktionskoeffizienten anhand des erzeugten Renderers zum Erzeugen von Mehrkanal-Audiodaten beinhaltet.
- Verfahren nach Anspruch 1, das ferner, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) einer Stereo-Lautsprechergeometrie entsprechen, das Erzeugen eines zweidimensionalen Stereo-Renderers beinhaltet.
- Verfahren nach Anspruch 1, das ferner, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) einer horizontalen Mehrkanal-Lautsprechergeometrie mit mehr als zwei Lautsprechern entsprechen, das Erzeugen eines horizontalen zweidimensionalen Mehrkanal-Renderers beinhaltet.
- Verfahren nach Anspruch 4, wobei das Erzeugen des horizontalen zweidimensionalen Mehrkanal-Renderers das Erzeugen eines unregelmäßigen horizontalen zweidimensionalen Mehrkanal-Renderers beinhaltet, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) eine unregelmäßige Lautsprechergeometrie anzeigen.
- Verfahren nach Anspruch 4, wobei das Erzeugen des horizontalen zweidimensionalen Mehrkanal-Renderers das Erzeugen eines regelmäßigen horizontalen zweidimensionalen Mehrkanal-Renderers beinhaltet, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) eine regelmäßige Lautsprechergeometrie anzeigen.
- Verfahren nach Anspruch 1, das ferner, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) einer dreidimensionalen Mehrkanal-Lautsprechergeometrie mit mehr als zwei Lautsprechern auf mehr als einer horizontalen Ebene entsprechen, das Erzeugen eines dreidimensionalen Mehrkanal-Renderers beinhaltet.
- Verfahren nach Anspruch 7, wobei das Erzeugen des dreidimensionalen Mehrkanal-Renderers das Erzeugen eines unregelmäßigen dreidimensionalen Mehrkanal-Renderers beinhaltet, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) eine unregelmäßige Lautsprechergeometrie anzeigen.
- Verfahren nach Anspruch 7, wobei das Erzeugen des dreidimensionalen Mehrkanal-Renderers das Erzeugen eines nahezu regelmäßigen dreidimensionalen Mehrkanal-Renderers beinhaltet, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) eine nahezu regelmäßige Lautsprechergeometrie anzeigen.
- Verfahren nach Anspruch 7, wobei das Erzeugen des dreidimensionalen Mehrkanal-Renderers das Erzeugen eines regelmäßigen dreidimensionalen Mehrkanal-Renderers beinhaltet, wenn die bestimmten lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) eine regelmäßige Lautsprechergeometrie anzeigen.
- Verfahren nach Anspruch 1, wobei das Erzeugen des Renderers Folgendes beinhaltet:Bestimmen einer zulässigen Reihenfolge von sphärischen Basisfunktionen, mit denen die Kugelflächenfunktionskoeffizienten assoziiert sind, wobei die zulässige Reihenfolge diejenigen der Kugelflächenfunktionskoeffizienten identifiziert, die angesichts der lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) zum Rendern benötigt werden; undErzeugen des Renderers auf der Basis der bestimmten zulässigen Reihenfolge.
- Verfahren nach Anspruch 11, wobei das Erzeugen des Renderers Folgendes beinhaltet:
Erzeugen des Renderers so, dass der Renderer nur diejenigen der Kugelflächenfunktionskoeffizienten rendert, die mit sphärischen Basisfunktionen mit einer Ordnung assoziiert sind, die gleich oder kleiner ist als die vorbestimmte zulässige Ordnung. - Verfahren nach Anspruch 1, wobei das Bestimmen der lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) das Empfangen von Eingängen von einem Zuhörer beinhaltet, wobei der Eingang die lokalen Lautsprechergeometrieinformationen (41) beschreibt.
- Gerät zum Erzeugen eines Audio-Renderers zur Wiedergabe von Kugelflächenfunktionskoeffizienten, wobei das Gerät Folgendes umfasst:
einen oder mehrere Prozessoren, konfiguriert zum:Bestimmen von lokalen Lautsprechergeometrie-Kategorisierungsinformationen (45) durch Kategorisieren einer lokalen Lautsprechergeometrie als eine aus einer Mehrzahl von Lautsprechergeometrien, wobei die Mehrzahl von Lautsprechergeometrien wenigstens eine oder mehrere aus einer Mono-Lautsprechergeometrie, einer Stereo-Lautsprechergeometrie, einer horizontalen Mehrkanal-Lautsprechergeometrie oder einer dreidimensionalen Mehrkanal-Lautsprechergeometrie beinhaltet, auf der Basis von lokalen Lautsprechergeometrieinformationen (41) von einem oder mehreren Lautsprechern, die zur Wiedergabe von Kugelflächenfunktionskoeffizienten benutzt werden, die für ein Schallfeld repräsentativ sind; undErzeugen eines zweidimensionalen oder eines dreidimensionalen Renderers (34) auf der Basis der Kategorisierungsinformationen (45) und der lokalen Lautsprechergeometrieinformationen (41). - Nichtflüchtiges computerlesbares Speichermedium, auf dem Befehle gespeichert sind, die bei Ausführung bewirken, dass ein oder mehrere Prozessoren das Verfahren nach einem der Ansprüche 1-13 durchführen.
Applications Claiming Priority (4)
Application Number | Priority Date | Filing Date | Title |
---|---|---|---|
US201361762302P | 2013-02-07 | 2013-02-07 | |
US201361829832P | 2013-05-31 | 2013-05-31 | |
US14/174,784 US9736609B2 (en) | 2013-02-07 | 2014-02-06 | Determining renderers for spherical harmonic coefficients |
PCT/US2014/015311 WO2014124264A1 (en) | 2013-02-07 | 2014-02-07 | Determining renderers for spherical harmonic coefficients |
Publications (2)
Publication Number | Publication Date |
---|---|
EP2954703A1 EP2954703A1 (de) | 2015-12-16 |
EP2954703B1 true EP2954703B1 (de) | 2019-12-18 |
Family
ID=51259222
Family Applications (2)
Application Number | Title | Priority Date | Filing Date |
---|---|---|---|
EP14707033.8A Active EP2954702B1 (de) | 2013-02-07 | 2014-02-07 | Abbildung virtueller lautsprecher auf physikalischen lautsprechern |
EP14707870.3A Active EP2954703B1 (de) | 2013-02-07 | 2014-02-07 | Bestimmung von renderern für kugelflächenfunktionskoeffizienten |
Family Applications Before (1)
Application Number | Title | Priority Date | Filing Date |
---|---|---|---|
EP14707033.8A Active EP2954702B1 (de) | 2013-02-07 | 2014-02-07 | Abbildung virtueller lautsprecher auf physikalischen lautsprechern |
Country Status (7)
Country | Link |
---|---|
US (2) | US9913064B2 (de) |
EP (2) | EP2954702B1 (de) |
JP (2) | JP6309545B2 (de) |
KR (2) | KR101877604B1 (de) |
CN (2) | CN104969577B (de) |
TW (2) | TWI611706B (de) |
WO (2) | WO2014124268A1 (de) |
Families Citing this family (120)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US8483853B1 (en) | 2006-09-12 | 2013-07-09 | Sonos, Inc. | Controlling and manipulating groupings in a multi-zone media system |
US9202509B2 (en) | 2006-09-12 | 2015-12-01 | Sonos, Inc. | Controlling and grouping in a multi-zone media system |
US8788080B1 (en) | 2006-09-12 | 2014-07-22 | Sonos, Inc. | Multi-channel pairing in a media system |
US8923997B2 (en) | 2010-10-13 | 2014-12-30 | Sonos, Inc | Method and apparatus for adjusting a speaker system |
US11265652B2 (en) | 2011-01-25 | 2022-03-01 | Sonos, Inc. | Playback device pairing |
US11429343B2 (en) | 2011-01-25 | 2022-08-30 | Sonos, Inc. | Stereo playback configuration and control |
US8938312B2 (en) | 2011-04-18 | 2015-01-20 | Sonos, Inc. | Smart line-in processing |
US9042556B2 (en) | 2011-07-19 | 2015-05-26 | Sonos, Inc | Shaping sound responsive to speaker orientation |
US8811630B2 (en) | 2011-12-21 | 2014-08-19 | Sonos, Inc. | Systems, methods, and apparatus to filter audio |
US9084058B2 (en) | 2011-12-29 | 2015-07-14 | Sonos, Inc. | Sound field calibration using listener localization |
US9729115B2 (en) | 2012-04-27 | 2017-08-08 | Sonos, Inc. | Intelligently increasing the sound level of player |
US9524098B2 (en) | 2012-05-08 | 2016-12-20 | Sonos, Inc. | Methods and systems for subwoofer calibration |
USD721352S1 (en) | 2012-06-19 | 2015-01-20 | Sonos, Inc. | Playback device |
US9668049B2 (en) | 2012-06-28 | 2017-05-30 | Sonos, Inc. | Playback device calibration user interfaces |
US9219460B2 (en) | 2014-03-17 | 2015-12-22 | Sonos, Inc. | Audio settings based on environment |
US9106192B2 (en) | 2012-06-28 | 2015-08-11 | Sonos, Inc. | System and method for device playback calibration |
US9690539B2 (en) | 2012-06-28 | 2017-06-27 | Sonos, Inc. | Speaker calibration user interface |
US9706323B2 (en) | 2014-09-09 | 2017-07-11 | Sonos, Inc. | Playback device calibration |
US9690271B2 (en) | 2012-06-28 | 2017-06-27 | Sonos, Inc. | Speaker calibration |
US9288603B2 (en) | 2012-07-15 | 2016-03-15 | Qualcomm Incorporated | Systems, methods, apparatus, and computer-readable media for backward-compatible audio coding |
US9473870B2 (en) * | 2012-07-16 | 2016-10-18 | Qualcomm Incorporated | Loudspeaker position compensation with 3D-audio hierarchical coding |
US8930005B2 (en) | 2012-08-07 | 2015-01-06 | Sonos, Inc. | Acoustic signatures in a playback system |
US8965033B2 (en) | 2012-08-31 | 2015-02-24 | Sonos, Inc. | Acoustic optimization |
US9008330B2 (en) | 2012-09-28 | 2015-04-14 | Sonos, Inc. | Crossover frequency adjustments for audio speakers |
US9913064B2 (en) | 2013-02-07 | 2018-03-06 | Qualcomm Incorporated | Mapping virtual speakers to physical speakers |
EP2765791A1 (de) * | 2013-02-08 | 2014-08-13 | Thomson Licensing | Verfahren und Vorrichtung zur Bestimmung der Richtungen dominanter Schallquellen bei einer Higher-Order-Ambisonics-Wiedergabe eines Schallfelds |
USD721061S1 (en) | 2013-02-25 | 2015-01-13 | Sonos, Inc. | Playback device |
EP2991384B1 (de) | 2013-04-26 | 2021-06-02 | Sony Corporation | Audioverarbeitungsvorrichtung, -verfahren und -programm |
WO2014175076A1 (ja) | 2013-04-26 | 2014-10-30 | ソニー株式会社 | 音声処理装置および音声処理システム |
US9412385B2 (en) * | 2013-05-28 | 2016-08-09 | Qualcomm Incorporated | Performing spatial masking with respect to spherical harmonic coefficients |
US9466305B2 (en) | 2013-05-29 | 2016-10-11 | Qualcomm Incorporated | Performing positional analysis to code spherical harmonic coefficients |
US10499176B2 (en) | 2013-05-29 | 2019-12-03 | Qualcomm Incorporated | Identifying codebooks to use when coding spatial components of a sound field |
CN105379311B (zh) * | 2013-07-24 | 2018-01-16 | 索尼公司 | 信息处理设备以及信息处理方法 |
WO2015054033A2 (en) * | 2013-10-07 | 2015-04-16 | Dolby Laboratories Licensing Corporation | Spatial audio processing system and method |
KR102231755B1 (ko) | 2013-10-25 | 2021-03-24 | 삼성전자주식회사 | 입체 음향 재생 방법 및 장치 |
US9489955B2 (en) | 2014-01-30 | 2016-11-08 | Qualcomm Incorporated | Indicating frame parameter reusability for coding vectors |
US9922656B2 (en) | 2014-01-30 | 2018-03-20 | Qualcomm Incorporated | Transitioning of ambient higher-order ambisonic coefficients |
US9226073B2 (en) | 2014-02-06 | 2015-12-29 | Sonos, Inc. | Audio output balancing during synchronized playback |
US9226087B2 (en) | 2014-02-06 | 2015-12-29 | Sonos, Inc. | Audio output balancing during synchronized playback |
US9264839B2 (en) | 2014-03-17 | 2016-02-16 | Sonos, Inc. | Playback device configuration based on proximity detection |
US10412522B2 (en) | 2014-03-21 | 2019-09-10 | Qualcomm Incorporated | Inserting audio channels into descriptions of soundfields |
WO2015147435A1 (ko) * | 2014-03-25 | 2015-10-01 | 인텔렉추얼디스커버리 주식회사 | 오디오 신호 처리 시스템 및 방법 |
US10770087B2 (en) | 2014-05-16 | 2020-09-08 | Qualcomm Incorporated | Selecting codebooks for coding vectors decomposed from higher-order ambisonic audio signals |
US9852737B2 (en) | 2014-05-16 | 2017-12-26 | Qualcomm Incorporated | Coding vectors decomposed from higher-order ambisonics audio signals |
US9620137B2 (en) | 2014-05-16 | 2017-04-11 | Qualcomm Incorporated | Determining between scalar and vector quantization in higher order ambisonic coefficients |
US9367283B2 (en) | 2014-07-22 | 2016-06-14 | Sonos, Inc. | Audio settings |
USD883956S1 (en) | 2014-08-13 | 2020-05-12 | Sonos, Inc. | Playback device |
US10127006B2 (en) | 2014-09-09 | 2018-11-13 | Sonos, Inc. | Facilitating calibration of an audio playback device |
US9952825B2 (en) | 2014-09-09 | 2018-04-24 | Sonos, Inc. | Audio processing algorithms |
US9910634B2 (en) | 2014-09-09 | 2018-03-06 | Sonos, Inc. | Microphone calibration |
US9891881B2 (en) | 2014-09-09 | 2018-02-13 | Sonos, Inc. | Audio processing algorithm database |
US9747910B2 (en) | 2014-09-26 | 2017-08-29 | Qualcomm Incorporated | Switching between predictive and non-predictive quantization techniques in a higher order ambisonics (HOA) framework |
EP3540732B1 (de) | 2014-10-31 | 2023-07-26 | Dolby International AB | Parametrische decodierung von mehrkanaligen audiosignalen |
WO2016077317A1 (en) * | 2014-11-11 | 2016-05-19 | Google Inc. | Virtual sound systems and methods |
EP3024253A1 (de) * | 2014-11-21 | 2016-05-25 | Harman Becker Automotive Systems GmbH | Audiosystem und Verfahren |
US9973851B2 (en) | 2014-12-01 | 2018-05-15 | Sonos, Inc. | Multi-channel playback of audio content |
US10664224B2 (en) | 2015-04-24 | 2020-05-26 | Sonos, Inc. | Speaker calibration user interface |
WO2016172593A1 (en) | 2015-04-24 | 2016-10-27 | Sonos, Inc. | Playback device calibration user interfaces |
US20170085972A1 (en) | 2015-09-17 | 2017-03-23 | Sonos, Inc. | Media Player and Media Player Design |
USD768602S1 (en) | 2015-04-25 | 2016-10-11 | Sonos, Inc. | Playback device |
USD906278S1 (en) | 2015-04-25 | 2020-12-29 | Sonos, Inc. | Media player device |
USD920278S1 (en) | 2017-03-13 | 2021-05-25 | Sonos, Inc. | Media playback device with lights |
USD886765S1 (en) | 2017-03-13 | 2020-06-09 | Sonos, Inc. | Media playback device |
US10248376B2 (en) | 2015-06-11 | 2019-04-02 | Sonos, Inc. | Multiple groupings in a playback system |
US10334387B2 (en) | 2015-06-25 | 2019-06-25 | Dolby Laboratories Licensing Corporation | Audio panning transformation system and method |
US9729118B2 (en) | 2015-07-24 | 2017-08-08 | Sonos, Inc. | Loudness matching |
US9538305B2 (en) | 2015-07-28 | 2017-01-03 | Sonos, Inc. | Calibration error conditions |
US12087311B2 (en) | 2015-07-30 | 2024-09-10 | Dolby Laboratories Licensing Corporation | Method and apparatus for encoding and decoding an HOA representation |
EP3329486B1 (de) * | 2015-07-30 | 2020-07-29 | Dolby International AB | Verfahren und vorrichtung zur erzeugung einer mezzanin-hoa-signalrepräsentation aus einer hoa-signalrepräsentation |
US9736610B2 (en) | 2015-08-21 | 2017-08-15 | Sonos, Inc. | Manipulation of playback device response using signal processing |
US9712912B2 (en) | 2015-08-21 | 2017-07-18 | Sonos, Inc. | Manipulation of playback device response using an acoustic filter |
JP6437695B2 (ja) | 2015-09-17 | 2018-12-12 | ソノズ インコーポレイテッド | オーディオ再生デバイスのキャリブレーションを容易にする方法 |
US9693165B2 (en) | 2015-09-17 | 2017-06-27 | Sonos, Inc. | Validation of audio calibration using multi-dimensional motion check |
USD1043613S1 (en) | 2015-09-17 | 2024-09-24 | Sonos, Inc. | Media player |
US10306392B2 (en) | 2015-11-03 | 2019-05-28 | Dolby Laboratories Licensing Corporation | Content-adaptive surround sound virtualization |
CN105392102B (zh) * | 2015-11-30 | 2017-07-25 | 武汉大学 | 用于非球面扬声器阵列的三维音频信号生成方法及系统 |
EP3188504B1 (de) | 2016-01-04 | 2020-07-29 | Harman Becker Automotive Systems GmbH | Multimedia-wiedergabe für eine vielzahl von empfängern |
WO2017118551A1 (en) * | 2016-01-04 | 2017-07-13 | Harman Becker Automotive Systems Gmbh | Sound wave field generation |
US9743207B1 (en) | 2016-01-18 | 2017-08-22 | Sonos, Inc. | Calibration using multiple recording devices |
US10003899B2 (en) | 2016-01-25 | 2018-06-19 | Sonos, Inc. | Calibration with particular locations |
US11106423B2 (en) | 2016-01-25 | 2021-08-31 | Sonos, Inc. | Evaluating calibration of a playback device |
US9886234B2 (en) | 2016-01-28 | 2018-02-06 | Sonos, Inc. | Systems and methods of distributing audio to one or more playback devices |
DE102016103209A1 (de) | 2016-02-24 | 2017-08-24 | Visteon Global Technologies, Inc. | System und Verfahren zur Positionserkennung von Lautsprechern und zur Wiedergabe von Audiosignalen als Raumklang |
US9860662B2 (en) | 2016-04-01 | 2018-01-02 | Sonos, Inc. | Updating playback device configuration information based on calibration data |
US9864574B2 (en) | 2016-04-01 | 2018-01-09 | Sonos, Inc. | Playback device calibration based on representation spectral characteristics |
US9763018B1 (en) | 2016-04-12 | 2017-09-12 | Sonos, Inc. | Calibration of audio playback devices |
FR3050601B1 (fr) * | 2016-04-26 | 2018-06-22 | Arkamys | Procede et systeme de diffusion d'un signal audio a 360° |
US9794710B1 (en) | 2016-07-15 | 2017-10-17 | Sonos, Inc. | Spatial audio correction |
US9860670B1 (en) | 2016-07-15 | 2018-01-02 | Sonos, Inc. | Spectral correction using spatial calibration |
US10372406B2 (en) | 2016-07-22 | 2019-08-06 | Sonos, Inc. | Calibration interface |
US10459684B2 (en) | 2016-08-05 | 2019-10-29 | Sonos, Inc. | Calibration of a playback device based on an estimated frequency response |
US10412473B2 (en) | 2016-09-30 | 2019-09-10 | Sonos, Inc. | Speaker grill with graduated hole sizing over a transition area for a media device |
USD851057S1 (en) | 2016-09-30 | 2019-06-11 | Sonos, Inc. | Speaker grill with graduated hole sizing over a transition area for a media device |
USD827671S1 (en) | 2016-09-30 | 2018-09-04 | Sonos, Inc. | Media playback device |
US10712997B2 (en) | 2016-10-17 | 2020-07-14 | Sonos, Inc. | Room association based on name |
WO2018073759A1 (en) * | 2016-10-19 | 2018-04-26 | Audible Reality Inc. | System for and method of generating an audio image |
CA3054237A1 (en) * | 2017-01-27 | 2018-08-02 | Auro Technologies Nv | Processing method and system for panning audio objects |
JP6543848B2 (ja) * | 2017-03-29 | 2019-07-17 | 本田技研工業株式会社 | 音声処理装置、音声処理方法及びプログラム |
US11277705B2 (en) | 2017-05-15 | 2022-03-15 | Dolby Laboratories Licensing Corporation | Methods, systems and apparatus for conversion of spatial audio format(s) to speaker signals |
US10015618B1 (en) * | 2017-08-01 | 2018-07-03 | Google Llc | Incoherent idempotent ambisonics rendering |
US10609485B2 (en) | 2017-09-29 | 2020-03-31 | Apple Inc. | System and method for performing panning for an arbitrary loudspeaker setup |
WO2019149337A1 (en) * | 2018-01-30 | 2019-08-08 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Apparatuses for converting an object position of an audio object, audio stream provider, audio content production system, audio playback apparatus, methods and computer programs |
CN112119646B (zh) * | 2018-05-22 | 2022-09-06 | 索尼公司 | 信息处理装置、信息处理方法以及计算机可读存储介质 |
WO2020030304A1 (en) * | 2018-08-09 | 2020-02-13 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | An audio processor and a method considering acoustic obstacles and providing loudspeaker signals |
US10299061B1 (en) | 2018-08-28 | 2019-05-21 | Sonos, Inc. | Playback device calibration |
US11206484B2 (en) | 2018-08-28 | 2021-12-21 | Sonos, Inc. | Passive speaker authentication |
WO2020044244A1 (en) | 2018-08-29 | 2020-03-05 | Audible Reality Inc. | System for and method of controlling a three-dimensional audio engine |
US11798569B2 (en) * | 2018-10-02 | 2023-10-24 | Qualcomm Incorporated | Flexible rendering of audio data |
US10739726B2 (en) | 2018-10-03 | 2020-08-11 | International Business Machines Corporation | Audio management for holographic objects |
KR102323529B1 (ko) | 2018-12-17 | 2021-11-09 | 한국전자통신연구원 | 복합 차수 앰비소닉을 이용한 오디오 신호 처리 방법 및 장치 |
US11968518B2 (en) | 2019-03-29 | 2024-04-23 | Sony Group Corporation | Apparatus and method for generating spatial audio |
CN113853803A (zh) | 2019-04-02 | 2021-12-28 | 辛格股份有限公司 | 用于空间音频渲染的系统和方法 |
US11122386B2 (en) * | 2019-06-20 | 2021-09-14 | Qualcomm Incorporated | Audio rendering for low frequency effects |
US10734965B1 (en) | 2019-08-12 | 2020-08-04 | Sonos, Inc. | Audio calibration of a portable playback device |
CN110751956B (zh) * | 2019-09-17 | 2022-04-26 | 北京时代拓灵科技有限公司 | 一种沉浸式音频渲染方法及系统 |
EP4085660A4 (de) * | 2019-12-30 | 2024-05-22 | Comhear Inc. | Verfahren zum bereitstellen eines räumlichen schallfeldes |
US11246001B2 (en) | 2020-04-23 | 2022-02-08 | Thx Ltd. | Acoustic crosstalk cancellation and virtual speakers techniques |
US11750971B2 (en) * | 2021-03-11 | 2023-09-05 | Nanning Fulian Fugui Precision Industrial Co., Ltd. | Three-dimensional sound localization method, electronic device and computer readable storage |
CN115346537A (zh) * | 2021-05-14 | 2022-11-15 | 华为技术有限公司 | 一种音频编码、解码方法及装置 |
CN115497485B (zh) * | 2021-06-18 | 2024-10-18 | 华为技术有限公司 | 三维音频信号编码方法、装置、编码器和系统 |
Family Cites Families (43)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
GB9204485D0 (en) | 1992-03-02 | 1992-04-15 | Trifield Productions Ltd | Surround sound apparatus |
US6072878A (en) * | 1997-09-24 | 2000-06-06 | Sonic Solutions | Multi-channel surround sound mastering and reproduction techniques that preserve spatial harmonics |
EP1275272B1 (de) * | 2000-04-19 | 2012-11-21 | SNK Tech Investment L.L.C. | Mehrkanalige raumklangsmischung und wiedergabeverfahren mit einhaltung der räumlichen harmonischen in drei dimensionen |
JP3624805B2 (ja) | 2000-07-21 | 2005-03-02 | ヤマハ株式会社 | 音像定位装置 |
US7113610B1 (en) | 2002-09-10 | 2006-09-26 | Microsoft Corporation | Virtual sound source positioning |
FR2847376B1 (fr) | 2002-11-19 | 2005-02-04 | France Telecom | Procede de traitement de donnees sonores et dispositif d'acquisition sonore mettant en oeuvre ce procede |
US20040264704A1 (en) * | 2003-06-13 | 2004-12-30 | Camille Huin | Graphical user interface for determining speaker spatialization parameters |
US8054980B2 (en) | 2003-09-05 | 2011-11-08 | Stmicroelectronics Asia Pacific Pte, Ltd. | Apparatus and method for rendering audio information to virtualize speakers in an audio system |
GB0419346D0 (en) | 2004-09-01 | 2004-09-29 | Smyth Stephen M F | Method and apparatus for improved headphone virtualisation |
US7928311B2 (en) | 2004-12-01 | 2011-04-19 | Creative Technology Ltd | System and method for forming and rendering 3D MIDI messages |
JP2008118166A (ja) | 2005-06-30 | 2008-05-22 | Pioneer Electronic Corp | スピーカエンクロージャ、これを備えたスピーカシステムおよび多チャンネルステレオシステム |
US7693709B2 (en) | 2005-07-15 | 2010-04-06 | Microsoft Corporation | Reordering coefficients for waveform coding or decoding |
JP4674505B2 (ja) | 2005-08-01 | 2011-04-20 | ソニー株式会社 | 音声信号処理方法、音場再現システム |
US9215544B2 (en) * | 2006-03-09 | 2015-12-15 | Orange | Optimization of binaural sound spatialization based on multichannel encoding |
DE102006053919A1 (de) * | 2006-10-11 | 2008-04-17 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Vorrichtung und Verfahren zum Erzeugen einer Anzahl von Lautsprechersignalen für ein Lautsprecher-Array, das einen Wiedergaberaum definiert |
DE102007059597A1 (de) | 2007-09-19 | 2009-04-02 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Eine Vorrichtung und ein Verfahren zur Ermittlung eines Komponentensignals in hoher Genauigkeit |
EP2056627A1 (de) | 2007-10-30 | 2009-05-06 | SonicEmotion AG | Verfahren und Vorrichtung für erhöhte Klangfeldwiedergabepräzision in einem bevorzugtem Zuhörbereich |
JP4338102B1 (ja) | 2008-08-25 | 2009-10-07 | 薫 長山 | スピーカーシステム |
US8391500B2 (en) * | 2008-10-17 | 2013-03-05 | University Of Kentucky Research Foundation | Method and system for creating three-dimensional spatial audio |
WO2010070225A1 (fr) | 2008-12-15 | 2010-06-24 | France Telecom | Codage perfectionne de signaux audionumeriques multicanaux |
EP2205007B1 (de) * | 2008-12-30 | 2019-01-09 | Dolby International AB | Verfahren und Vorrichtung zur Kodierung dreidimensionaler Hörbereiche und zur optimalen Rekonstruktion |
GB2478834B (en) * | 2009-02-04 | 2012-03-07 | Richard Furse | Sound system |
WO2010092014A2 (en) * | 2009-02-11 | 2010-08-19 | Basf Se | Pesticidal mixtures |
GB0906269D0 (en) | 2009-04-09 | 2009-05-20 | Ntnu Technology Transfer As | Optimal modal beamformer for sensor arrays |
JP2010252220A (ja) * | 2009-04-20 | 2010-11-04 | Nippon Hoso Kyokai <Nhk> | 3次元音響パンニング装置およびそのプログラム |
US8971551B2 (en) | 2009-09-18 | 2015-03-03 | Dolby International Ab | Virtual bass synthesis using harmonic transposition |
EP2486561B1 (de) | 2009-10-07 | 2016-03-30 | The University Of Sydney | Rekonstruktion eines aufgezeichneten schallfelds |
US20110091055A1 (en) * | 2009-10-19 | 2011-04-21 | Broadcom Corporation | Loudspeaker localization techniques |
WO2011117399A1 (en) | 2010-03-26 | 2011-09-29 | Thomson Licensing | Method and device for decoding an audio soundfield representation for audio playback |
WO2011152044A1 (ja) * | 2010-05-31 | 2011-12-08 | パナソニック株式会社 | 音響再生装置 |
US9271081B2 (en) | 2010-08-27 | 2016-02-23 | Sonicemotion Ag | Method and device for enhanced sound field reproduction of spatially encoded audio input signals |
EP2450880A1 (de) | 2010-11-05 | 2012-05-09 | Thomson Licensing | Datenstruktur für Higher Order Ambisonics-Audiodaten |
JP2012104871A (ja) * | 2010-11-05 | 2012-05-31 | Sony Corp | 音響制御装置及び音響制御方法 |
TWI517028B (zh) | 2010-12-22 | 2016-01-11 | 傑奧笛爾公司 | 音訊空間定位和環境模擬 |
EP2541547A1 (de) * | 2011-06-30 | 2013-01-02 | Thomson Licensing | Verfahren und Vorrichtung zum Ändern der relativen Standorte von Schallobjekten innerhalb einer Higher-Order-Ambisonics-Wiedergabe |
EP2692155B1 (de) | 2012-03-22 | 2018-05-16 | Dirac Research AB | Entwurf für eine audiovorkompensierungssteuerung mit einem variablen satz unterstützender lautsprecher |
US9473870B2 (en) * | 2012-07-16 | 2016-10-18 | Qualcomm Incorporated | Loudspeaker position compensation with 3D-audio hierarchical coding |
CN104584588B (zh) | 2012-07-16 | 2017-03-29 | 杜比国际公司 | 用于渲染音频声场表示以供音频回放的方法和设备 |
CA2893729C (en) * | 2012-12-04 | 2019-03-12 | Samsung Electronics Co., Ltd. | Audio providing apparatus and audio providing method |
US9913064B2 (en) | 2013-02-07 | 2018-03-06 | Qualcomm Incorporated | Mapping virtual speakers to physical speakers |
US10499176B2 (en) | 2013-05-29 | 2019-12-03 | Qualcomm Incorporated | Identifying codebooks to use when coding spatial components of a sound field |
EP2866475A1 (de) | 2013-10-23 | 2015-04-29 | Thomson Licensing | Verfahren und Vorrichtung zur Decodierung einer Audioschallfelddarstellung für Audiowiedergabe mittels 2D-Einstellungen |
US20150264483A1 (en) | 2014-03-14 | 2015-09-17 | Qualcomm Incorporated | Low frequency rendering of higher-order ambisonic audio data |
-
2014
- 2014-02-06 US US14/174,775 patent/US9913064B2/en not_active Expired - Fee Related
- 2014-02-06 US US14/174,784 patent/US9736609B2/en active Active
- 2014-02-07 WO PCT/US2014/015315 patent/WO2014124268A1/en active Application Filing
- 2014-02-07 WO PCT/US2014/015311 patent/WO2014124264A1/en active Application Filing
- 2014-02-07 TW TW103104152A patent/TWI611706B/zh not_active IP Right Cessation
- 2014-02-07 KR KR1020157023104A patent/KR101877604B1/ko active IP Right Grant
- 2014-02-07 TW TW103104151A patent/TWI538531B/zh not_active IP Right Cessation
- 2014-02-07 EP EP14707033.8A patent/EP2954702B1/de active Active
- 2014-02-07 JP JP2015557125A patent/JP6309545B2/ja not_active Expired - Fee Related
- 2014-02-07 CN CN201480007510.XA patent/CN104969577B/zh not_active Expired - Fee Related
- 2014-02-07 EP EP14707870.3A patent/EP2954703B1/de active Active
- 2014-02-07 CN CN201480006477.9A patent/CN104956695B/zh not_active Expired - Fee Related
- 2014-02-07 JP JP2015557126A patent/JP6284955B2/ja not_active Expired - Fee Related
- 2014-02-07 KR KR1020157023103A patent/KR20150115822A/ko active IP Right Grant
Non-Patent Citations (1)
Title |
---|
None * |
Also Published As
Publication number | Publication date |
---|---|
CN104956695B (zh) | 2017-06-06 |
CN104956695A (zh) | 2015-09-30 |
KR20150115823A (ko) | 2015-10-14 |
TWI611706B (zh) | 2018-01-11 |
US9913064B2 (en) | 2018-03-06 |
CN104969577A (zh) | 2015-10-07 |
WO2014124268A1 (en) | 2014-08-14 |
JP6284955B2 (ja) | 2018-02-28 |
TWI538531B (zh) | 2016-06-11 |
TW201436587A (zh) | 2014-09-16 |
WO2014124264A1 (en) | 2014-08-14 |
KR20150115822A (ko) | 2015-10-14 |
CN104969577B (zh) | 2017-05-10 |
EP2954702A1 (de) | 2015-12-16 |
US20140219456A1 (en) | 2014-08-07 |
US20140219455A1 (en) | 2014-08-07 |
JP2016509820A (ja) | 2016-03-31 |
EP2954703A1 (de) | 2015-12-16 |
US9736609B2 (en) | 2017-08-15 |
EP2954702B1 (de) | 2019-06-05 |
JP6309545B2 (ja) | 2018-04-11 |
JP2016509819A (ja) | 2016-03-31 |
TW201436588A (zh) | 2014-09-16 |
KR101877604B1 (ko) | 2018-07-12 |
Similar Documents
Publication | Publication Date | Title |
---|---|---|
EP2954703B1 (de) | Bestimmung von renderern für kugelflächenfunktionskoeffizienten | |
EP2954521B1 (de) | Signalisierung von audiowiedergabeinformationen in einem bitstrom | |
US10477310B2 (en) | Ambisonic signal generation for microphone arrays | |
US20150264483A1 (en) | Low frequency rendering of higher-order ambisonic audio data | |
EP3668124A1 (de) | Bildschirmbezogene anpassung von hoa-inhalt | |
KR101818877B1 (ko) | 고차 앰비소닉 오디오 렌더러들에 대한 희소성 정보의 획득 | |
US11122386B2 (en) | Audio rendering for low frequency effects | |
KR101941764B1 (ko) | 고차 앰비소닉 오디오 렌더러들에 대한 대칭성 정보의 획득 | |
US9466302B2 (en) | Coding of spherical harmonic coefficients | |
EP3987824B1 (de) | Audiowiedergabe für niedrigfrequente effekte |
Legal Events
Date | Code | Title | Description |
---|---|---|---|
PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
17P | Request for examination filed |
Effective date: 20150716 |
|
AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
AX | Request for extension of the european patent |
Extension state: BA ME |
|
DAX | Request for extension of the european patent (deleted) | ||
STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
17Q | First examination report despatched |
Effective date: 20171004 |
|
GRAP | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOSNIGR1 |
|
STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: GRANT OF PATENT IS INTENDED |
|
INTG | Intention to grant announced |
Effective date: 20190711 |
|
GRAS | Grant fee paid |
Free format text: ORIGINAL CODE: EPIDOSNIGR3 |
|
GRAA | (expected) grant |
Free format text: ORIGINAL CODE: 0009210 |
|
STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE PATENT HAS BEEN GRANTED |
|
AK | Designated contracting states |
Kind code of ref document: B1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
REG | Reference to a national code |
Ref country code: CH Ref legal event code: EP |
|
REG | Reference to a national code |
Ref country code: IE Ref legal event code: FG4D |
|
REG | Reference to a national code |
Ref country code: DE Ref legal event code: R096 Ref document number: 602014058517 Country of ref document: DE |
|
REG | Reference to a national code |
Ref country code: AT Ref legal event code: REF Ref document number: 1215926 Country of ref document: AT Kind code of ref document: T Effective date: 20200115 |
|
REG | Reference to a national code |
Ref country code: NL Ref legal event code: MP Effective date: 20191218 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: BG Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20200318 Ref country code: FI Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: GR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20200319 Ref country code: NO Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20200318 Ref country code: LV Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: SE Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: LT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: DE Payment date: 20200210 Year of fee payment: 7 |
|
REG | Reference to a national code |
Ref country code: LT Ref legal event code: MG4D |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: HR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: RS Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: AL Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: EE Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: NL Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: RO Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: CZ Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: PT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20200513 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: IS Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20200418 Ref country code: SM Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: SK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
REG | Reference to a national code |
Ref country code: DE Ref legal event code: R097 Ref document number: 602014058517 Country of ref document: DE |
|
REG | Reference to a national code |
Ref country code: CH Ref legal event code: PL |
|
REG | Reference to a national code |
Ref country code: AT Ref legal event code: MK05 Ref document number: 1215926 Country of ref document: AT Kind code of ref document: T Effective date: 20191218 |
|
PLBE | No opposition filed within time limit |
Free format text: ORIGINAL CODE: 0009261 |
|
STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT |
|
REG | Reference to a national code |
Ref country code: BE Ref legal event code: MM Effective date: 20200229 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: DK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: MC Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: ES Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: LU Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20200207 |
|
26N | No opposition filed |
Effective date: 20200921 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: AT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: LI Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20200229 Ref country code: SI Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: CH Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20200229 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: IT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: IE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20200207 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: BE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20200229 Ref country code: PL Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: FR Payment date: 20210120 Year of fee payment: 8 |
|
REG | Reference to a national code |
Ref country code: DE Ref legal event code: R119 Ref document number: 602014058517 Country of ref document: DE |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: DE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20210901 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: TR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: MT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 Ref country code: CY Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: MK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20191218 |
|
PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: FR Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20220228 |
|
PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: GB Payment date: 20240111 Year of fee payment: 11 |