EP4627806A1 - Apparatus, methods and computer programs for spatial audio processing - Google Patents
Apparatus, methods and computer programs for spatial audio processingInfo
- Publication number
- EP4627806A1 EP4627806A1 EP23809114.4A EP23809114A EP4627806A1 EP 4627806 A1 EP4627806 A1 EP 4627806A1 EP 23809114 A EP23809114 A EP 23809114A EP 4627806 A1 EP4627806 A1 EP 4627806A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- metadata
- audio
- audio signal
- spatial
- microphones
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/005—Circuits for transducers for combining the signals of two or more microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2430/00—Signal processing covered by H04R, not provided for in its groups
- H04R2430/20—Processing of the output signals of the acoustic transducers of an array for obtaining a desired directivity characteristic
- H04R2430/21—Direction finding using differential microphone array [DMA]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R25/00—Electric hearing aids
- H04R25/40—Arrangements for obtaining a desired directivity characteristic
- H04R25/407—Circuits for combining signals of a plurality of transducers
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
Definitions
- Examples of the disclosure relate to apparatus, methods and computer programs for spatial audio processing. Some relate to apparatus, methods and computer programs for distributed spatial audio processing.
- an apparatus comprising means for: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
- the metadata indicative of a direction of arrival for an audio source may be used to process at least one further audio signal from at least one further microphone to generate the further metadata.
- the further audio signal and the transmitted at least one of the obtained audio signals may be configured to be processed to determine a direction of arrival and if the direction of arrival is to the front or back of the microphones.
- the further audio signal and the transmitted at least one of the obtained audio signals may be configured to be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
- An indication of confidence in a determined direction of arrival for the audio source may be transmitted.
- the means may be for determining tracking information for the two or more microphones and enabling the tracking information to be transmitted.
- the direction of arrival may be determined for one or more dominant audio sources.
- a method comprising: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
- a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
- the spatial metadata may comprise information that enables the spatial characteristics of the audio source to be recreated by a playback device.
- the at least one audio signal and the at least one further audio signal may be received separately.
- At least one of the at least one audio signal and the at least one further audio signal may be received by a wireless communication link.
- the apparatus may be comprised within a device and the device may comprise at least one of: a user device; a mobile phone; a processing device, a capturing device, a playback device.
- examples of the disclosure there may be provided a method comprising: receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
- FIG. 1 shows an example system
- FIG. 5 shows an example method
- FIG. 7 shows an example headset
- FIG. 8 shows an example system
- FIG. 9 shows an example system
- FIG. 10 shows an example system
- FIG. 11 shows an example system
- FIG. 12 shows an example capturing device
- FIG. 13 shows example system
- FIGS 14A to 14E show an example test set up and results; and FIG. 15 shows an example apparatus.
- the first device 103 comprises two or more microphones 105, an apparatus 107 and a transceiver 109. Only components that are referred to in this description are shown in Fig. 1.
- the first device 103 could comprise additional components in examples of the disclosure.
- the microphones 105 can comprise any means that can be configured to detect acoustic signals.
- the microphones 105 can be configured to capture acoustic signals from one or more sound sources.
- the microphones 105 can be configured to detect acoustic sound signals and convert the acoustic signals into an output electric signal.
- the microphones 105 provide microphone signals as an output.
- the microphone signals can comprise audio signals.
- the first device 103 comprises two microphones 105.
- the first device 103 could comprise more than two microphones 105.
- the microphones 105 can be positioned in or on the first device 103 so as to enable spatial audio to be captured.
- the microphones 105 can be located at different positions in or on the first device 103.
- the microphones 105 can be located at opposite ends or on opposite sides of the first device 103.
- the first device 103 is configured so that the audio signals from the microphones 105 are provided to the apparatus 107 as an input. This enables the apparatus 107 to obtain two or more audio signals from two or more microphones 105.
- the apparatus 107 can comprise a controller comprising a processor and memory. Examples of an apparatus 107 are shown in Fig. 15.
- the apparatus 107 can be configured to enable control of the first device 103.
- the apparatus 107 can be configured to perform spatial processing of the audio signals obtained from the microphones 105.
- the apparatus 107 can comprise a time alignment mechanism.
- the time alignment mechanism can be configured to synchronise output signals with obtained metadata such that after data transmission to the second device 111 audio signals and metadata received at device 111 are synchronised. Synchronisation can be based on known delays caused by audio signal and metadata transmission and/or processing and can require delaying either audio signals or metadata. In some examples it is possible to create time offset information for alignment of audio and metadata information in the second device 111. In this description it can be assumed that synchronisation is part of audio signal and metadata encoding and decoding processes.
- the transceiver 109 can be configured to enable wireless communications.
- the transceiver can enable low power wireless communications such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC) or any other suitable protocol.
- the second device 111 can be configured to receive the signals from the first device 103.
- the second device 111 could be an audio playback device or an audio capture device or a processing device or any other suitable type of device.
- the second device 111 could have more computational resources than the first device 103.
- the second device 111 comprises a transceiver 113 and an apparatus 115. Only components that are referred to in this description are shown in Fig. 1.
- the second device 111 could comprise additional components in examples of the disclosure.
- the transceiver 113 can comprise any means that can enable data to be received from the first device 103.
- the data that is received can comprise audio signals captured by the microphones 105 of the first device 103, metadata obtained from processing of the audio signals that is performed by the apparatus 107 of the first device 103, and/or any other suitable data.
- the transceiver 113 can be configured to enable wireless communications.
- the transceiver can enable low power wireless communications such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC) or any other suitable protocol.
- the second device 111 is configured so that the transceiver 113 can provide input signals to the apparatus 115.
- the apparatus 115 can comprise a controller comprising a processor and memory. Examples of an apparatus 115 are shown in Fig. 15.
- the apparatus 115 can be configured to enable control of the second device 111.
- the apparatus 115 can be configured to perform spatial processing of the audio signals and the metadata received from the first device 103.
- system 101 can comprise additional components that are not shown in Fig. 1.
- system 101 could comprise one or more further apparatus that can also be configured to obtain audio signals. The audio signals can then be provided to the second device 111 and used for spatial processing.
- the time domain signals can be converted to the frequency domain using any suitable transforms.
- the time domain signals can be converted to the frequency domain using Short-Time Fourier Transform (STFT), Quadrature Mirror Filter (QMF) or any other suitable means.
- STFT Short-Time Fourier Transform
- QMF Quadrature Mirror Filter
- the resulting time-frequency domain microphone signals can be denoted as Si(b, ri), where i is the microphone channel index, b is the frequency bin index, and n is the temporal frame index.
- the value of b is in the range 0, ... , B - 1 , where B is the number of bin indexes at every time index n.
- the respective subbands comprise one or more frequency bins.
- a given subband k has a lowest bin b k i ow and a highest bin b kihigh .
- the widths of the subbands can be selected based on properties of human hearing, for example equivalent rectangular bandwidth (ERB) or Bark scale can be used.
- the spatial analysis can be used to determine the direction of a dominant sound source. This can be done using signals from pairs of microphones. The most dominant direction for respective temporal frame indices is estimated. This can be done by searching a time shift t k that maximizes the correlation between the two microphone channels for the subband k. Si(b, ri) can be shifted by T samples as follows:
- the angle of the first direction can be defined as
- Information from the analysis of the other pairs of microphones can be used to remove the sign ambiguity in e ⁇ k, n). That is, it can be used to resolve whether the dominant direction is to the front of the first and second microphones or if it is to the rear of the first and second microphones.
- Fig. 2 schematically shows how the analysis between respective pairs of microphones can be used to resolve this directional ambiguity.
- three microphones 105 are shown. It is assumed that all of the microphones 105 are within the same horizontal plane.
- the dominant direction of arrival is on the same side as third microphone 105-3. In this case the dominant direction would be in front of the first microphone 105-1 and the second microphone 105-2. In this case the correct angle would be a.
- the spatial analysis process can use inference between pairs of microphones to determine the correct angle, d ⁇ k. n) -> e ⁇ k. n).
- an energy ratio r ⁇ k. n) corresponding to angle e ⁇ k, ri) can be estimated.
- the energy ratio can be estimated using a normalized correlation value c(/c, r) or any other suitable means.
- the normalized correlation value could be determined using:
- r ⁇ k. n is between -1 and 1 , and can be further limited between 0 and 1.
- different parts of this spatial analysis can be distributed between different devices within the system 101.
- the direction analysis for the first pair of microphones can be done in the first device 103 and the directional analysis between a second pair of microphones can be done in the second device 111.
- the overall conclusion of the sound source direction can be done when the outputs of the different analysis are combined.
- This reasoning is described for free field conditions that also apply for devices with acoustically non-transparent mechanics.
- the distance metric used in the above reasoning relates to acoustic distance considering the acoustic propagation of the acoustic waveform including also diffraction from product mechanics and therefore is not limited to free field radiation.
- Fig. 3 shows an example method according to examples of the disclosure. The method could be implemented using an apparatus 107 of a first device 103 as shown in Fig. 1 or by using any other suitable means.
- the method comprises, at block 301 obtaining two or more audio signals from two or more microphones 105.
- the microphones 105 can be spatially positioned as so as to enable spatial information to be obtained.
- the microphones 105 could be positioned at different ends or on different sides of the device 103.
- processing the obtained audio signal can comprise determining whether the direction of arrival is to the front or the rear of the two or more microphones 105.
- the processing that is performed initially could provide an indication of whether the direction of arrival is at the front or at the back but need not provide an indication of an angle to the right or left.
- the method comprises enabling transmission of at least one of the obtained audio signals with the metadata.
- the apparatus 107 can control the transceiver 109 to control the transmission of the audio signals and the metadata.
- the at least one of the obtained audio signals and the metadata can be transmitted via a wireless communication link.
- the processing device 111 can also be configured to receive at least one further audio signal from at least one further microphone.
- the further microphone that is used to obtain the further audio signal can be positioned in a different location to the two or more microphones 105 that are used to obtain the two or more audio signals.
- the further microphone could be part of a different device to the first device 103. For instance, if the first device 103 is an earpiece for a left ear then the further device could be an ear piece for the right ear. In such examples the audio signals from the further microphone could be transmitted from the different device to the second device 111.
- the further microphone 105 could be part of the second device 111 so that the further audio signals do not need to be transmitted via a wireless connection.
- the audio signal and the metadata and the further audio signals could be transmitted together.
- the respective audio signals could be obtained by microphones 105 in different parts of the same first device 103.
- the respective audio signals and any metadata that has been generated from the audio signals can be encoded and transmitted together.
- the obtained audio signal and the metadata indicative of a direction of arrival are configured to be processed to obtain further metadata.
- the obtained audio signal and the metadata indicative of a direction of arrival can be processed by the second device 111 to which they are transmitted, or by any other suitable device.
- the second device 111 can be configured to use the audio signals and the metadata indicative of the direction and the further audio signals to perform spatial processing.
- the second device 111 can be configured to perform the parts of the spatial processing that have not been performed by the first device 103.
- the second device 111 can be configured to resolve ambiguity in the angles determined by the first device 103, to perform energy estimations, and/or to perform any other suitable parts of the spatial processing.
- the metadata indicates a pair of angles the spatial processing that is performed by the second device 111 can resolve between the pair of angles. For example, it can determine if a direction of arrival should be to the front or the back of the microphones 105 of the first device 103.
- the first device 103 can perform spatial analysis using signals from a first microphone 105 and a second microphone 105.
- the second device 111 or other suitable device, can complete the spatial analysis using signals from the first microphone 105 and a third microphone. Therefore, the first device 103 only needs to transmit the signals from the first microphone 105 to the second device 111. The signals from the second microphone do not need to be transmitted to the second device 111.
- the further audio signal and the transmitted at least one of the obtained audio signals can be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
- the audio signals could be converted from a binaural format to a stereo format or between any other suitable formats.
- the first device 103 could comprise tracking means that are configured to track motion of the first device 103. Information relating to the motion of the first device 103 could then be included with the transmitted audio signals and the metadata.
- the metadata can comprise information that is obtained by performing part of a spatial analysis process.
- the metadata can be indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal.
- the metadata can indicate a pair of directions but does not resolve between the pair of directions. For example, it does not indicate if the direction of arrival for a sound source is to the front of the first device 103 or to the back of the first device 103.
- the method also comprises, at block 405, processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
- the spatial metadata comprises information that enables the spatial characteristics of the audio source to be recreated by a playback device.
- the processing that is performed by the second device 111 can comprise parts of the spatial processing that have not been performed by the first device 103.
- the processing performed by the second device 111 can comprise resolving ambiguity in the angles determined by the first device 103, performing energy estimations, and/or performing any other suitable parts of the spatial processing.
- the metadata indicates a pair of angles the spatial processing that is performed by the second device 111 can resolve between the pair of angles. For example, it can determine if a direction of arrival should be to the front or the back of the microphones 105 of the first device 103.
- the further audio signal and the audio signal associated with the metadata can be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
- the audio signals could be converted from a binaural format to a stereo format or between any other suitable formats.
- the processing that is performed can comprise part of a parametric spatial analysis.
- the audio signals that are received at block 501 can comprise an audio signal from a microphone 105 at the front of the first device 103 and an audio signal from a microphone 105 at the back of the first device 103.
- the processing that is performed can comprise a front-back analysis.
- the front-back analysis can determine a whether a direction of arrival is to the front of the first device 103 or to the rear of the first device 103.
- the metadata that is indicated at block 501 can indicate whether the direction of arrival is to the front or to the back of the first device 103 for given time intervals. Single value such as 0 or 1 can be used to indicate if the direction of arrival is to the front or to the back.
- the metadata that is generated by the first device 103 might not be sufficient to enable the spatial characteristics of the audio sources to be recreated by a playback device. However, the metadata that is generated by the first device 103 can be used to process further audio signals to obtain the metadata that is sufficient to enable the spatial characteristics of the audio sources to be recreated by a playback device.
- the method comprises encoding and transmitting the audio signals and the metadata. Any suitable protocols can be used for the encoding and the transmission of the data stream comprising the audio signals and the metadata.
- the audio signals can be transmitted on their own.
- the audio signals can be encoded with one or more further audio signals and can be transmitted together.
- one of the audio signals that was used for the front back analysis could be transmitted with an audio signal obtained from a different side of the first device 103.
- the transmitted audio signals could therefore comprise an audio signal captured by a microphone on the left-hand side of the first device 103 and an audio signal captured by a microphone 105 on the right-hand side of the first device 103.
- the method comprises receiving and decoding the data stream at the second device 111.
- the second device can receive the respective signals and the associated metadata in a single data stream.
- the different audio signals can be received in different data streams.
- the method comprises processing the decoded audio signals to generate further metadata.
- the further metadata can be generated by performing parts of the parametric spatial analysis that are not performed by the first device 103. For instance, in some examples the first device 103 analyses a signal from a microphone 105 at the front of the first device 103 and an audio signal from a microphone 105 at the back of the first device 103 to perform a front-back analysis. In such cases the second device 111 could then analyse a signal from a microphone 105 at the right hand side of the first device 103 and an audio signal from a microphone 105 at the left hand side of the first device 103 to perform a left-right analysis. The leftright analysis can determine a dominant direction of arrival for a sound source but does not determine if that angle is to the front of the first device 103 or to the back of the first device 103.
- the further metadata that is generated at block 107 would therefore indicate a direction of arrival at respective time intervals.
- the angles could be automatically assigned to a front sector because, at this point in the processing, it is not known if the direction of arrival is to the front or to the back of the first device 103.
- the respective metadata is combined to generate complete spatial metadata. That is, by combining the metadata generated by the first device 103 and the metadata generated by the second device 111 it can be determined if the directions determined in the left-right analysis should be in a front sector or a rear sector.
- the complete spatial metadata can comprise sufficient information that enables the spatial characteristics of the audio sources to be recreated by a playback device.
- the method comprises encoding and transmitting the audio signals and the complete metadata.
- Any suitable protocols can be used for the encoding and the transmission of the data stream comprising the audio signals and complete metadata.
- the encoding used could comprise Immersive Voice and Audio Services (IVAS) or any other suitable protocol.
- IVAS Immersive Voice and Audio Services
- Fig. 6 schematically shows an example system 101 that could be used in some examples of the disclosure.
- the system 101 comprises a first device 103, a second device 111 and a further device 601 .
- the first device 103 comprises a first ear piece of a headset and the further device 601 comprises a second ear piece of the headset.
- the first device 103 could be the left ear piece and the further device 601 could be the right ear piece.
- the respective ear pieces are in use they are positioned on either side of a user’s head.
- Both the first device 103 and the further device 601 comprise microphones 105.
- the first device 105 comprises two microphones 105 and the further device 601 comprises just one microphone 105.
- Other numbers and arrangements of the microphones 105 could be used in examples of the disclosure.
- the microphones of the first device 103 are positioned to enable spatial audio signals to be captured.
- Fig. 7 shows an example of an earpiece and an example arrangement of the microphones 105 that could be used in some examples.
- the first device 103 also comprises an apparatus 107.
- the apparatus 107 is configured to obtain audio signals from the respective microphones 105 of the first device 103 and perform part of the spatial analysis process.
- the apparatus 107 can be configured to perform front-back analysis that determines if a dominant direction for an audio source is to the front of the first device 103 or to the rear of the first device 103.
- the further device 601 is configured to encode and transmit a further audio signal 607.
- the further device 601 is a right ear piece and so the further audio signal 607 is a right audio signal 607.
- the right audio signal 607 can be transmitted by a wireless communication link or by any other suitable means.
- the right audio signal 607 could be transmitted together with the left audio signals 605.
- the respective earpieces could be part of the same headset.
- the respective ear pieces might not have any physical connection to each other and so the right audio signal 607 and the left audio signal 605 could be transmitted separately via different communication links.
- the system 101 is configured so that the metadata 603, the left audio signal 605 and the right audio signal 607 are received by the second device 111.
- the second device 111 comprises an apparatus 115 which can be configured to process the left audio signal 605 and the right audio signal 607.
- the apparatus 115 can use the respective audio signals 605, 607 to perform a left-right analysis.
- the outcome of the left-right analysis can therefore be combined with the metadata to determine complete metadata for the respective audio signals 605, 607.
- the second device 111 can therefore complete the spatial processing that is started by the first device 103.
- the complete metadata that is generated by the second device 111 is sufficient to be used with the audio signals enable reproduction of the spatial characteristics of the audio sources by a playback device.
- the complete spatial metadata can be used for beam-forming, head tracking, orientation tracking, or for any other suitable spatial process.
- Fig. 7 shows an example earpiece and an example arrangement of the microphones 105 that could be used to implement examples of the disclosure.
- the earpiece is a left ear piece.
- the earpiece could be a first device 103 as shown in Figs. 1 and 6.
- the earpiece comprises a first microphone 105-1 and second microphone 105-2.
- the first microphone 105-1 is positioned at a first end of the earpiece and is located in the user’s ear 701 when the earpiece is in use.
- the second microphone 105-2 is positioned at a second end of the earpiece.
- the second microphone 105-2 can be directed towards the user’s mouth.
- Fig. 8 schematically shows another example system 101 that could be used in some examples of the disclosure.
- This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a second device 111 and a further device 601 .
- the further device 601 comprises two microphones 105 that are configured to obtain audio signals that can be used for spatial audio analysis.
- one microphone 105 is positioned at the front of the further device 601 and one microphone 105 is positioned at the back of the further device 601 .
- the apparatus 801 of the further device 601 is configured to generate metadata 803 similar to the metadata 603 that is generated by the first device 103.
- the metadata 803 from the further device 601 indicates whether the direction of arrival is to the front or the back of the further device 601.
- the respective audio signals 605, 607 and the complete metadata can then be provided to an encoder 609 and then transmitted or stored as appropriate.
- Fig. 9 schematically shows another example system 101 that could be used in some examples of the disclosure.
- This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a second device 111 and a further device 601 where both the first device 103 and the further device 601 comprise at least two microphones 105 and generate metadata 603, 803 from a partial spatial analysis at the respective devices 103, 601.
- the confidence information can comprise a value.
- the value can be within any suitable range.
- the confidence value can be in the range [0,1],
- the confidence information can be encoded with the metadata 603, 803 and transmitted with the audio signals 605, 607 and the associated metadata 603, 607.
- the second device 111 can therefore receive the confidence information with the metadata 603, 803 and the associated audio signals 605, 607.
- the second device 111 can be configured to take the confidence information into account when selecting which metadata 603, 803 to use to complete the spatial analysis. This can be useful if there is for example wind noise or scratching noises present which can affect the reliability of the spatial information obtained from the audio signals.
- Fig. 10 schematically shows another example system 101 that could be used in some examples of the disclosure.
- This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a further device 601 , a second device 111 and an intermediate device 1001.
- the intermediate device 1001 is provided between the first device 103 and the second device 111 so that signals from the first device 103 are received by the intermediate device 1001 before they are forwarded on to the second device 111.
- the intermediate device could be a capturing device.
- the intermediate device 1001 could be a device comprising a camera that is being used to capture videos and corresponding audio.
- the second device 111 could be a playback device that is configured to process the audio signals and playback the spatial audio so that it can be heard by a user of the playback device.
- the first device 103 and the further device 601 can be configured to obtain audio signals 605, 607 and generate metadata 603 as described above.
- the first device 103 comprises two microphones 105 and the further device 601 comprises one microphone 105 similar to the arrangement as shown in Fig. 6.
- Other configurations or arrangements could be used in other examples.
- the example devices 130, 601 shown in Figs. 8 and 9 could be used or any suitable variations or configurations of this.
- the first device 103 is configured to perform part of the spatial analysis process.
- the apparatus 107 of the first device 103 is configured to perform front- back analysis to determine if a dominant direction of arrival for a sound source is to the front of the first device 103 or to the rear of the first device 103.
- the first device 103 transmits the left audio signal 605 and the metadata 603 generated from the part of the spatial analysis process to the intermediate device 1001.
- the further device 601 also transmits the right audio signal 607 to the intermediate device 1001.
- the intermediate device 1001 is configured to receive the metadata 603, the left audio signal 605 and the right audio signal 607. In this example the intermediate device 1001 does not perform any of the spatial processing. Instead, the intermediate device 1001 encodes the metadata 603, the left audio signal 605 and the right audio signal 607 into a bitstream 611 and enables the bitstream 611 to be transmitted to a second device 111.
- the intermediate device 1001 could be a capturing device such as a mobile phone or any other suitable type of device.
- the intermediate device 1001 could be configured to communicate in a cellular network or via any other suitable wireless communications network.
- the intermediate device 1001 comprises a stereo encoder 1003 and a metadata encoder 1005.
- the stereo encoder 1003 is configured to encode the respective audio signals 605, 607
- the metadata encoder 1005 is configured to encode the metadata 603.
- the bit rate that is needed for the metadata 603 is low compared to the bit rate that would be needed for full spatial metadata such as 360 degree directional information and energy ratios.
- the second device 111 comprises a stereo decoding module 1007, a metadata decoding module 1009 and a spatial processing module 1011.
- the respective modules 1007, 1009, 1011 could be provided by an apparatus 115 or by any other suitable means.
- Fig. 10 can be used to avoid unnecessary processing if head tracking or spatial processing is not needed.
- the spatial processing can be performed by the playback device rather than a capture device and so the spatial processing is only performed if it is needed.
- Fig. 11 schematically shows another example system 101 that could be used in some examples of the disclosure.
- This system 101 is similar to the system shown in Fig. 9 and comprises a first device 103, a second device 111 and a further device 601.
- the system 101 of Fig. 11 differs from the system 101 of Fig. 9 in that in Fig. 11 the further device 601 is configured to obtain head tracking information 1101 and transmit the headtracking information 1101 with the right audio signals 607 and the metadata 803 and confidence information.
- the second device 111 receives the head tracking information 1101 and the respective audio signals 605, 607 and the metadata 693, 803 and confidence information from first device 103 and the further device 601 .
- the apparatus 115 of the second device 111 can be configured to use the received information to perform the spatial analysis.
- the apparatus 115 can also be configured to use the head tracking information 1101 to adjust the metadata to take into account the orientation of the user’s head.
- the compensation for the user’s head position using the head tracking information can be made by the second device 111.
- This can be an audio capture device.
- the compensation could be made by a different device such as a playback device.
- Fig. 12 schematically shows an example capturing device 1201 that could be used in some examples of the disclosure.
- the capturing device 1201 can be a first device 103 that can be configured to perform spatial analysis process.
- the capturing device 1201 could be a user electronic device such as mobile phone or any other suitable type of device.
- the binaural signals do not have to originate from a headset or other similar device. Instead signals obtained from any suitable microphone array can be synthesized into spatial audio signals.
- the example capturing device 1201 of Fig. 12 comprises a microphone array 1203, a spatial analysis module 1205, a binaural synthesis module 1207 and an encoder 1209.
- the spatial analysis module 1205 and the binaural synthesis module 1207 could be provided by an apparatus 107 or by any other suitable means.
- the microphone array 1203 comprises a plurality of microphones 105.
- the microphones 105 can be positioned in different locations within or around the capturing device 1201 so as to enable spatial audio capture.
- the audio signals from the microphone array 1203 are provided to the spatial analysis module 1205.
- the spatial analysis module 1205 can be configured to perform the spatial analysis process.
- Spatial analysis 1205 can be configured to produce spatial metadata 1215.
- the spatial analysis module 1205 can be configured to produce complete spatial metadata, however only a subset of the produced metadata needs to be provided from the spatial analysis module 1205.
- the spatial metadata can contain part of the analysed spatial information.
- the spatial analysis module 1205 could be configured to generate front -back metadata.
- Other types of metadata could be generated in some examples.
- the spatial analysis module 1205 could generate left-right metadata in some examples.
- the encoder 1209 is configured to encode the binaural signal 1211 and the metadata 603.
- the encoder 1209 could be an IVAS encoder or any other suitable type of encoder.
- Fig. 13 schematically shows another example system 101 that could be used in some examples of the disclosure.
- the first device 103 that performs the first part of the spatial processing does not need to be a headset or other similar device.
- the first device 103 could be a spatial voice conferencing server or any other suitable type of device.
- the spatial rendering module 1303 is configured to perform part of a spatial analysis process on the audio objects.
- the spatial rendering module 1303 is configured to perform spatial signal rendering using the spatial locations defined by the location definition module 1301.
- the spatial rendering module 1303 can also be configured to perform front-back analysis of the audio objects and provide the metadata indicative of the output of the analysis.
- the audio signals that have been rendered by the spatial rendering module 1303 and the metadata can then be provided to the encoder 1305.
- the encoder 1305 could be an IVAS encoder or any other suitable type of encoder.
- the stereo decoding module 1007 is configured to decode the received audio signals and the metadata decoding module 1009 is configured to decode the received metadata.
- the decoded audio signals and the decoded metadata can then be used by the spatial processing module 1011 to complete the spatial processing.
- the spatial processing that is performed by the second device 111 can complete the spatial processing that is started by the first device 103 so as to enable reproduction of the spatial characteristics of the audio sources by a playback device.
- Figs 14A to 14E show an example test set up and results that were obtained using the test set up and examples of the disclosure.
- Fig. 14A shows the example test set up.
- an artificial head 1401 is configured with two microphones 105 attached to the left ear and one microphone attached to the right ear. Only the microphones 105 attached to the left ear are shown in the view of Fig. 14A.
- the microphones are spatially separated so as to enable spatial audio signals to be captured.
- a test signal was recorded in which a sound source was originally located to the front of the artificial head 1401 and rotated around the artificial head 1401 in anticlockwise direction with about a 1.5 meter distance to the artificial head 1401.
- Fig. 14 B shows the three signals recorded by the respective microphones 105.
- Fig. 14. D shows the outcome of the left-right analysis. In this case it is not known whether the sound source is to the front or the back and so all the sources are located in the front. That is all of the angles have a value between -90 and 90 degrees.
- the final stage of the processing is performed by combining the front-back analysis and the left-right analysis.
- the front-back estimate is used to determine if the sound source is to the front or back.
- the results of the combination are shown in Fig. 14E. These clearly show that the sound source has travelled in a circle around the artificial head 1401.
- This spatial information can now be used do spatial processing such as audio focussing or headtracking or any other suitable type of processing.
- the spatial processing can be performed by a capture device or a playback device or any other suitable type of device.
- the memory 1505 stores a computer program 1507 comprising computer program instructions (computer program code) that controls the operation of the controller 1501 when loaded into the processor 1503.
- the computer program instructions, of the computer program 1507 provide the logic and routines that enables the controller 1501. to perform the methods illustrated in the accompanying Figs.
- the processor 1503 by reading the memory 1505 is able to load and execute the computer program 1507.
- the apparatus 107 comprises: at least one processor 1503; and at least one memory 1505 storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: obtaining 301 two or more audio signals from two or more microphones; processing 303 the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission 305 of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
- the apparatus 115 When the apparatus 115 is configured for use in a second device 111 the apparatus 115 comprises: at least one processor 1503; and at least one memory 1505 storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving 401 at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving 403 at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate 405 spatial metadata.
- the computer program 1507 can be transmitted to the controller 1501 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.
- a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.
- the computer program 1507 comprises computer program instructions for causing an apparatus 107 to perform at least the following or for performing at least the following: obtaining 301 two or more audio signals from two or more microphones; processing 303 the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission 305 of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
- the computer program instructions can be comprised in a computer program 1507, a non-transitory computer readable medium, a computer program product, a machine- readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1507.
- processor 1503 is illustrated as a single component/circuitry it can be implemented as one or more separate components/circuitry some or all of which can be integrated/removable.
- the processor 1503 can be a single core or multi-core processor.
- references to ‘computer-readable storage medium’, ‘computer program product’, ‘tangibly embodied computer program’ etc. or a ‘controller’, ‘computer’, ‘processor’ etc. should be understood to encompass not only computers having different architectures such as single /multi- processor architectures and sequential (Von Neumann)/parallel architectures but also specialized circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry.
- References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc.
- circuitry may refer to one or more or all of the following:
- any portions of hardware processor(s) with software including digital signal processor(s)
- software including digital signal processor(s)
- memory or memories that work together to cause an apparatus, such as a mobile phone or server, to perform various functions
- circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
- the blocks illustrated in Figs. 3 to 5 can represent steps in a method and/or sections of code in the computer program 1507.
- the illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted.
- connection means operationally connected/coupled/in communication.
- intervening components can exist (including no intervening components), i.e., so as to provide direct or indirect connection/coupling/communication. Any such intervening components can include hardware and/or software components.
Landscapes
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Otolaryngology (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
Examples of the disclosure relate to apparatus, methods and computer programs for spatial audio processing. In examples of the disclosure an apparatus is configured to receive at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal. The apparatus is also configured to receive at least one further audio signal that has been captured by at least one further microphone and process the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
Description
TITLE
Apparatus, Methods and Computer Programs for Spatial Audio Processing
TECHNOLOGICAL FIELD
Examples of the disclosure relate to apparatus, methods and computer programs for spatial audio processing. Some relate to apparatus, methods and computer programs for distributed spatial audio processing.
BACKGROUND
Spatial audio enables spatial properties of a sound scene to be reproduced for a user so that the user can perceive the spatial properties of the sound scene. This can provide an immersive audio experience for a user or could be used for other applications.
BRIEF SUMMARY
According to various, but not necessarily all, examples of the disclosure there may be provided an apparatus comprising means for: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
The spatial metadata may comprise information that enables the spatial characteristics of the audio source to be recreated by a playback device.
Processing the obtained audio signal may comprise determining whether the direction of arrival is to the front or the rear of the two or more microphones.
The metadata indicative of a direction of arrival for an audio source may be used to process at least one further audio signal from at least one further microphone to generate the further metadata.
The further microphone that is used to obtain the further audio signal may be positioned in a different location to the two or more microphones that are used to obtain the two or more audio signals.
The at least one of the obtained audio signals and the metadata may be transmitted separately from the further audio signal.
The further audio signal and the transmitted at least one of the obtained audio signals may be configured to be processed to determine a direction of arrival and if the direction of arrival is to the front or back of the microphones.
The further audio signal and the transmitted at least one of the obtained audio signals may be configured to be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
An indication of confidence in a determined direction of arrival for the audio source may be transmitted.
The means may be for determining tracking information for the two or more microphones and enabling the tracking information to be transmitted.
The direction of arrival may be determined for one or more dominant audio sources.
The at least one of the obtained audio signals and the metadata may be transmitted via a wireless communication link.
The apparatus may be comprised within a device and the device may comprise at least one of: a headset, ear pieces, a wearable device.
According to various, but not necessarily all, examples of the disclosure there may be provided a method comprising: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
According to various, but not necessarily all, examples of the disclosure there may be provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
According to various, but not necessarily all, examples of the disclosure there may be provided an apparatus comprising means for: receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal;
receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
Generating the spatial metadata may comprise obtaining further metadata by processing the at least one audio signal and the at least one further audio signal and using the metadata and the further metadata to generate spatial metadata.
The metadata may be indicative of a direction of arrival for an audio source indicates whether the direction of arrival is to the front or the rear of the microphones that were used to capture the at least one audio signal.
The spatial metadata may comprise information that enables the spatial characteristics of the audio source to be recreated by a playback device.
The at least one audio signal and the at least one further audio signal may be received separately.
Processing the at least one audio signal and the at least one further audio signal may comprise using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
At least one of the at least one audio signal and the at least one further audio signal may be received by a wireless communication link.
The apparatus may be comprised within a device and the device may comprise at least one of: a user device; a mobile phone; a processing device, a capturing device, a playback device.
According to various, but not necessarily all, examples of the disclosure there may be provided a method comprising:
receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
According to various, but not necessarily all, examples of the disclosure there may be provided a computer program comprising instructions which, when executed by an apparatus cause the apparatus to perform: receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all of the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all of the features, in any combination, may be implemented by/comprised in/performable by an apparatus, a method, and/or computer program instructions as desired, and as appropriate.
BRIEF DESCRIPTION
Some examples will now be described with reference to the accompanying drawings in which:
FIG. 1 shows an example system;
FIG. 2 shows estimating a direction of arrival using three microphones;
FIG. 3 shows an example method;
FIG. 4 shows an example method;
FIG. 5 shows an example method;
FIG. 6 shows an example system;
FIG. 7 shows an example headset;
FIG. 8 shows an example system;
FIG. 9 shows an example system;
FIG. 10 shows an example system;
FIG. 11 shows an example system;
FIG. 12 shows an example capturing device;
FIG. 13 shows example system;
FIGS 14A to 14E show an example test set up and results; and FIG. 15 shows an example apparatus.
The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures.
DETAILED DESCRIPTION
Examples of the disclosure relate to spatial processing where different parts of the spatial processing can be performed by different parts of a system. In examples of the disclosure a first device, such as an ear piece or headset, can perform some spatial processing on audio signals captured by microphones of the first device. The spatial processing can determine metadata indicative of a direction of arrival, or at least partially indicative of a position, of a sound source. The metadata can then be provided to a processing device and used for further spatial processing. The further spatial processing could complete a process and/or could more accurately determine a direction of arrival. Examples of the disclosure can reduce the data that needs to be transmitted between the respective devices of the system. As some of the processing can be carried out at the first device this enables metadata to be sent rather than all of the audio signals that have been captured.
Fig. 1 schematically shows an example system 101 that can be used to implement examples of the disclosure. In this example the system 101 comprises a first device 103 and a second device 111. The system 101 can also comprise additional devices and components that are not shown in Fig. 1.
In the example of Fig. 1 the first device 103 comprises two or more microphones 105, an apparatus 107 and a transceiver 109. Only components that are referred to in this description are shown in Fig. 1. The first device 103 could comprise additional components in examples of the disclosure.
The first device 103 could be a wearable device such as a headset, ear piece or other suitable type of device. The first device 103 can be a small or compact device so that only a small amount of processing and power resources are available to the first device 103. The amount of processing and power resources that are available at the first device are small compared to the resources available at the second device 111.
The microphones 105 can comprise any means that can be configured to detect acoustic signals. The microphones 105 can be configured to capture acoustic signals from one or more sound sources. The microphones 105 can be configured to detect acoustic sound signals and convert the acoustic signals into an output electric signal. The microphones 105 provide microphone signals as an output. The microphone signals can comprise audio signals.
In the example of Fig. 1 the first device 103 comprises two microphones 105. In some examples the first device 103 could comprise more than two microphones 105. The microphones 105 can be positioned in or on the first device 103 so as to enable spatial audio to be captured. The microphones 105 can be located at different positions in or on the first device 103. For example, the microphones 105 can be located at opposite ends or on opposite sides of the first device 103.
The first device 103 is configured so that the audio signals from the microphones 105 are provided to the apparatus 107 as an input. This enables the apparatus 107 to obtain two or more audio signals from two or more microphones 105.
The apparatus 107 can comprise a controller comprising a processor and memory. Examples of an apparatus 107 are shown in Fig. 15. The apparatus 107 can be configured to enable control of the first device 103. In examples of the disclosure the apparatus 107 can be configured to perform spatial processing of the audio signals obtained from the microphones 105.
The first device 103 is configured so that the apparatus 107 can provide input signals to and receive output signals from the transceiver 109. The transceiver 109 can comprise any means that can enable data to be transmitted from the first device 103. The data that is transmitted can comprise audio signals captured by the microphones 105, metadata obtained from processing of the audio signals that is performed by the apparatus 107, and/or any other suitable data.
The apparatus 107 can comprise a time alignment mechanism. The time alignment mechanism can be configured to synchronise output signals with obtained metadata such that after data transmission to the second device 111 audio signals and metadata received at device 111 are synchronised. Synchronisation can be based on known delays caused by audio signal and metadata transmission and/or processing and can require delaying either audio signals or metadata. In some examples it is possible to create time offset information for alignment of audio and metadata information in the second device 111. In this description it can be assumed that synchronisation is part of audio signal and metadata encoding and decoding processes.
The transceiver 109 can be configured to enable wireless communications. In some examples the transceiver can enable low power wireless communications such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC) or any other suitable protocol.
The second device 111 can be configured to receive the signals from the first device 103. The second device 111 could be an audio playback device or an audio capture device or a processing device or any other suitable type of device. The second device 111 could have more computational resources than the first device 103.
The second device 111 comprises a transceiver 113 and an apparatus 115. Only components that are referred to in this description are shown in Fig. 1. The second device 111 could comprise additional components in examples of the disclosure.
The transceiver 113 can comprise any means that can enable data to be received from the first device 103. The data that is received can comprise audio signals captured by the microphones 105 of the first device 103, metadata obtained from processing of the audio signals that is performed by the apparatus 107 of the first device 103, and/or any other suitable data.
The transceiver 113 can be configured to enable wireless communications. In some examples the transceiver can enable low power wireless communications such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC) or any other suitable protocol.
The second device 111 is configured so that the transceiver 113 can provide input signals to the apparatus 115. The apparatus 115 can comprise a controller comprising a processor and memory. Examples of an apparatus 115 are shown in Fig. 15. The apparatus 115 can be configured to enable control of the second device 111. In examples of the disclosure the apparatus 115 can be configured to perform spatial processing of the audio signals and the metadata received from the first device 103.
In some examples the system 101 can comprise additional components that are not shown in Fig. 1. For example, the system 101 could comprise one or more further apparatus that can also be configured to obtain audio signals. The audio signals can then be provided to the second device 111 and used for spatial processing.
The spatial processing that is performed by the respective apparatus 107, 115 can comprise any suitable processing. The spatial processing can comprise parametric spatial analysis or any other suitable process. In examples of the disclosure the spatial processing can be distributed so that part of the spatial processing is performed at the
first device 103 and part of the spatial processing is performed at the second device 111.
In examples of the disclosure the spatial processing can be performed in the frequency domain. Before the transform into the frequency domain the audio signals are obtained in the time domain. The time domain audio signals can be denoted Si(t), where t is the time index and / is the microphone channel index.
The time domain signals can be converted to the frequency domain using any suitable transforms. In some examples the time domain signals can be converted to the frequency domain using Short-Time Fourier Transform (STFT), Quadrature Mirror Filter (QMF) or any other suitable means.
The resulting time-frequency domain microphone signals can be denoted as Si(b, ri), where i is the microphone channel index, b is the frequency bin index, and n is the temporal frame index. The value of b is in the range 0, ... , B - 1 , where B is the number of bin indexes at every time index n. The frequency bins can be further combined into subbands k = 0, K - 1. The respective subbands comprise one or more frequency bins. A given subband k has a lowest bin bk iow and a highest bin bkihigh. The widths of the subbands can be selected based on properties of human hearing, for example equivalent rectangular bandwidth (ERB) or Bark scale can be used.
The spatial analysis can be used to determine the direction of a dominant sound source. This can be done using signals from pairs of microphones. The most dominant direction for respective temporal frame indices is estimated. This can be done by searching a time shift tk that maximizes the correlation between the two microphone channels for the subband k. Si(b, ri) can be shifted by T samples as follows:
The spatial analysis is configured to find the delay
for each subband k which maximises the correlation between two microphone channels:
In the above equation, the optimal delay is searched between a first microphone (microphone 1) and a second microphone (microphone 2). Re indicates the real part of the result, and * is the complex conjugate of the signal. The delay search range parameter Dmax is defined based on the distance between microphones. The value of Tk is searched only on the range which is physically possible considering the distance between the microphones and the speed of sound.
The angle of the first direction can be defined as
As shown, there is still uncertainty of the sign of the angle. That is, it is unclear whether the direction of arrival is to the front of the microphones (or an axis defied relative to the microphones) or to the rear of the microphones.
Information from the analysis of the other pairs of microphones can be used to remove the sign ambiguity in e^k, n). That is, it can be used to resolve whether the dominant direction is to the front of the first and second microphones or if it is to the rear of the first and second microphones.
Fig. 2 schematically shows how the analysis between respective pairs of microphones can be used to resolve this directional ambiguity. In the example of Fig. 2 three microphones 105 are shown. It is assumed that all of the microphones 105 are within the same horizontal plane.
Signals from the first microphone 105-1 and the second microphone 105-2 are used to determine the angle of the first direction as described above. This results in two possible angles for the dominant direction a and -a. That is, it is unresolved as to whether the dominant direction is in front of the first and second microphones 105-1 , 105-2 or to the back of the first and second microphones 105-1 , 105-2. To address this the signals received from the first microphone 105-1 and the third microphone 105-3
are analysed to determine if the acoustic signal arrives first at the first microphone 105- 1 or the third microphone 105-3.
If the signal arrives at the third microphone 105-3 before it arrives at the first microphone 105-1 then the dominant direction of arrival is on the same side as third microphone 105-3. In this case the dominant direction would be in front of the first microphone 105-1 and the second microphone 105-2. In this case the correct angle would be a.
Conversely if the signal arrives at the first microphone 105-1 before it arrives at the third microphone 105-3 then the dominant direction of arrival is on the opposite side as third microphone 105-3. In this case the dominant direction would be to the back of the first microphone 105-1 and the second microphone 105-2. In this case the correct angle would be -a.
Using this logic the spatial analysis process can use inference between pairs of microphones to determine the correct angle, d^k. n) -> e^k. n).
Once the correct angle has been determined an energy ratio r^k. n) corresponding to angle e^k, ri) can be estimated. The energy ratio can be estimated using a normalized correlation value c(/c, r) or any other suitable means. The normalized correlation value could be determined using:
The value of r^k. n) is between -1 and 1 , and can be further limited between 0 and 1.
In examples of the disclosure different parts of this spatial analysis can be distributed between different devices within the system 101. For example, the direction analysis for the first pair of microphones can be done in the first device 103 and the directional analysis between a second pair of microphones can be done in the second device 111. The overall conclusion of the sound source direction can be done when the outputs of the different analysis are combined.
This reasoning is described for free field conditions that also apply for devices with acoustically non-transparent mechanics. In general the distance metric used in the above reasoning relates to acoustic distance considering the acoustic propagation of the acoustic waveform including also diffraction from product mechanics and therefore is not limited to free field radiation.
Fig. 3 shows an example method according to examples of the disclosure. The method could be implemented using an apparatus 107 of a first device 103 as shown in Fig. 1 or by using any other suitable means.
The method comprises, at block 301 obtaining two or more audio signals from two or more microphones 105. The microphones 105 can be spatially positioned as so as to enable spatial information to be obtained. For example, the microphones 105 could be positioned at different ends or on different sides of the device 103.
At block 303 the method comprises processing the obtained audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones. Any suitable process can be used to determine the direction of arrival. In some examples the spatial analysis process described above could be used. Other processes or variations on this process could be used in some examples.
The processing that is performed at block 303 can be partial or limited spatial analysis. That is only some of the parts of the spatial analysis process are performed. In this example a direction of arrival could be estimated but it might not be resolved as to whether the direction should be in the positive direction or the negative direction. That is a pair of directions could be determined but the processing at the first device 103 does not resolve between the pair of directions. The output of the processing could be an angle but it is not determined whether the angle should be to the front of the microphones 105 or to the back of the microphones 105.
In some examples of the disclosure processing the obtained audio signal can comprise determining whether the direction of arrival is to the front or the rear of the two or more microphones 105. The processing that is performed initially could provide an indication
of whether the direction of arrival is at the front or at the back but need not provide an indication of an angle to the right or left.
The apparatus 107 can also be configured to generate metadata indicative of the determined direction information. The metadata can give an indication of the angle that has been determined. The metadata could indicate that it has not been resolved whether the angle should be positive or negative.
At block 305 the method comprises enabling transmission of at least one of the obtained audio signals with the metadata. The apparatus 107 can control the transceiver 109 to control the transmission of the audio signals and the metadata. The at least one of the obtained audio signals and the metadata can be transmitted via a wireless communication link.
The metadata can be transmitted in any suitable format. In some examples it could be transmitted as a single bit. For instance, a 1 could indicate that a direction of arrival is to the front and a 0 could indicate that a direction of arrival is to the back. This could result in the metadata only requiring a small amount of data to be transmitted. For example, it could add only an additional 24 bits per frame.
The audio signals and the metadata can be combined and transmitted together. The audio signals can be transmitted to a second device 111 which can be as shown in Fig. 1. The second device 111 could be a processing device or any other suitable type of device.
The processing device 111 can also be configured to receive at least one further audio signal from at least one further microphone. The further microphone that is used to obtain the further audio signal can be positioned in a different location to the two or more microphones 105 that are used to obtain the two or more audio signals. In some examples the further microphone could be part of a different device to the first device 103. For instance, if the first device 103 is an earpiece for a left ear then the further device could be an ear piece for the right ear. In such examples the audio signals from the further microphone could be transmitted from the different device to the second device 111.
In some examples the further microphone 105 could be part of the second device 111 so that the further audio signals do not need to be transmitted via a wireless connection.
The audio signal and the metadata can be transmitted separately to the further audio signals. For example, if the further audio signals are obtained by a different device or a different part of the first device 103 then different transceivers 109 and different communication links can be used to transmit the respective signals. The second device 111 can therefore also receive the audio signal and the metadata separately to the further audio signals.
In some examples the audio signal and the metadata and the further audio signals could be transmitted together. For example, the respective audio signals could be obtained by microphones 105 in different parts of the same first device 103. In such examples the respective audio signals and any metadata that has been generated from the audio signals can be encoded and transmitted together.
The obtained audio signal and the metadata indicative of a direction of arrival are configured to be processed to obtain further metadata. The obtained audio signal and the metadata indicative of a direction of arrival can be processed by the second device 111 to which they are transmitted, or by any other suitable device.
The metadata indicative of a direction of arrival for an audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source. The spatial metadata comprises information that enables the spatial characteristics of the audio source to be recreated by a playback device.
The second device 111 , or any other suitable device, can be configured to use the audio signals and the metadata indicative of the direction and the further audio signals to perform spatial processing. The second device 111 can be configured to perform the parts of the spatial processing that have not been performed by the first device 103. For example, the second device 111 can be configured to resolve ambiguity in the angles determined by the first device 103, to perform energy estimations, and/or
to perform any other suitable parts of the spatial processing. Where the metadata indicates a pair of angles the spatial processing that is performed by the second device 111 can resolve between the pair of angles. For example, it can determine if a direction of arrival should be to the front or the back of the microphones 105 of the first device 103.
In examples of the disclosure only a subset of the audio signals that are obtained from the microphones of the first device 103 need to be transmitted to the second device 111 for processing. Only the signals that are needed for parts of the spatial processing that are to be performed by the second device 111 need to be transmitted to the second device 111. For example, the first device 103 can perform spatial analysis using signals from a first microphone 105 and a second microphone 105. The second device 111 , or other suitable device, can complete the spatial analysis using signals from the first microphone 105 and a third microphone. Therefore, the first device 103 only needs to transmit the signals from the first microphone 105 to the second device 111. The signals from the second microphone do not need to be transmitted to the second device 111.
In some examples other spatial processes can be performed on the respective audio signals. For example, the further audio signal and the transmitted at least one of the obtained audio signals can be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format. For instance, the audio signals could be converted from a binaural format to a stereo format or between any other suitable formats.
In some examples other information can be transmitted with the audio signals and the metadata. For instance, in some examples an indication of the confidence in a determined direction of arrival can be transmitted. In some examples the first device 103 could comprise tracking means that are configured to track motion of the first device 103. Information relating to the motion of the first device 103 could then be included with the transmitted audio signals and the metadata.
The tracking means could be head tracking means or orientation tracking means or any other suitable type of tracking means. Head tracking means can be configured to
determine an angular position or 3D orientation of a user’s head. For example, head tracking means can be used to determine a direction in which a user is facing. Head tracking means can be used to compensate for the user’s head position and create a more stable spatial audio stream for the user.
Orientation tracking means can be used to determine an orientation of a device such as the first device 103. If the device 103 is being worn by a user then the orientation tracking means will give information about the orientation of a user’s body. Changes in the orientation of a user’s body might not be compensated for when providing a spatial audio stream to a user. However, they can be taken into account for audio zooming or for other suitable purposes.
In the example of Fig. 3 the at least one of the obtained audio signals and the metadata can be sent to a second device 111 for processing. In other examples there could be one or more intermediate devices between the first device 103 and the device that completes the spatial processing. In such examples the intermediate devices that receive the at least one of the obtained audio signals and the metadata could forward the audio signals and the metadata to the processing device. The intermediate device could also receive one or more of the further audio signals.
In some examples the further audio signals could be transmitted without any metadata. In other examples the further audio signals could be transmitted with metadata. The metadata could comprise information indicative of a direction of a sound source, confidence information, tracking information or any other suitable information.
In some examples the metadata could be transmitted without the audio signals. In such cases a first device can process audio signals obtained from two or more microphones and then process these to obtain the metadata. The metadata could then be transmitted on its own without the audio signals. The second device 111 that performs the spatial processing could then receive two or more audio signals and use the metadata to process these. These examples could enable low quality microphones to be used in the first device 103 to obtain the metadata while higher quality microphones can be used to obtain transport audio signals.
Fig. 4 shows another example method according to examples of the disclosure. The method could be implemented using an apparatus 115 of a second device 111 as shown in Fig. 1 or by using any other suitable means.
The method comprises, at block 401 , receiving at least one audio signal with associated metadata. The at least one audio signal and the associated metadata can be received from a first device 103 or from any other suitable device. The at least one audio signal and the associated metadata can be received via a wireless communication link or by any other suitable means.
The metadata can be associated with the at least one audio signal in that the metadata can be derived, at least in part, by processing the at least one audio signal. The metadata can comprise spatial information relating to audio sources represented by the at least one audio signal.
The metadata can comprise information that is obtained by performing part of a spatial analysis process. The metadata can be indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal. The metadata can indicate a pair of directions but does not resolve between the pair of directions. For example, it does not indicate if the direction of arrival for a sound source is to the front of the first device 103 or to the back of the first device 103.
In examples of the disclosure the metadata can be obtained from analysing two or more audio signals. Only a subset of the audio signals that were analysed to obtain the audio data need to be transmitted to the second device 111 because the metadata comprises some of the spatial information.
The method comprises, at block 403, receiving at least one further audio signal that has been captured by at least one further microphone. The further microphone that is used to obtain the further audio signal can be positioned in a different location to the device from which the audio signal and the associated metadata are received. In some examples the further microphone could be part of a different device to the first device 103. For instance, if the first device 103 is an earpiece for a left ear then the further device could be an ear piece for the right ear. In such examples the audio signals from
the further microphone could be transmitted from the different device to the second device 111.
In some examples the further microphone 105 could be part of the second device 111 so that the further audio signals do not need to be received via a wireless connection.
In some examples the audio signal and the metadata can be received separately to the further audio signals. For example, if the further audio signals are obtained by a different device or a different part of the first device 103 then different communication links can be used to receive the respective signals.
In some examples the audio signal and the metadata and the further audio signals could be received together. For example, the respective audio signals could be obtained by microphones 105 in different parts of the same first device 103. In such examples the respective audio signals and any metadata that has been generated from the audio signals can be received in the same signal.
The method also comprises, at block 405, processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata. The spatial metadata comprises information that enables the spatial characteristics of the audio source to be recreated by a playback device. The processing that is performed by the second device 111 can comprise parts of the spatial processing that have not been performed by the first device 103. For example, the processing performed by the second device 111 can comprise resolving ambiguity in the angles determined by the first device 103, performing energy estimations, and/or performing any other suitable parts of the spatial processing. Where the metadata indicates a pair of angles the spatial processing that is performed by the second device 111 can resolve between the pair of angles. For example, it can determine if a direction of arrival should be to the front or the back of the microphones 105 of the first device 103.
Other types of processing could be performed by the second device 111 in some examples using the respective audio signals. For example, the further audio signal and the audio signal associated with the metadata can be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second,
different spatial audio format. For instance, the audio signals could be converted from a binaural format to a stereo format or between any other suitable formats.
Fig. 5 shows another example method according to examples of the disclosure. The method could be implemented using the system 101 as shown in Fig. 1 or by using any other suitable means.
In the example of Fig. 5 blocks 501 and 503 of the method can be performed by a first device 103 and blocks 505 to 511 can be performed by a second device 111 or other suitable device. Other ways of distributing the respective parts of the process could be used in other examples.
At block 501 the method comprises receiving two or more audio signals and processing the audio signals to generate metadata.
The audio signals can be received from two or more microphones 105 within the same first device 103 or within the same part of the first device 103. For instance, if the first device 103 is an ear piece then the microphones 105 could be comprised within the ear piece. If the first device 103 is a head set the microphones could be comprised within the same part of the headset, for example they could be located in the region around a first ear. Other types of first devices 103 could be used in other examples. For instance, the first devices 103 could be teleconferencing devices or any other suitable devices comprising two or more microphones 105.
The processing that is performed can comprise part of a parametric spatial analysis.
In some examples the audio signals that are received at block 501 can comprise an audio signal from a microphone 105 at the front of the first device 103 and an audio signal from a microphone 105 at the back of the first device 103. The processing that is performed can comprise a front-back analysis. The front-back analysis can determine a whether a direction of arrival is to the front of the first device 103 or to the rear of the first device 103.
The metadata that is indicated at block 501 can indicate whether the direction of arrival is to the front or to the back of the first device 103 for given time intervals. Single value such as 0 or 1 can be used to indicate if the direction of arrival is to the front or to the back.
The metadata that is generated by the first device 103 might not be sufficient to enable the spatial characteristics of the audio sources to be recreated by a playback device. However, the metadata that is generated by the first device 103 can be used to process further audio signals to obtain the metadata that is sufficient to enable the spatial characteristics of the audio sources to be recreated by a playback device.
At block 503 the method comprises encoding and transmitting the audio signals and the metadata. Any suitable protocols can be used for the encoding and the transmission of the data stream comprising the audio signals and the metadata.
In some examples the audio signals can be transmitted on their own. In some examples the audio signals can be encoded with one or more further audio signals and can be transmitted together. For instance, one of the audio signals that was used for the front back analysis could be transmitted with an audio signal obtained from a different side of the first device 103. The transmitted audio signals could therefore comprise an audio signal captured by a microphone on the left-hand side of the first device 103 and an audio signal captured by a microphone 105 on the right-hand side of the first device 103.
At block 505 the method comprises receiving and decoding the data stream at the second device 111. In some examples the second device can receive the respective signals and the associated metadata in a single data stream. In other examples the different audio signals can be received in different data streams.
At block 507 the method comprises processing the decoded audio signals to generate further metadata. The further metadata can be generated by performing parts of the parametric spatial analysis that are not performed by the first device 103.
For instance, in some examples the first device 103 analyses a signal from a microphone 105 at the front of the first device 103 and an audio signal from a microphone 105 at the back of the first device 103 to perform a front-back analysis. In such cases the second device 111 could then analyse a signal from a microphone 105 at the right hand side of the first device 103 and an audio signal from a microphone 105 at the left hand side of the first device 103 to perform a left-right analysis. The leftright analysis can determine a dominant direction of arrival for a sound source but does not determine if that angle is to the front of the first device 103 or to the back of the first device 103.
The further metadata that is generated at block 107 would therefore indicate a direction of arrival at respective time intervals. The angles could be automatically assigned to a front sector because, at this point in the processing, it is not known if the direction of arrival is to the front or to the back of the first device 103.
At block 509 the respective metadata is combined to generate complete spatial metadata. That is, by combining the metadata generated by the first device 103 and the metadata generated by the second device 111 it can be determined if the directions determined in the left-right analysis should be in a front sector or a rear sector.
The complete spatial metadata can comprise sufficient information that enables the spatial characteristics of the audio sources to be recreated by a playback device.
At block 511 the method comprises encoding and transmitting the audio signals and the complete metadata. Any suitable protocols can be used for the encoding and the transmission of the data stream comprising the audio signals and complete metadata. For example, the encoding used could comprise Immersive Voice and Audio Services (IVAS) or any other suitable protocol.
In some examples the encoded audio signals and complete metadata could be stored in the second device 111 instead of, or in addition to, being transmitted.
Fig. 6 schematically shows an example system 101 that could be used in some examples of the disclosure. In this example the system 101 comprises a first device 103, a second device 111 and a further device 601 .
In this example the first device 103 comprises a first ear piece of a headset and the further device 601 comprises a second ear piece of the headset. The first device 103 could be the left ear piece and the further device 601 could be the right ear piece. When the respective ear pieces are in use they are positioned on either side of a user’s head.
Both the first device 103 and the further device 601 comprise microphones 105. In the example of Fig. 6 the first device 105 comprises two microphones 105 and the further device 601 comprises just one microphone 105. Other numbers and arrangements of the microphones 105 could be used in examples of the disclosure.
The microphones of the first device 103 are positioned to enable spatial audio signals to be captured. Fig. 7 shows an example of an earpiece and an example arrangement of the microphones 105 that could be used in some examples.
In the example of Fig. 6 the microphones 105 of the first device 103 are positioned so that one microphone 105 is at the front of the first device 103 and one microphone 105 is at the back of the first device 103.
The first device 103 also comprises an apparatus 107. The apparatus 107 is configured to obtain audio signals from the respective microphones 105 of the first device 103 and perform part of the spatial analysis process. In the example of Fig. 6 the apparatus 107 can be configured to perform front-back analysis that determines if a dominant direction for an audio source is to the front of the first device 103 or to the rear of the first device 103.
The apparatus 107 is configured to generate metadata that indicates whether the direction of arrival is to the front or the back of the first device 103.
The first device 103 is configured to encode and transmit the metadata 603 and one or more of the audio signals 605. In this example the first device 103 is a left ear piece and so the audio signal 605 is a left audio signal 605. The metadata 603 and the left audio signal 605 can be transmitted by a wireless communication link or by any other suitable means.
In the example of Fig. 6 two audio signals are obtained by the microphones 105 in the first device 103. However, only one of the audio signals needs to be transmitted with the metadata 603. The metadata 603 with one of the audio signals provides sufficient information for the second device 111 to complete the spatial processing. This therefore reduces the data that needs to be transmitted via the communication link.
The further device 601 is configured to encode and transmit a further audio signal 607. In this example the further device 601 is a right ear piece and so the further audio signal 607 is a right audio signal 607. The right audio signal 607 can be transmitted by a wireless communication link or by any other suitable means. In some examples the right audio signal 607 could be transmitted together with the left audio signals 605. For instance, the respective earpieces could be part of the same headset. In other examples the respective ear pieces might not have any physical connection to each other and so the right audio signal 607 and the left audio signal 605 could be transmitted separately via different communication links.
The system 101 is configured so that the metadata 603, the left audio signal 605 and the right audio signal 607 are received by the second device 111. The second device 111 comprises an apparatus 115 which can be configured to process the left audio signal 605 and the right audio signal 607. In this example the apparatus 115 can use the respective audio signals 605, 607 to perform a left-right analysis.
The left-right analysis 115 that is performed by the second device 111 can determine a direction of arrival for a sound source but not if that direction of arrival is to the front or to the rear of the first device 103.
The outcome of the left-right analysis can therefore be combined with the metadata to determine complete metadata for the respective audio signals 605, 607. The second
device 111 can therefore complete the spatial processing that is started by the first device 103. The complete metadata that is generated by the second device 111 is sufficient to be used with the audio signals enable reproduction of the spatial characteristics of the audio sources by a playback device. The complete spatial metadata can be used for beam-forming, head tracking, orientation tracking, or for any other suitable spatial process.
The respective audio signals 605, 607 and the complete metadata can then be provided to an encoder 609. The encoder 609 could be an IVAS encoder or any other suitable type of encoder.
The encoder 609 is configured to encode the audio signals 605, 607 and complete metadata into a bit stream 611 that can then be transmitted to a playback device or stored in a suitable location.
Fig. 7 shows an example earpiece and an example arrangement of the microphones 105 that could be used to implement examples of the disclosure.
In this example the earpiece is a left ear piece. The earpiece could be a first device 103 as shown in Figs. 1 and 6. The earpiece comprises a first microphone 105-1 and second microphone 105-2. The first microphone 105-1 is positioned at a first end of the earpiece and is located in the user’s ear 701 when the earpiece is in use. The second microphone 105-2 is positioned at a second end of the earpiece. The second microphone 105-2 can be directed towards the user’s mouth.
In this example the spacing between the microphones 105-1 , 105-2, is sufficient to enable front-back analysis to be performed using the audio signals provided by the respective microphones 105-1 , 105-2.
In this example both of the microphones 105 that are used for the front-back analysis are located on the same side of the user’s head. This means that the front-back analysis is not affected by the presence of the user’s head and makes the analysis robust against Head Related Transfer Function (HRTF) shaping.
Other types of first devices 103 and further devices 601 could be used in examples of the disclosure, for example the first device 103 and the further device 601 could comprise teleconferencing devices or any other suitable devices comprising one or more microphones.
Fig. 8 schematically shows another example system 101 that could be used in some examples of the disclosure. This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a second device 111 and a further device 601 .
The system 101 of Fig. 8 differs from the system 101 of Fig. 6 in that in Fig. 8 both the first device 103 and the further device 601 can be configured to obtain two or more audio signals and to begin the process of the spatial analysis on those obtained signals before they are transmitted to the second device 111.
In the example of Fig. 8 the further device 601 comprises two microphones 105 that are configured to obtain audio signals that can be used for spatial audio analysis. In the example of Fig. 8 one microphone 105 is positioned at the front of the further device 601 and one microphone 105 is positioned at the back of the further device 601 .
The further device 601 also comprises an apparatus 801. The apparatus 801 can comprise a memory and processor. The apparatus 801 can be as shown in Fig. 15. The apparatus 801 can be the same as, or similar to, the apparatus 107 of the first device 103.
The apparatus 801 of the further device 601 is configured to obtain audio signals from the respective microphones 105 of the further device 601 and perform part of the spatial analysis process. In the example of Fig. 8 the apparatus 108 can be configured to perform the same, or a similar, front-back analysis as is performed in the first device 103. This can enable an indication of whether a sound source is to the front of the first device 103 or to the rear of the first device 103 to be determined for the further device 601.
The apparatus 801 of the further device 601 is configured to generate metadata 803 similar to the metadata 603 that is generated by the first device 103. The metadata
803 from the further device 601 indicates whether the direction of arrival is to the front or the back of the further device 601.
The further device 601 is configured to encode and transmit the metadata 803 and one or more of the audio signals 607 in a similar manner to, or in the same manner as, the first device 103. In this example the further device 601 is a right ear piece and so the audio signal 607 is a right audio signal 607.
Therefore, in the example of Fig. 8, both the left audio signal 605 and the right audio signal 607 are transmitted with metadata 603, 803. The metadata 603, 803 that is transmitted with the respective audio signals 605, 607 is generated by spatial processing performed at the device 103, 601 that has been used to capture the audio signal 605, 607.
The second device 111 is configured to receive both the left audio signal 605 and the associated metadata 603 and the right audio signals 607 and the associated metadata 803. The apparatus 115 of the second device 115 can use the respective audio signals 605, 607 to perform a left-right analysis and to determine complete metadata for the respective audio signals 605, 607.
The respective audio signals 605, 607 and the complete metadata can then be provided to an encoder 609 and then transmitted or stored as appropriate.
In this case the metadata 803 that is generated by the further device 601 can be a duplicate of the metadata 603 that is generated by the first device 103. In some examples the second device 111 can make a decision on which of the metadata 603, 803 to use to complete the spatial analysis. The decision on which of the metadata 603, 803 to use can be based on the reliability of the respective metadata 603, 803. In some examples the metadata 603, 803 that is most reliable can be determined by which side of the devices 103, 601 the dominant sound source is on. If it is determined that the sound source in on the left hand side of the devices 103, 601 then the analysis from the first device 103, which is the left earpiece, will be more reliable and should be used. If the audio source is on the left head side of the user’s head then the sound signals that are detected by the microphones 105 in the further device 601 will have
been attenuated by the user’s head and are more likely to be confused with other sound sources.
In some examples the second device 111 could be configured to compare the information in the metadata 603 from the first device 103 with information in the metadata 803 from the further device 601 and determine the reliability of the metadata 603, 803 based on a level of similarity. If the respective metadata 603, 803 comprises similar information then they can be considered to be more reliable. If the respective metadata 603, 803 do not comprise similar information then they can be considered to be less reliable and the second device 111 can adapt the spatial analysis process to take this into account. For example, the energy ratios can be made smaller to reduce potential artifacts.
Fig. 9 schematically shows another example system 101 that could be used in some examples of the disclosure. This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a second device 111 and a further device 601 where both the first device 103 and the further device 601 comprise at least two microphones 105 and generate metadata 603, 803 from a partial spatial analysis at the respective devices 103, 601.
The example system 101 of Fig. 9 differs from the example system of Fig. 8 in that confidence information is determined for the metadata generated by the first device 103 and the further device 601 . The confidence information can comprise a parameter that indicates how reliable a direction estimate, or any other suitable information, in the metadata 603, 803 is.
In some examples the confidence information can comprise a value. The value can be within any suitable range. In some examples the confidence value can be in the range [0,1],
Any suitable process or method can be used to determine the confidence information. For example, it can be derived from a strength of correlation between the microphone signals that have been used to generate the respective metadata 603, 803 or from any other suitable methods.
The confidence information can be encoded with the metadata 603, 803 and transmitted with the audio signals 605, 607 and the associated metadata 603, 607. The second device 111 can therefore receive the confidence information with the metadata 603, 803 and the associated audio signals 605, 607.
The second device 111 can be configured to take the confidence information into account when selecting which metadata 603, 803 to use to complete the spatial analysis. This can be useful if there is for example wind noise or scratching noises present which can affect the reliability of the spatial information obtained from the audio signals.
Fig. 10 schematically shows another example system 101 that could be used in some examples of the disclosure. This system 101 is similar to the system shown in Fig. 6 and comprises a first device 103, a further device 601 , a second device 111 and an intermediate device 1001.
In this example the intermediate device 1001 is provided between the first device 103 and the second device 111 so that signals from the first device 103 are received by the intermediate device 1001 before they are forwarded on to the second device 111. In some examples the intermediate device could be a capturing device. For example, the intermediate device 1001 could be a device comprising a camera that is being used to capture videos and corresponding audio. The second device 111 could be a playback device that is configured to process the audio signals and playback the spatial audio so that it can be heard by a user of the playback device.
In the system 101 of Fig. 10 the first device 103 and the further device 601 can be configured to obtain audio signals 605, 607 and generate metadata 603 as described above. In this example the first device 103 comprises two microphones 105 and the further device 601 comprises one microphone 105 similar to the arrangement as shown in Fig. 6. Other configurations or arrangements could be used in other examples. For instance, the example devices 130, 601 shown in Figs. 8 and 9 could be used or any suitable variations or configurations of this.
The first device 103 is configured to perform part of the spatial analysis process. In this example the apparatus 107 of the first device 103 is configured to perform front- back analysis to determine if a dominant direction of arrival for a sound source is to the front of the first device 103 or to the rear of the first device 103.
The first device 103 transmits the left audio signal 605 and the metadata 603 generated from the part of the spatial analysis process to the intermediate device 1001. The further device 601 also transmits the right audio signal 607 to the intermediate device 1001.
The intermediate device 1001 is configured to receive the metadata 603, the left audio signal 605 and the right audio signal 607. In this example the intermediate device 1001 does not perform any of the spatial processing. Instead, the intermediate device 1001 encodes the metadata 603, the left audio signal 605 and the right audio signal 607 into a bitstream 611 and enables the bitstream 611 to be transmitted to a second device 111.
The intermediate device 1001 could be a capturing device such as a mobile phone or any other suitable type of device. The intermediate device 1001 could be configured to communicate in a cellular network or via any other suitable wireless communications network.
In the example of Fig. 10 the intermediate device 1001 comprises a stereo encoder 1003 and a metadata encoder 1005. The stereo encoder 1003 is configured to encode the respective audio signals 605, 607, and the metadata encoder 1005 is configured to encode the metadata 603. The bit rate that is needed for the metadata 603 is low compared to the bit rate that would be needed for full spatial metadata such as 360 degree directional information and energy ratios.
The bitstream 611 comprising the encoded audio signals and the encoded metadata is sent from the intermediate device 1001 to the second device 111. The bitstream 611 can be transmitted using any suitable communication link. The communication link could be a wireless communication link such as cellular communication link or any other suitable type of communication link.
The second device 111 could be a playback device or any other suitable type of device. The second device 111 can be configured to receive the bitstream 611 and perform any suitable spatial processing on the respective audio signals.
In the example of Fig. 10 the second device 111 comprises a stereo decoding module 1007, a metadata decoding module 1009 and a spatial processing module 1011. The respective modules 1007, 1009, 1011 could be provided by an apparatus 115 or by any other suitable means.
The stereo decoding module 1007 is configured to decode the received audio signals and the metadata decoding module 1009 is configured to decode the received metadata. The decoded audio signals and the decoded metadata can then be used by the spatial processing module 1011 to complete the spatial processing. The spatial processing that is performed by the second device 111 can complete the spatial processing that is started by the first device 103 so as to enable reproduction of the spatial characteristics of the audio sources by a playback device.
The example of Fig. 10 can be used to avoid unnecessary processing if head tracking or spatial processing is not needed. In this example the spatial processing can be performed by the playback device rather than a capture device and so the spatial processing is only performed if it is needed.
Variations to the example of Fig. 10 could be made. For instance, in some examples the functions of the intermediate device 1001 and the functions of the second device 111 could all be performed by a single device or entity. In such cases the operations of stereo encoding 1003 and metadata encoding 1005 that are performed by the intermediate device 1001 could be used when the data is being stored. The operations of stereo decoding 1007, metadata decoding 1009, and spatial processing can be used to enable spatial processing when the signal is to be played back.
Fig. 11 schematically shows another example system 101 that could be used in some examples of the disclosure. This system 101 is similar to the system shown in Fig. 9 and comprises a first device 103, a second device 111 and a further device 601.
The system 101 of Fig. 11 differs from the system 101 of Fig. 9 in that in Fig. 11 the further device 601 is configured to obtain head tracking information 1101 and transmit the headtracking information 1101 with the right audio signals 607 and the metadata 803 and confidence information.
The second device 111 receives the head tracking information 1101 and the respective audio signals 605, 607 and the metadata 693, 803 and confidence information from first device 103 and the further device 601 . The apparatus 115 of the second device 111 can be configured to use the received information to perform the spatial analysis. The apparatus 115 can also be configured to use the head tracking information 1101 to adjust the metadata to take into account the orientation of the user’s head.
In the example of Fig. 11 the compensation for the user’s head position using the head tracking information can be made by the second device 111. This can be an audio capture device. In some examples the compensation could be made by a different device such as a playback device.
Fig. 12 schematically shows an example capturing device 1201 that could be used in some examples of the disclosure. In this example the capturing device 1201 can be a first device 103 that can be configured to perform spatial analysis process. The capturing device 1201 could be a user electronic device such as mobile phone or any other suitable type of device.
In the example of Fig. 12 the binaural signals do not have to originate from a headset or other similar device. Instead signals obtained from any suitable microphone array can be synthesized into spatial audio signals.
The example capturing device 1201 of Fig. 12 comprises a microphone array 1203, a spatial analysis module 1205, a binaural synthesis module 1207 and an encoder 1209. The spatial analysis module 1205 and the binaural synthesis module 1207 could be provided by an apparatus 107 or by any other suitable means.
The microphone array 1203 comprises a plurality of microphones 105. The microphones 105 can be positioned in different locations within or around the capturing device 1201 so as to enable spatial audio capture.
The audio signals from the microphone array 1203 are provided to the spatial analysis module 1205. The spatial analysis module 1205 can be configured to perform the spatial analysis process. Spatial analysis 1205 can be configured to produce spatial metadata 1215.
The spatial analysis module 1205 can be configured to produce complete spatial metadata, however only a subset of the produced metadata needs to be provided from the spatial analysis module 1205. For example, the spatial metadata can contain part of the analysed spatial information. For example, the spatial analysis module 1205 could be configured to generate front -back metadata. Other types of metadata could be generated in some examples. For instance, the spatial analysis module 1205 could generate left-right metadata in some examples.
The binaural synthesis module 1207 is configured to synthesize the audio signals into a binaural format to generate a binaural signal 1211. Binaural synthesis requires complete spatial metadata. The spatial metadata that is used for binaural synthesis may contain more information than spatial metadata 1215.
The encoder 1209 is configured to encode the binaural signal 1211 and the metadata 603. The encoder 1209 could be an IVAS encoder or any other suitable type of encoder.
The encoder 1209 is configured to encode the binaural signals into a bit stream 1213 that can then be transmitted to a second device 111 to enable the spatial processing to be completed.
The completion of the spatial processing can comprise any suitable steps or processes. In some examples the process can comprise using metadata indicative of the front or back direction to position sounds to loudspeakers in the front or back. If
headtracking is being used with binaural signals then analysis and resynthesis based on the head tracking information can be used.
Fig. 13 schematically shows another example system 101 that could be used in some examples of the disclosure. In this example the first device 103 that performs the first part of the spatial processing does not need to be a headset or other similar device. In this example the first device 103 could be a spatial voice conferencing server or any other suitable type of device.
In this example the first device 103 comprises a location definition module 1301 , a spatial rendering module 1303 and an encoder 1305. The location definition module 1301 and the spatial rendering module 1303 can be provided by an apparatus 107 or by any other suitable means.
In this example the first device 103 is configured to receive a plurality of audio objects. The audio objects are provided to the location definition module 1301. The location definition module 1301 is configured to define a spatial location for each of the audio objects.
The spatial rendering module 1303 is configured to perform part of a spatial analysis process on the audio objects. The spatial rendering module 1303 is configured to perform spatial signal rendering using the spatial locations defined by the location definition module 1301. The spatial rendering module 1303 can also be configured to perform front-back analysis of the audio objects and provide the metadata indicative of the output of the analysis.
The audio signals that have been rendered by the spatial rendering module 1303 and the metadata can then be provided to the encoder 1305. The encoder 1305 could be an IVAS encoder or any other suitable type of encoder.
The encoder 1305 is configured to encode the audio signals and the metadata into a bit stream 611 that can then be transmitted to a playback device or stored in a suitable location.
The bit stream 611 is transmitted to a second device 111. The second device 111 could be a playback device or any other suitable type of device. The second device 111 can be configured to receive the bitstream 611 and perform any suitable spatial processing on the respective audio signals.
In the example of Fig. 13 the second device 111 comprises a stereo decoding module 1007, a metadata decoding module 1009 and a spatial processing module 1011. The respective modules 1007, 1009, 1011 could be provided by an apparatus 115 or by any other suitable means.
The stereo decoding module 1007 is configured to decode the received audio signals and the metadata decoding module 1009 is configured to decode the received metadata. The decoded audio signals and the decoded metadata can then be used by the spatial processing module 1011 to complete the spatial processing. The spatial processing that is performed by the second device 111 can complete the spatial processing that is started by the first device 103 so as to enable reproduction of the spatial characteristics of the audio sources by a playback device.
Figs 14A to 14E show an example test set up and results that were obtained using the test set up and examples of the disclosure.
Fig. 14A shows the example test set up. In this set up an artificial head 1401 is configured with two microphones 105 attached to the left ear and one microphone attached to the right ear. Only the microphones 105 attached to the left ear are shown in the view of Fig. 14A. The microphones are spatially separated so as to enable spatial audio signals to be captured.
Using this set up, a test signal was recorded in which a sound source was originally located to the front of the artificial head 1401 and rotated around the artificial head 1401 in anticlockwise direction with about a 1.5 meter distance to the artificial head 1401. Fig. 14 B shows the three signals recorded by the respective microphones 105.
The first plot 1403 shows the signal obtained from the front microphone on the left hand side. The second plot 1405 shows the signal obtained from the back microphone
on the left hand side. The third plot 1407 shows the signal obtained from the microphone on the right hand side.
The signals from the microphones 105 on the left hand side were used for front-back analysis. The results of this analysis are shown in Fig. 14. C. For clarity the plots in Figs. 14C to 14E only show the results for the analysis at intervals of 0.5 seconds and for only one frequency domain subband.
In Fig. 14C it can be seen that between 6 and 13.5 seconds the sound source is considered to be in the back but at other times it is considered to be in front. In this case the front-back spatial analysis means that the direction of the audio source is always either 0 or 180 degrees.
The signals from the front microphone on the left hand side and the microphone from the right hand side were used for the left-right analysis. Fig. 14. D shows the outcome of the left-right analysis. In this case it is not known whether the sound source is to the front or the back and so all the sources are located in the front. That is all of the angles have a value between -90 and 90 degrees.
The final stage of the processing is performed by combining the front-back analysis and the left-right analysis. The front-back estimate is used to determine if the sound source is to the front or back. The results of the combination are shown in Fig. 14E. These clearly show that the sound source has travelled in a circle around the artificial head 1401. This spatial information can now be used do spatial processing such as audio focussing or headtracking or any other suitable type of processing. The spatial processing can be performed by a capture device or a playback device or any other suitable type of device.
Examples of the disclosure therefore enable distributed spatial processing. By performing some of the spatial processing in a first device such as a headset the amount of data that needs to be transmitted from the first device 103 to the second device 111 is reduced. For example, in example of the disclosure only two audio signals and the metadata need to be sent to the second device. This is significantly
less data than three audio signals and also means that only two channels are needed between the microphone devices 101 , 601 and the second device 111.
In examples of the disclosure only a part of the spatial processing is performed in the first device 103. This keeps the computational load required for the first device to a low level.
Fig. 15 schematically illustrates an apparatus 107/115 that can be used to implement examples of the disclosure. In this example the apparatus 107/115 comprises a controller 1501. The controller 1501 can be a chip or a chip-set. In some examples the controller 1501 can be provided within a first device 103 or a second device 105 or any other suitable type of device.
In the example of Fig. 15 the implementation of the controller 1501 can be as controller circuitry. In some examples the controller 1501 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware).
As illustrated in Fig. 15 the controller 1501 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 1507 in a general-purpose or special-purpose processor 1503 that may be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 1503.
The processor 1503 is configured to read from and write to the memory 1505. The processor 1503 can also comprise an output interface via which data and/or commands are output by the processor 1503 and an input interface via which data and/or commands are input to the processor 1503.
The memory 1505 stores a computer program 1507 comprising computer program instructions (computer program code) that controls the operation of the controller 1501 when loaded into the processor 1503. The computer program instructions, of the computer program 1507, provide the logic and routines that enables the controller 1501. to perform the methods illustrated in the accompanying Figs. The processor
1503 by reading the memory 1505 is able to load and execute the computer program 1507.
When the apparatus 107 is configured for use in a first device 103 the apparatus 107 comprises: at least one processor 1503; and at least one memory 1505 storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: obtaining 301 two or more audio signals from two or more microphones; processing 303 the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission 305 of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
When the apparatus 115 is configured for use in a second device 111 the apparatus 115 comprises: at least one processor 1503; and at least one memory 1505 storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving 401 at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving 403 at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate 405 spatial metadata.
As illustrated in Fig. 15, the computer program 1507 can arrive at the controller 1501 via any suitable delivery mechanism 1511. The delivery mechanism 1511 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory
computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid-state memory, an article of manufacture that comprises or tangibly embodies the computer program 1507. The delivery mechanism can be a signal configured to reliably transfer the computer program 1507. The controller 1501 can propagate or transmit the computer program 1507 as a computer data signal. In some examples the computer program 1507 can be transmitted to the controller 1501 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.
When the computer program 1507 is configured for use in a first device 103 the computer program 1507 comprises computer program instructions for causing an apparatus 107 to perform at least the following or for performing at least the following: obtaining 301 two or more audio signals from two or more microphones; processing 303 the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission 305 of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
When the computer program 1507 is configured for use in a second device 103 the computer program 1507 comprises computer program instructions for causing an apparatus 115 to perform at least the following or for performing at least the following: receiving 401 at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving 403 at least one further audio signal that has been captured by at least one further microphone; and
processing the at least one audio signal and the at least one further audio signal using the metadata to generate 405 spatial metadata.
The computer program instructions can be comprised in a computer program 1507, a non-transitory computer readable medium, a computer program product, a machine- readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1507.
Although the memory 1505 is illustrated as a single component/circuitry it can be implemented as one or more separate components/circuitry some or all of which can be integrated/removable and/or can provide permanent/semi-permanent/ dynamic/cached storage.
Although the processor 1503 is illustrated as a single component/circuitry it can be implemented as one or more separate components/circuitry some or all of which can be integrated/removable. The processor 1503 can be a single core or multi-core processor.
References to ‘computer-readable storage medium’, ‘computer program product’, ‘tangibly embodied computer program’ etc. or a ‘controller’, ‘computer’, ‘processor’ etc. should be understood to encompass not only computers having different architectures such as single /multi- processor architectures and sequential (Von Neumann)/parallel architectures but also specialized circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc.
As used in this application, the term ‘circuitry’ may refer to one or more or all of the following:
(a) hardware-only circuitry implementations (such as implementations in only analog and/or digital circuitry) and
(b) combinations of hardware circuits and software, such as (as applicable):
(i) a combination of analog and/or digital hardware circuit(s) with software/firmware and
(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory or memories that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and
(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (for example, firmware) for operation, but the software may not be present when it is not needed for operation.
This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
The blocks illustrated in Figs. 3 to 5 can represent steps in a method and/or sections of code in the computer program 1507. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted.
The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to “comprising only one...” or by using “consisting”.
In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected/coupled/in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., so as to provide direct or indirect
connection/coupling/communication. Any such intervening components can include hardware and/or software components.
As used herein, the term "determine/determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, " determine/determining" can include resolving, selecting, choosing, establishing, and the like.
In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’ or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all of the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example.
Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
Features described in the preceding description may be used in combinations other than the combinations explicitly described above.
Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not.
The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a/an/the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning.
The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and also to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result.
In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described.
The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such
alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure.
Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance it should be understood that the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and/or shown in the drawings whether or not emphasis has been placed thereon. l/we claim:
Claims
1 . An apparatus comprising means for: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
2. An apparatus as claimed in claim 1 , wherein the spatial metadata comprises information that enables the spatial characteristics of the audio source to be recreated by a playback device.
3. An apparatus as claimed in any preceding claim, wherein processing the obtained audio signal comprises determining whether the direction of arrival is to the front or the rear of the two or more microphones.
4. An apparatus as claimed in any preceding claim, wherein the metadata indicative of a direction of arrival for an audio source is used to process at least one further audio signal from at least one further microphone to generate the further metadata.
5. An apparatus as claimed in claim 4, wherein the further microphone that is used to obtain the further audio signal is positioned in a different location to the two or more microphones that are used to obtain the two or more audio signals.
6. An apparatus as claimed in any of claims 4 to 5, wherein the at least one of the obtained audio signals and the metadata is transmitted separately from the further audio signal.
7. An apparatus as claimed in any of claims 4 to 6, wherein the further audio signal and the transmitted at least one of the obtained audio signals are configured to be processed to determine a direction of arrival and if the direction of arrival is to the front or back of the microphones.
8. An apparatus as claimed in any of claims 4 to 7, wherein the further audio signal and the transmitted at least one of the obtained audio signals are configured to be processed using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
9. An apparatus as claimed in any preceding claim, wherein an indication of confidence in a determined direction of arrival for the audio source is transmitted.
10. An apparatus as claimed in any preceding claim, wherein the means are for determining tracking information for the two or more microphones and enabling the tracking information to be transmitted.
11. An apparatus as claimed in any preceding claim, wherein the direction of arrival is determined for one or more dominant audio sources.
12. An apparatus as claimed in any preceding claim, wherein the at least one of the obtained audio signals and the metadata are transmitted via a wireless communication link.
13. A device comprising an apparatus as claimed in any preceding claim, wherein the device comprises at least one of: a headset, ear pieces, a wearable device.
14. A method comprising: obtaining two or more audio signals from two or more microphones; processing the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enabling transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be
processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
15. An apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain two or more audio signals from two or more microphones; process the obtained two or more audio signals to obtain metadata indicative of a direction of arrival for an audio source relative to the two or more microphones; and enable transmission of at least one of the obtained audio signals with the metadata wherein the transmitted at least one audio signal is configured to be processed to obtain further metadata, wherein the metadata indicative of the direction of arrival for the audio source and the further metadata are configured to be combined to generate spatial metadata for the audio source.
16. An apparatus comprising means for: receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
17. An apparatus as claimed in claim 16, wherein generating the spatial metadata comprises obtaining further metadata by processing the at least one audio signal and the at least one further audio signal and using the metadata and the further metadata to generate spatial metadata.
18. An apparatus as claimed in any of claims 16 to 17, wherein the metadata is indicative of a direction of arrival for an audio source indicates whether the direction of
arrival is to the front or the rear of the microphones that were used to capture the at least one audio signal.
19. An apparatus as claimed in any of claims 16 to 18, wherein the spatial metadata comprises information that enables the spatial characteristics of the audio source to be recreated by a playback device.
20. An apparatus as claimed in any of claims 16 to 19, wherein the at least one audio signal and the at least one further audio signal are received separately.
21. An apparatus as claimed in any of claims 16 to 20, wherein processing the at least one audio signal and the at least one further audio signal comprises using the metadata to convert the respective audio signals from a first spatial audio format to a second, different spatial audio format.
22. An apparatus as claimed in any of claims 16 to 21 , wherein at least one of the at least one audio signal and the at least one further audio signal are received by a wireless communication link.
23. A device comprising an apparatus as claimed in any of claims 16 to 22, wherein the device comprises at least one of: a user device; a mobile phone; a processing device, a capturing device, a playback device.
24. A method comprising: receiving at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receiving at least one further audio signal that has been captured by at least one further microphone; and processing the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
25. An apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive at least one audio signal with associated metadata wherein the metadata is indicative of a direction of arrival for an audio source relative to microphones that were used to capture the at least one audio signal; receive at least one further audio signal that has been captured by at least one further microphone; and process the at least one audio signal and the at least one further audio signal using the metadata to generate spatial metadata.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GBGB2218136.6A GB202218136D0 (en) | 2022-12-02 | 2022-12-02 | Apparatus, methods and computer programs for spatial audio processing |
| PCT/EP2023/081119 WO2024115062A1 (en) | 2022-12-02 | 2023-11-08 | Apparatus, methods and computer programs for spatial audio processing |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4627806A1 true EP4627806A1 (en) | 2025-10-08 |
Family
ID=84926537
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23809114.4A Pending EP4627806A1 (en) | 2022-12-02 | 2023-11-08 | Apparatus, methods and computer programs for spatial audio processing |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4627806A1 (en) |
| CN (1) | CN120380777A (en) |
| GB (1) | GB202218136D0 (en) |
| WO (1) | WO2024115062A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2549532A (en) * | 2016-04-22 | 2017-10-25 | Nokia Technologies Oy | Merging audio signals with spatial metadata |
| GB2556093A (en) * | 2016-11-18 | 2018-05-23 | Nokia Technologies Oy | Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices |
| GB2559765A (en) * | 2017-02-17 | 2018-08-22 | Nokia Technologies Oy | Two stage audio focus for spatial audio processing |
-
2022
- 2022-12-02 GB GBGB2218136.6A patent/GB202218136D0/en not_active Ceased
-
2023
- 2023-11-08 EP EP23809114.4A patent/EP4627806A1/en active Pending
- 2023-11-08 WO PCT/EP2023/081119 patent/WO2024115062A1/en not_active Ceased
- 2023-11-08 CN CN202380082691.1A patent/CN120380777A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN120380777A (en) | 2025-07-25 |
| GB202218136D0 (en) | 2023-01-18 |
| WO2024115062A1 (en) | 2024-06-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11950063B2 (en) | Apparatus, method and computer program for audio signal processing | |
| EP3141001B1 (en) | System, apparatus and method for consistent acoustic scene reproduction based on adaptive functions | |
| JP6703525B2 (en) | Method and device for enhancing sound source | |
| US11632643B2 (en) | Recording and rendering audio signals | |
| CN114424588B (en) | Directional estimation enhancement for parametric spatial audio capture using wideband estimation | |
| CN113597776B (en) | Wind noise reduction in parametric audio | |
| US12439220B2 (en) | Apparatus, methods and computer programs for enabling reproduction of spatial audio signals | |
| US12587781B2 (en) | Parametric spatial audio rendering with near-field effect | |
| CN102969003A (en) | Camera sound extraction method and device | |
| WO2022263710A1 (en) | Apparatus, methods and computer programs for obtaining spatial metadata | |
| EP4627806A1 (en) | Apparatus, methods and computer programs for spatial audio processing | |
| CN114097029A (en) | Packet loss concealment for DirAC-based spatial audio coding | |
| TW202448192A (en) | Apparatus and method for binaural pose correction | |
| US12532144B2 (en) | Apparatus, methods and computer programs for processing audio signals | |
| US20240048902A1 (en) | Pair Direction Selection Based on Dominant Audio Direction | |
| CN115942168B (en) | Spatial audio capture | |
| US20250203310A1 (en) | Spatial Audio Processing | |
| CN117711428A (en) | Apparatus, method and computer program for spatially processing an audio scene |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250702 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |