EP2766904A1 - An audio scene mapping apparatus - Google Patents
An audio scene mapping apparatusInfo
- Publication number
- EP2766904A1 EP2766904A1 EP11873915.0A EP11873915A EP2766904A1 EP 2766904 A1 EP2766904 A1 EP 2766904A1 EP 11873915 A EP11873915 A EP 11873915A EP 2766904 A1 EP2766904 A1 EP 2766904A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio
- audio signal
- frequency range
- signal
- ieast
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R5/00—Stereophonic arrangements
- H04R5/027—Spatial or constructional arrangements of microphones, e.g. in dummy heads
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/02—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/173—Transcoding, i.e. converting between two coded representations avoiding cascaded coding-decoding
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B20/00—Signal processing not specific to the method of recording or reproducing; Circuits therefor
- G11B20/10—Digital recording or reproducing
Definitions
- the present application relates to apparatus for the processing of audio and additionally audio-visual signals.
- the invention further relates to, but is not limited to, apparatus for processing audio and additionally audio-visual signals from mobile devices.
- Multiple 'feeds' may be found in sharing services for video and audio signals (such as those employed by YouTube).
- Such systems which are known and are widely used to share user generated content recorded and uploaded or up- streamed to a server and then downloaded or down-streamed to a viewing/listening user.
- Such systems rely on users recording and uploading or up- streaming a recording of an event using the recording facilities at hand to the user. This may typically be in the form of the camera and microphone arrangement of a mobile device such as a mobile phone.
- the viewing/listening end user may then select one of the up-streamed or uploaded data to view or listen.
- an apparatus comprising at least one processor and at least one memory including computer code for one or more programs, the at least one memory and the computer code configured to with the at least one processor cause the apparatus to at least perform: receiving at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range; scaling the at least two audio signals; and combining the at least two audio signals to generate a combined audio signal representation of the audio source.
- Receiving the at least two signals from at least two recording apparatus may further cause the apparatus to perform demultiplexing from each signal at least one of the audio signals.
- Scaling each audio signal dependent on the at least one parameter associated with each audio signal may cause the apparatus to perform scaling the at least one audio signal dependent on the estimation of an energy of the audio signal; and an estimation of the ratio of the energy of the audio signal frequency range with respect to a full band energy value.
- Each audio signal may be a frequency domain representation audio signal comprising at least one frequency sample value, wherein combining the at least two audio signals to generate a combined audio signal representation of the audio source may cause the apparatus to append the at least two audio signals.
- the apparatus may be further caused to demultiplex from each encapsulated signal a frequency range indicator associated with each audio signal, wherein appending the at least two audio signals may further cause the apparatus to append the at least two audio signals dependent on the frequency range indicator.
- the apparatus may be further caused to perform a frequency to time domain transform on the at least two appended audio signals to generate a time domain representation of the audio source.
- the at least two recording apparatus may comprise recording apparatus within a first device.
- the apparatus may be further caused to perform decoding from each signal at least one audio signal.
- apparatus comprising at least one processor and at least one memory including computer code for one or more programs, the at least one memory and the computer code configured to with the at least one processor cause the apparatus to at least perform: capturing at least one audio signal representing at least one audio source within an audio scene; selecting a first frequency range from the audio signal; outputting the first frequency range from the audio signal to a further apparatus.
- the apparatus may be further caused to perform: estimating the energy of the at least one audio signal; estimating the energy of the first frequency range from the audio signal; outputting with the first frequency range from the audio signal an energy indicator comprising at least one of: the energy of the at least one audio signal; the energy of the first frequency range from the audio signal; the ratio of the energy of the first frequency range from the audio signal to the energy of the at least one audio signal.
- the apparatus may be further caused to perform: outputting with the first frequency range from the audio signal an indicator of the first frequency range.
- Selecting a first frequency range from the audio signal may cause the apparatus to further perform: time to frequency domain transforming the audio signal representing at least one audio source within an audio scene; and selecting at least one frequency domain representation from the audio signal in the frequency domain, the at least one frequency domain representation being associated with the first frequency range.
- Outputting the first frequency range from the audio signal to a further apparatus may further cause the apparatus to perform encapsulating the audio signal in a encapsulated signal format.
- a method comprising: receiving at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range; scaling the at least two audio signals; and combining the at least two audio signals to generate a combined audio signal representation of the audio source.
- Receiving the at least two signals from at least two recording apparatus may further comprise demultiplexing from each signal at least one of the audio signals.
- the method may further comprise demultiplexing from each signal at least one parameter associated with each audio signal.
- Scaling the at least two audio signals may further comprise scaling each audio signal dependent on the at least one parameter associated with each audio signal.
- Scaling each audio signal dependent on the at least one parameter associated with each audio signal may comprise scaling the at least one audio signal dependent on the estimation of an energy of the audio signal; and an estimation of the ratio of the energy of the audio signal frequency range with respect to a full band energy value.
- Each audio signal may be a frequency domain representation audio signal comprising at least one frequency sample value, wherein combining the at least two audio signals to generate a combined audio signal representation of the audio source may comprise appending the at least two audio signals.
- the method may further comprise demultiplexing from each encapsulated signal a frequency range indicator associated with each audio signal, wherein appending the at least two audio signals may further comprise appending the at least two audio signals dependent on the frequency range indicator.
- the method may further comprise frequency to time domain transforming the at least two appended audio signals to generate a time domain representation of the audio source.
- the at least two recording apparatus may comprise recording apparatus within a first device.
- the method may further comprise decoding from each signal at least one audio signal.
- a method comprising: capturing at least one audio signal representing at least one audio source within an audio scene; selecting a first frequency range from the audio signal; and outputting the first frequency range from the audio signal to a further apparatus.
- the method may further comprise: estimating the energy of the at least one audio signal; estimating the energy of the first frequency range from the audio signal; outputting with the first frequency range from the audio signal an energy indicator including at least one of: the energy of the at least one audio signal; the energy of the first frequency range from the audio signal; the ratio of the energy of the first frequency range from the audio signal to the energy of the at least one audio signal.
- the method may further comprise outputting with the first frequency range from the audio signal an indicator of the first frequency range.
- Selecting a first frequency range from the audio signal may further comprise: time to frequency domain transforming the audio signal representing at least one audio source within an audio scene; and selecting at least one frequency domain representation from the audio signal in the frequency domain, the at least one frequency domain representation being associated with the first frequency range.
- Outputting the first frequency range from the audio signal to a further apparatus may further comprise encapsulating the audio signal in an encapsulated signal format.
- apparatus comprising: means for receiving at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range; means for scaling the at least two audio signals; and means for combining the at least two audio signals to generate a combined audio signal representation of the audio source.
- the means for receiving the at least two signals from at least two recording apparatus may further comprise means for demultiplexing from each signal at least one of the audio signals.
- the apparatus may further comprise means for demultiplexing from each signal at least one parameter associated with each audio signal.
- the means for scaling the at least two audio signals may further comprise means for scaling each audio signal dependent on the at least one parameter associated with each audio signal.
- the means for scaling each audio signal dependent on the at least one parameter associated with each audio signal may comprise means for scaling the at least one audio signal dependent on: the estimation of an energy of the audio signal; and an estimation of the ratio of the energy of the audio signal frequency range with respect to a full band energy value.
- Each audio signal may be a frequency domain representation audio signal comprising at least one frequency sample value, wherein the means for combining the at least two audio signals to generate a combined audio signal representation of the audio source may comprise means for appending the at least two audio signals.
- the apparatus may further comprise means for demultiplexing from each encapsulated signal a frequency range indicator associated with each audio signal, wherein the means for appending the at least two audio signals may further comprise means for appending the at least two audio signals dependent on the frequency range indicator.
- the apparatus may further comprise means for frequency to time domain transforming the at least two appended audio signals to generate a time domain representation of the audio source.
- the at least two recording apparatus may comprise recording apparatus within a first device.
- the apparatus may further comprise means for decoding from each signal at least one audio signal.
- apparatus comprising: means for capturing at least one audio signal representing at least one audio source within an audio scene; means for selecting a first frequency range from the audio signal; and means for outputting the first frequency range from the audio signal to a further apparatus.
- the apparatus may further comprise: means for estimating the energy of the at least one audio signal; means for estimating the energy of the first frequency range from the audio signal; means for outputting with the first frequency range from the audio signal an energy indicator including at least one of: the energy of the at least one audio signal; the energy of the first frequency range from the audio signal; the ratio of the energy of the first frequency range from the audio signal to the energy of the at least one audio signal.
- the apparatus may further comprise means for outputting with the first frequency range from the audio signal an indicator of the first frequency range.
- the means for selecting a first frequency range from the audio signal may comprise: means for time to frequency domain transforming the audio signal representing at least one audio source within an audio scene; and means for selecting at least one frequency domain representation from the audio signal in the frequency domain, the at least one frequency domain representation being associated with the first frequency range.
- the means for outputting the first frequency range from the audio signal to a further apparatus may further comprise means for encapsulating the audio signal in an encapsulated signal format.
- an apparatus comprising: an input configured to receive at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range; a scaler configured to scale the at least two audio signals; and a combiner configured to combine the at least two audio signals to generate a combined audio signal representation of the audio source.
- the input configured to receive the at least two signals from at least two recording apparatus may further comprise a demultplexer configured to demultiplex from each signal at least one of the audio signals.
- the demultiplexer may further demultiplex from each signal at least one parameter associated with each audio signal.
- the scaler may be further configured to scale each audio signal dependent on the at least one parameter associated with each audio signal.
- the scaler scaling each audio signal dependent on the at least one parameter associated with each audio signal may further be configured to scale the at least one audio signal dependent on the estimation of an energy of the audio signal; and an estimation of the ratio of the energy of the audio signal frequency range with respect to a full band energy value.
- Each audio signal may be a frequency domain representation audio signal comprising at least one frequency sample value, wherein the combiner may comprise a sample combiner configured to append samples for the at least two audio signals.
- the demultiplexer may further be configured to demultiplex from each encapsulated signal a frequency range indicator associated with each audio signal, wherein the sample combiner may be configured to append the at least two audio signals dependent on the frequency range indicator.
- the apparatus may comprise a frequency to time domain transformer configured to generate a time domain representation of the audio source.
- the at least two recording apparatus may comprise recording apparatus within a first device.
- the apparatus may further comprise a decoder configured to decode each at least one audio signal.
- apparatus comprising: a microphone configured to capture at least one audio signal representing at least one audio source within an audio scene; a selector configured to select a first frequency range from the audio signal; and a multiplexer configured to output the first frequency range from the audio signal to a further apparatus.
- the apparatus may further comprise an energy estimator configured to estimate the energy of the at least one audio signal and estimate the energy of the first frequency range from the audio signal; the multiplexer further configured to output with the first frequency range from the audio signal an energy indicator including at least one of: the energy of the at least one audio signal; the energy of the first frequency range from the audio signal; the ratio of the energy of the first frequency range from the audio signal to the energy of the at least one audio signal.
- an energy estimator configured to estimate the energy of the at least one audio signal and estimate the energy of the first frequency range from the audio signal
- the multiplexer further configured to output with the first frequency range from the audio signal an energy indicator including at least one of: the energy of the at least one audio signal; the energy of the first frequency range from the audio signal; the ratio of the energy of the first frequency range from the audio signal to the energy of the at least one audio signal.
- the multiplexer may further be configured to output with the first frequency range from the audio signal an indicator of the first frequency range.
- the apparatus may further comprise a time to frequency domain transformer configured to time to frequency domain transform the audio signal representing at least one audio source within an audio scene; and the selector may be configured to select at least one frequency domain representation from the audio signal in the frequency domain, the at least one frequency domain representation being associated with the first frequency range.
- the multiplexer may be configured to encapsulate the audio signal in a encapsulated signal format.
- a computer program product stored on a medium may cause an apparatus to perform the method as described herein.
- An electronic device may comprise apparatus as described herein.
- a chipset may comprise apparatus as described herein.
- Embodiments of the present application aim to address problems associated with the state of the art.
- Figure 1 shows schematically a multi-user free-viewpoint service sharing system which may encompass embodiments of the application
- FIG. 2 shows schematically an apparatus suitable for being employed in embodiments of the application
- Figure 3 shows schematically an audio scene audio spectrum represented recording apparatus according to some embodiments of the application
- Figure 4 shows schematically an example audio scene representation map
- Figure 5 shows schematically the recording apparatus according to some embodiments of the application
- Figure 6 shows a flow diagram of the operation of the recording apparatus according to some embodiments of the application
- Figure 7 shows schematically the audio scene server according to some embodiments of the application.
- Figure 8 shows a flow diagram of the operation of the audio scene server according to some embodiments.
- the concept of this application is related to assisting in the production of immersive person to person communication and can include video (and in some embodiments synthetic or computer generated content).
- the recording apparatus within an event space or event scene can be configured to record or capture the audio data occurring within the audio visual event scene being monitored or listened to.
- the event space is typically of a limited physical size where all of the recording apparatus are recording substantially the same audio source.
- the same audio signal is being recorded from different positions, locations or orientations significantly close enough to each other that there would be no significant quality improvement between the audio capture apparatus. For example a concert being recorded by multiple apparatus located in the same area of the concert hall. In most cases the additional recording or capture apparatus would not bring significant additional benefit to the end user experience of the audio scene.
- each of which is processing, encoding and uploading audio signals to an audio event or audio scene server creates significant processing and bandwidth load across the recording apparatus and between the recording apparatus and the audio scene server.
- a significant proportion of the load could be considered to be therefore redundant.
- the event scene Figure 1 comprises a first capture apparatus A 303, a second capture apparatus B 305, a third capture apparatus C 307, a fourth capture apparatus D 309 and a fifth capture apparatus E 311 .
- FIG. 3 an example audio spectrum of the audio source denoting the event scene 301 is shown.
- the source audio spectrum 201 is shown comprising a spectrum from 0 to FHz.
- each of the five capture apparatus A to E capture audio spectra is shown in Figure 3.
- Figure 3 shows a first audio spectrum associated with the recording apparatus A, a second audio spectrum associated with the second recording apparatus B, a third audio spectrum associated with the third recording apparatus C, a fourth audio spectrum 309 associated with the fourth recording apparatus D, and a fifth audio spectrum 311 associated with the fifth recording apparatus E.
- Each spectrum, in such an example, where the capture apparatus are physically located near to each other would be substantially the same as each other.
- audio scene recording or capturing for multiple users comprises for each of the capture apparatus recording or capturing a frequency range of the audio scene, processing a selected or determined portion of the audio scene spectrum from the captured spectrum frequency, and then at the server rendering the frequency regions from multiple apparatus to obtain a full spectrum signal for the captured audio scene.
- the coding part can be computationally very efficient and implement simple coding audio algorithms whilst the signal quality at the rendering side can be maximised whilst the total bitrate required for the audio scene reduced.
- the audio scene server or receiver can be located in the network.
- a recording apparatus or device can act as an audio scene server and be located either within the scene or outside of the physical location of audio scene.
- the audio space 1 can have located within it at least one recording or capturing device or apparatus 19 which are arbitrarily positioned within the audio space to record suitable audio scenes.
- the apparatus 19 shown in Figure 1 are represented as microphones with a polar gain pattern 101 showing the directional audio capture gain associated with each apparatus.
- the apparatus 19 in Figure 1 are shown such that some of the apparatus are capable of attempting to capture the audio scene or activity 103 within the audio space.
- the activity 103 can be any event the user of the apparatus wishes to capture. For example the event could be a music event or audio of a news worthy event.
- the apparatus 19 although being shown having a directional microphone gain pattern 101 would be appreciated that in some embodiments the microphone or microphone array of the recording apparatus 19 has a omnidirectional gain or different gain profile to that shown in Figure 1 .
- Each recording apparatus 19 can in some embodiments transmit or alternatively store for later consumption the captured audio signals via a transmission channel 107 to an audio scene server 109.
- the recording apparatus 19 in some embodiments can encode the audio signal to compress the audio signal in a known way in order to reduce the bandwidth required in "uploading" the audio signal to the audio scene server 109.
- the recording apparatus 19 in some embodiments can be configured to estimate and upload via the transmission channel 107 to the audio scene server 109 an estimation of the location and/or the orientation or direction of the apparatus.
- the position information can be obtained, for example, using GPS coordinates, cell-ID or a-GPS or any other suitable location estimation methods and the orientation/direction can be obtained, for example using a digital compass, accelerometer, or gyroscope information.
- the recording apparatus 19 can be configured to capture or record one or more audio signals for example the apparatus in some embodiments have multiple microphones each configured to capture the audio signal from different directions. In such embodiments the recording device or apparatus 19 can record and provide more than one signal from different the direction/orientations and further supply position/direction information for each signal.
- an audio or sound source can be defined as each of the captured or audio recorded signal.
- each audio source can be defined as having a position or location which can be an absolute or relative value.
- the audio source can be defined as having a position relative to a desired listening location or position.
- the audio source can be defined as having an orientation, for example where the audio source is a beamformed processed combination of multiple microphones in the recording apparatus, or a directional microphone.
- the orientation may have both a directionality and a range, for example defining the 3dB gain range of a directional microphone.
- step 1001 The capturing and encoding of the audio signal and the estimation of the position/direction of the apparatus is shown in Figure 1 by step 1001 .
- the audio scene server 109 furthermore can in some embodiments communicate via a further transmission channel 1 1 1 to a listening device 1 13.
- the listening device 1 13 which is represented in Figure 1 by a set of headphones, can prior to or during downloading via the further transmission channel 1 1 1 select a listening point, in other words select a position such as indicated in Figure 1 by the selected listening point 105.
- the listening device 1 13 can communicate via the further transmission channel 1 1 1 to the audio scene server 109 the request.
- the selection of a listening position by the listening device 1 13 is shown in Figure 1 by step 1005.
- the audio scene server 109 can as discussed above in some embodiments receive from each of the recording apparatus 19 an approximation or estimation of the location and/or direction of the recording apparatus 19.
- the audio scene server 109 can in some embodiments from the various captured audio signals from recording apparatus 19 produce a composite audio signal representing the desired listening position and the composite audio signal can be passed via the further transmission channel 1 1 1 to the listening device 1 13.
- the generation or supply of a suitable audio signal based on the selected listening position indicator is shown in Figure 1 by step 1007.
- the listening device 1 13 can request a multiple channel audio signal or a mono-channel audio signal. This request can in some embodiments be received by the audio scene server 109 which can generate the requested multiple channel data.
- the audio scene server 109 in some embodiments can receive each uploaded audio signal and can keep track of the positions and the associated direction/orientation associated with each audio source, In some embodiments the audio scene server 109 can provide a high level coordinate system which corresponds to locations where the uploaded/upstreamed content source is available to the listening device 1 13. The "high level" coordinates can be provided for example as a map to the listening device 1 13 for selection of the listening position.
- the listening device end user or an application used by the end user
- the audio scene server 109 can in some embodiments receive the selection/determination and transmit the downmixed signal corresponding to the specified location to the listening device.
- the listening device/end user can be configured to select or determine other aspects of the desired audio signal, for example signal quality, number of channels of audio desired, etc.
- the audio scene server 109 can provide in some embodiments a selected set of downmixed signals which correspond to listening points neighbouring the desired location/direction and the listening device 1 13 selects the audio signal desired.
- Figure 2 shows a schematic block diagram of an exemplary apparatus or electronic device 10, which may be used to record (or operate as a recording device 19) or listen (or operate as a listening device 1 13) to the audio signals (and similarly to record or view the audio-visual images and data). Furthermore in some embodiments the apparatus or electronic device can function as the audio scene server 109.
- the electronic device 10 may for example be a mobile terminal or user equipment of a wireless communication system when functioning as the recording device or listening device 1 13.
- the apparatus can be an audio player or audio recorder, such as an MP3 player, a media recorder/player (also known as an MP4 player), or any suitable portable device suitable for recording audio or audio/video camcorder/memory audio or video recorder.
- the apparatus 10 can in some embodiments comprise an audio subsystem.
- the audio subsystem for example can comprise in some embodiments a microphone or array of microphones 1 1 for audio signal capture.
- the microphone or array of microphones can be a solid state microphone, in other words capable of capturing audio signals and outputting a suitable digital format signal.
- the microphone or array of microphones 1 1 can comprise any suitable microphone or audio capture means, for example a condenser microphone, capacitor microphone, electrostatic microphone, Electret condenser microphone, dynamic microphone, ribbon microphone, carbon microphone, piezoelectric microphone, or microelectrical-mechanical system (MEMS) microphone.
- MEMS microelectrical-mechanical system
- the microphone 1 1 or array of microphones can in some embodiments output the audio captured signal to an analogue-to-digital converter (ADC) 14.
- the apparatus can further comprise an analogue-to-digital converter (ADC) 14 configured to receive the analogue captured audio signal from the microphones and outputting the audio captured signal in a suitable digital form.
- ADC analogue-to-digital converter
- the analogue-to-digital converter 14 can be any suitable analogue-to- digital conversion or processing means.
- the apparatus 10 audio subsystem further comprises a digital-to-analogue converter 32 for converting digital audio signals from a processor 21 to a suitable analogue format.
- the digital-to-analogue converter (DAC) or signal processing means 32 can in some embodiments be any suitable DAC technology.
- the audio subsystem can comprise in some embodiments a speaker 33.
- the speaker 33 can in some embodiments receive the output from the digital- to-analogue converter 32 and present the analogue audio signal to the user.
- the speaker 33 can be representative of a headset, for example a set of headphones, or cordless headphones.
- the apparatus 10 is shown having both audio capture and audio presentation components, it would be understood that in some embodiments the apparatus 10 can comprise one or the other of the audio capture and audio presentation parts of the audio subsystem such that in some embodiments of the apparatus the microphone (for audio capture) or the speaker (for audio presentation) are present.
- the apparatus 10 comprises a processor 21 .
- the processor 21 is coupled to the audio subsystem and specifically in some examples the analogue-to-digital converter 14 for receiving digital signals representing audio signals from the microphone 1 1 , and the digital-to-analogue converter (DAC) 12 configured to output processed digital audio signals.
- the processor 21 can be configured to execute various program codes.
- the implemented program codes can comprise for example audio classification and audio scene mapping code routines.
- the program codes can be configured to perform audio scene event detection and device selection indicator generation, wherein the audio scene server 109 can be configured to determine events from multiple received audio recordings to assist the user in selecting an audio recording which is meaningful and does not require the listener to carry out undue searching of all of the audio recordings.
- the apparatus further comprises a memory 22.
- the processor is coupled to memory 22.
- the memory can be any suitable storage means.
- the memory 22 comprises a program code section 23 for storing program codes implementable upon the processor 21 .
- the memory 22 can further comprise a stored data section 24 for storing data, for example data that has been encoded in accordance with the application or data to be encoded via the application embodiments as described later.
- the implemented program code stored within the program code section 23, and the data stored within the stored data section 24 can be retrieved by the processor 21 whenever needed via the memory-processor coupling.
- the apparatus 10 can comprise a user interface 15.
- the user interface 15 can be coupled in some embodiments to the processor 21 .
- the processor can control the operation of the user interface and receive inputs from the user interface 15.
- the user interface 15 can enable a user to input commands to the electronic device or apparatus 10, for example via a keypad, and/or to obtain information from the apparatus 10, for example via a display which is part of the user interface 15.
- the user interface 15 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the apparatus 10 and further displaying information to the user of the apparatus 10.
- the apparatus further comprises a transceiver 13, the transceiver in such embodiments can be coupled to the processor and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network.
- the transceiver 13 or any suitable transceiver or transmitter and/or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
- the coupling can, as shown in Figure 1 , be the transmission channel 107 (where the apparatus is functioning as the recording device 19 or audio scene server 109) or further transmission channel 1 1 1 (where the device is functioning as the listening device 1 13 or audio scene server 109).
- the transceiver 13 can communicate with further devices by any suitable known communications protocol, for example in some embodiments the transceiver 13 or transceiver means can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).
- UMTS universal mobile telecommunications system
- WLAN wireless local area network
- IRDA infrared data communication pathway
- the apparatus comprises a position sensor 16 configured to estimate the position of the apparatus 10.
- the position sensor 16 can in some embodiments be a satellite positioning sensor such as a GPS (Global Positioning System), GLONASS or Galileo receiver.
- GPS Global Positioning System
- GLONASS Galileo receiver
- the positioning sensor can be a cellular ID system or an assisted GPS system.
- the apparatus 10 further comprises a direction or orientation sensor.
- the orientation/direction sensor can in some embodiments be an electronic compass, accelerometer, a gyroscope or be determined by the motion of the apparatus using the positioning estimate. It is to be understood again that the structure of the electronic device 10 could be supplemented and varied in many ways.
- the above apparatus 10 in some embodiments can be operated as an audio scene server 109.
- the audio scene server 109 can comprise a processor, memory and transceiver combination.
- the capture apparatus can in some embodiments comprise a microphone or some suitable means for capturing at least one audio signal representing at least one audio source within an audio scene as described herein.
- the microphone (which may be part of an audio sub-system 1 1 14) can be configured to capture the audio signal either in the digital domain or to capture the audio signal in the analogue domain and convert the analogue domain audio signal into the digital domain.
- step 501 the operation of capturing the audio signal is shown by step 501 .
- the recording apparatus 19 and as shown in Figure 3 and 4 by apparatus 303, 305, 307, 309, and 31 1 can in some embodiments comprise a transformer 401 or suitable means for time to frequency domain transforming the audio signal.
- the transformer 401 can be configured to apply a transform to the captured digitised time domain microphone signal (x).
- bin is the frequency bin index
- L is the time frame index
- T is the hop size between successive segments
- TF() is the time to frequency operator.
- the time to frequency operator can be any suitable time to frequency transformation operator.
- the time to frequency operator can be a modified discrete cosine transform (MDCT), modified discrete sine transform (MDST), a Fast Fourier transform (FFT) a discrete cosine transform (DCT) quadrature mirror filter (QMF) and complex valued QMF.
- MDCT modified discrete cosine transform
- MDST modified discrete sine transform
- FFT Fast Fourier transform
- DCT discrete cosine transform
- QMF quadrature mirror filter
- complex valued QMF complex valued QMF.
- x jn (n) w(n) ⁇ x(n + 1 ⁇ ⁇ )
- win(n) is a /V-point analysis window, such as sinusoidal, Hanning, Hamming, Welch, Bartlett, Kaiser or Kaiser-Bessel Derived (KBD) window.
- the application of the time to frequency domain transformation can in some embodiments be employed on a frame by frame basis where the size of the frame is of a short duration. In some embodiments the frame can be 20ms and typically less than 50ms.
- the transformer 401 can in some embodiments output the frequency domain samples to an energy determiner 403.
- the recording apparatus 19 can comprise an energy determiner 403, or suitable means for estimating the energy of the at least one audio signal configured to receive the frequency domain samples and generate an energy value determination.
- the energy determiner 403 can in some embodiments be configured to determine the total energy of the signal segment being processed. This can mathematically be represented as:
- the energy determiner 403 can be configured to determine the ratio of the energy for the selected frequency region with respect to the overall energy. This ratio can be represented mathematically as 7 ( freqEnd- ⁇
- ⁇ k freqSlart where FreqStart and FreqEnd describe the start and end bin indices of the selected frequency region respectively.
- the operation of determining the energy of the signal is shown with regards to Figure 6 by step 505.
- the boundaries of the frequency region with respect to each of the recording apparatus can be in some embodiments freely selected or in some further embodiments can follow some predefined boundaries.
- the frequency range boundaries used for dividing the spectrum between the capture apparatus can be matched to utilise human auditory modelling parameters. As the human auditory system operates on a pseudo logarithmic scale this can be represented as non-uniform frequency bands used as they more closely reflect the auditory sensitivity.
- the non-uniform bands can follow the boundaries of the equivalent rectangular bandwidth (ERB) which is a measure used in psycho- acoustics giving an approximation to the bandwidth of the filters in human hearing.
- ERB equivalent rectangular bandwidth
- the non-uniform bands can follow the boundaries of the bark bands.
- freqEnd fiOffset[fldx + ⁇ ] where fldx describes the frequency band index for the selected frequency region.
- the frequency band index can in some embodiments be received by the recording device or apparatus from the network, randomly selected or be selected according to any suitable method either internally (with respect to the recording apparatus) or system wide and instructed. The exact method for determining the selected frequency region however will not be discussed with regards to this application in order to simplify the understanding of this application.
- FIG. 3 an example spectrum division or region selection for the event scene 301 recording apparatus A 303, B 305, C 307, D 309 and E 31 1 and the frequency band index for the selected frequency region are shown.
- apparatus A is assigned the frequency region 203 OHz to fi Hz 202
- recording apparatus B 305 is assigned the frequency region 205 from fi 202 to f 2 204
- recording apparatus C 307 is assigned the frequency region 207 from f 2 204 to f 3 206
- recording apparatus D 309 is assigned the frequency region 209 from f 3 206 to f 4 208
- recording apparatus E 31 1 is assigned the frequency region 21 1 from f 4 208 to F 210.
- the determined energy indices, and furthermore in some embodiments the frequency selection information, can be passed to the sample selector 405.
- the recording apparatus comprises a sample selector 405 or suitable means for selecting a first frequency range from the audio signal for example means for example in the frequency domain representation means for selecting at least one frequency domain representation from the audio signal in the frequency domain.
- the sample selector is configured to filter or select frequency samples which lie within the determined frequency selection range.
- the sample selector 405 can be mathematically represented as follows:
- the recording apparatus 19 can further comprise a multiplexer 407 or suitable means for encapsulating the audio signal in an encapsulated signal format.
- the multiplexer 407 is configured to receive the selected samples, together with various parameters, and multiplex or encode the values into a suitable encoded form to be transferred to the audio scene server 109.
- the captured signal components and related subcomponents are multiplexed to create an encapsulation format for transmission and storage.
- the encapsulation format can for example comprise the following elements:
- step 509 The generation or multiplexing the samples and other parameters into an encapsulated file format for transmission to the audio scene server 109 is shown with respect to Figure 6 by step 509.
- the multiplexer can then output the multiplexed signal to the audio scene server 109 by a suitable means for outputting the first frequency range from the audio signal to a further apparatus.
- the output of the multiplexed signal to the audio scene server can be shown with regards to Figure 6 in step 51 1. It would be understood that the operations of capturing, transforming, energy determination, sample selection, multiplexing and outputting can be performed for each time frame index and for each of the recording apparatus. With respect to Figures 7 and 8 the audio scene server 109 and the operation of the audio scene server according to some embodiments of the application are further described.
- the audio scene server or apparatus can comprise means for receiving at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range.
- the audio scene server 109 can in some embodiments comprise a demultiplexer block 601.
- the demultiplexer block 601 can be configured to receive from each of the recording apparatus #1 to #N the multiplexed audio samples.
- the operation of receiving the multiplexed audio samples from the recording apparatus is shown in Figure 8 by step 701.
- the demultiplexer block 601 or suitable means for demultiplexing from each signal at least one of the audio signals and furthermore in some embodiments means for demultiplexing from each signal at least one parameter associated with each audio signal can further be considered to comprise multiple demultiplexer blocks each associated with a recording apparatus.
- demultiplexer, block 6011 is associated with the multiplexed audio samples from recording device or apparatus #1 and demultiplexer block 601 n is associated with the recording apparatus #N.
- Each of the demultiplexer blocks 601 n can be in some embodiments considered to be operating in parallel with each other either partially, completely or be considered to be time division processes running on the same demultiplexer or demultiplexer processor.
- the demultiplexer block 601 n can be configured to extract the captured signal components and related subcomponents from the received encapsulated format.
- the following elements can for example be retrieved from the multiplexed audio signal such as: E x , the energy of the captured signal segment; eRatio, the energy ratio of the selected frequency region with respect to the full band energy; fldx, the selected frequency region; and Y d , the frequency samples associated with the selected frequency region of the apparatus.
- the demultiplexer 601 is configured to reverse the operation of the multiplexer found within each of the recording apparatus.
- the demultiplexer 601 can be configured to output the reference identification values, and the samples to a sample scaler 603.
- the audio scene server 109 further comprises a sample scaler 603 or suitable means for scaling the at least two audio signals.
- the means for scaling the at least two audio signals comprises means for scaling each audio signal dependent on the at least one parameter associated with each audio signal.
- the sample scaler 603 can be configured in some embodiments to decode the received frequency samples to the full band frequency spectrum.
- freqEnd + 1]
- Y,d is initially a zero valued vector or a null vector and the value id describes the corresponding device index of the received data for the rendering side.
- the corresponding device indices can be 0, 1 , and 2.
- sample scaler 603 can be configured in some embodiments to scale the frequency samples for each received device according to the energy of the signal segment and energy ratio of the selected frequency region with respect to the full band energy of a reference capture apparatus.
- capture apparatus related scaling can be represented mathematically by:
- the scaling with respect to a reference device provides a greater weighting to the captured content of the reference device with respect to other capture apparatus or devices.
- the ambience of the audio scene can be focussed on the captured content surrounding the reference device rather than generally across all of the devices.
- the sample scaler 603 can then be configured to output scaled sample values to a sample combiner 605.
- the apparatus comprises a sample combiner or suitable means for combining the at least two audio signals to generate a combined audio signal representation of the audio source.
- the means for combining in some embodiments can comprise means for appending the at least two audio signals.
- the sample combiner 605 can be configured to receive the scaled samples from each of the sample scaler sub-devices 603i to 603 n where there are N recording apparatus inputs and combine these in such a way to generate a single sample stream output.
- NID describes the number of apparatus of devices from which the data has been received.
- the combined samples can then be passed to an inverse transformer 607.
- the audio scene apparatus comprises an inverse transformer 607 or suitable means for frequency to time domain transforming the at least two appended audio signals to generate a time domain representation of the audio source.
- the inverse transformer 607 can be considered to be any inverse transformation associated with the time to frequency domain transformation described with respect to the recording apparatus.
- the time-to-frequency domain transform was a modified discrete cosine transform (MDCT)
- the inverse transform 607 can comprise an inverse modified discrete cosine transform (IMDCT).
- the encoding and decoding of the captured content can be applied such that backwards compatibility to existing audio codecs such as advanced audio coding (AAC) is achieved. For example in some embodiments this can be useful where the individually recorded content is consumed also individually. In other words a multi-device rendering is not performed.
- AAC advanced audio coding
- variations can be performed such that existing coding solutions provide well-established and tested methods for coding.
- the signal processing flow can be
- the additional subcomponents as described herein can be embedded into the audio codec bitstreams. These codecs typically offer ancillary data sections where various user data can be stored. In some further embodiments the additional subcomponents can be stored in separate files and not part multiplexed as of the coded bit stream.
- embodiments may also be applied to audio-video signals where the audio signal components of the recorded data are processed in terms of the determining of the base signal and the determination of the time alignment factors for the remaining signals and the video signal components may be synchronised using the above embodiments of the invention.
- the video parts may be synchronised using the audio synchronisation information.
- user equipment is intended to cover any suitable type of wireless user equipment, such as mobile telephones, portable data processing devices or portable web browsers.
- elements of a public land mobile network may also comprise apparatus as described above.
- the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
- some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
- firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
- While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general hardware or controller or other computing devices, or some combination
- the embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware.
- any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions.
- the software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
- the memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
- the data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
- Embodiments of the inventions may be practiced in various components such as integrated circuit modules.
- the design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
- Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules.
- the resultant design in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Signal Processing For Digital Recording And Reproducing (AREA)
Abstract
Description
Claims
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/IB2011/054565 WO2013054159A1 (en) | 2011-10-14 | 2011-10-14 | An audio scene mapping apparatus |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| EP2766904A1 true EP2766904A1 (en) | 2014-08-20 |
| EP2766904A4 EP2766904A4 (en) | 2015-07-29 |
Family
ID=48081432
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP11873915.0A Withdrawn EP2766904A4 (en) | 2011-10-14 | 2011-10-14 | AUDIO SCENE MAPPING APPARATUS |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US9392363B2 (en) |
| EP (1) | EP2766904A4 (en) |
| WO (1) | WO2013054159A1 (en) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10175931B2 (en) * | 2012-11-02 | 2019-01-08 | Sony Corporation | Signal processing device and signal processing method |
| US9602916B2 (en) | 2012-11-02 | 2017-03-21 | Sony Corporation | Signal processing device, signal processing method, measurement method, and measurement device |
| EP3284227B1 (en) * | 2015-04-16 | 2023-04-05 | Andrew Wireless Systems GmbH | Uplink signal combiners for mobile radio signal distribution systems using ethernet data networks |
| US10573291B2 (en) | 2016-12-09 | 2020-02-25 | The Research Foundation For The State University Of New York | Acoustic metamaterial |
| JP7020432B2 (en) | 2017-01-31 | 2022-02-16 | ソニーグループ株式会社 | Signal processing equipment, signal processing methods and computer programs |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EA011601B1 (en) * | 2005-09-30 | 2009-04-28 | Скуэрхэд Текнолоджи Ас | A method and a system for directional capturing of an audio signal |
| US20070081663A1 (en) * | 2005-10-12 | 2007-04-12 | Atsuhiro Sakurai | Time scale modification of audio based on power-complementary IIR filter decomposition |
| JP5337941B2 (en) | 2006-10-16 | 2013-11-06 | フラウンホッファー−ゲゼルシャフト ツァ フェルダールング デァ アンゲヴァンテン フォアシュンク エー.ファオ | Apparatus and method for multi-channel parameter conversion |
| US20110191112A1 (en) * | 2007-11-27 | 2011-08-04 | Nokia Corporation | Encoder |
| EP2396637A1 (en) * | 2009-02-13 | 2011-12-21 | Nokia Corp. | Ambience coding and decoding for audio applications |
| CA2996784A1 (en) | 2009-06-01 | 2010-12-09 | Music Mastermind, Inc. | System and method of receiving, analyzing, and editing audio to create musical compositions |
| JP5793675B2 (en) | 2009-07-31 | 2015-10-14 | パナソニックIpマネジメント株式会社 | Encoding device and decoding device |
| CN102630385B (en) | 2009-11-30 | 2015-05-27 | 诺基亚公司 | Method, device and system for audio scaling processing in audio scene |
| US9332346B2 (en) | 2010-02-17 | 2016-05-03 | Nokia Technologies Oy | Processing of multi-device audio capture |
-
2011
- 2011-10-14 WO PCT/IB2011/054565 patent/WO2013054159A1/en not_active Ceased
- 2011-10-14 US US14/351,326 patent/US9392363B2/en active Active
- 2011-10-14 EP EP11873915.0A patent/EP2766904A4/en not_active Withdrawn
Also Published As
| Publication number | Publication date |
|---|---|
| US9392363B2 (en) | 2016-07-12 |
| EP2766904A4 (en) | 2015-07-29 |
| US20150043756A1 (en) | 2015-02-12 |
| WO2013054159A1 (en) | 2013-04-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9820037B2 (en) | Audio capture apparatus | |
| CN109313907B (en) | Merge audio signals with spatial metadata | |
| US10097943B2 (en) | Apparatus and method for reproducing recorded audio with correct spatial directionality | |
| US20130226324A1 (en) | Audio scene apparatuses and methods | |
| CN111542877B (en) | Determination of spatial audio parameter encoding and associated decoding | |
| US20160155455A1 (en) | A shared audio scene apparatus | |
| EP2666160A1 (en) | An audio scene processing apparatus | |
| US20210400413A1 (en) | Ambience Audio Representation and Associated Rendering | |
| WO2013088208A1 (en) | An audio scene alignment apparatus | |
| WO2010125228A1 (en) | Encoding of multiview audio signals | |
| US9392363B2 (en) | Audio scene mapping apparatus | |
| EP2786594A1 (en) | Signal processing for audio scene rendering | |
| US12315523B2 (en) | Multichannel audio encode and decode using directional metadata | |
| US20150142454A1 (en) | Handling overlapping audio recordings | |
| CN115580822A (en) | Spatial audio capture, transmission and reproduction | |
| WO2012098427A1 (en) | An audio scene selection apparatus | |
| WO2012171584A1 (en) | An audio scene mapping apparatus | |
| US20150310869A1 (en) | Apparatus aligning audio signals in a shared audio scene | |
| WO2014083380A1 (en) | A shared audio scene apparatus | |
| CN103180907B (en) | audio scene device | |
| WO2015028715A1 (en) | Directional audio apparatus | |
| WO2013030623A1 (en) | An audio scene mapping apparatus | |
| WO2014016645A1 (en) | A shared audio scene apparatus |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20140410 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAX | Request for extension of the european patent (deleted) | ||
| RA4 | Supplementary search report drawn up and despatched (corrected) |
Effective date: 20150630 |
|
| RIC1 | Information provided on ipc code assigned before grant |
Ipc: G10L 19/16 20130101ALI20150624BHEP Ipc: G11B 20/10 20060101AFI20150624BHEP Ipc: H04R 5/027 20060101ALI20150624BHEP Ipc: H04R 3/00 20060101ALI20150624BHEP Ipc: G10L 19/02 20130101ALI20150624BHEP |
|
| RAP1 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: NOKIA TECHNOLOGIES OY |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20170503 |