EP4727164A1 - Audio apparatus and method of operation therefor - Google Patents

Audio apparatus and method of operation therefor

Info

Publication number
EP4727164A1
EP4727164A1 EP24205998.8A EP24205998A EP4727164A1 EP 4727164 A1 EP4727164 A1 EP 4727164A1 EP 24205998 A EP24205998 A EP 24205998A EP 4727164 A1 EP4727164 A1 EP 4727164A1
Authority
EP
European Patent Office
Prior art keywords
audio
capture
scene
signals
video
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24205998.8A
Other languages
German (de)
French (fr)
Inventor
Christiaan Varekamp
Sam Martin JELFS
Cornelis Pieter Janse
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Koninklijke Philips NV
Original Assignee
Koninklijke Philips NV
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Koninklijke Philips NV filed Critical Koninklijke Philips NV
Priority to EP24205998.8A priority Critical patent/EP4727164A1/en
Priority to PCT/EP2025/077735 priority patent/WO2026077738A1/en
Publication of EP4727164A1 publication Critical patent/EP4727164A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/20Arrangements for obtaining desired frequency or directional characteristics
    • H04R1/32Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only
    • H04R1/40Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only by combining a number of identical transducers
    • H04R1/406Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only by combining a number of identical transducers microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/005Circuits for transducers for combining the signals of two or more microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/15Aspects of sound capture and related signal processing for recording or reproduction

Landscapes

  • Health & Medical Sciences (AREA)
  • Otolaryngology (AREA)
  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • General Health & Medical Sciences (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

An audio apparatus comprises a video signal receiver (201) receiving video signals from cameras capturing a scene, and an audio signal receiver (203) receives audio signals from microphones capturing audio of the scene. An object selector (205) selects an object of the scene by selecting an image object in a video signal and an object position determiner (207) determines an object position for the object in the scene from detections of the object in images of the video signals and from known poses of the cameras. The distance from the object position to the audio capture poses does not exceed five times the largest distance between the poses. An audio mixer (209) generates an audio mix signal by mixing the audio signals. At least some mixing weights for the audio signals are determined as a function of the first object position relative to audio capture poses for the audio signals.

Description

    FIELD OF THE INVENTION
  • The invention relates to generation of an audio mix signal from a plurality of audio signals and in particular, but not exclusively, to generation of an audio signal capturing audio in a scene using a large and distributed audio capture pose arrangement.
  • BACKGROUND OF THE INVENTION
  • Capturing live scenes has become increasingly important in the last decades. For example, capturing audio and video for large events, such as sports or concerts, has become widespread. Further, the demands and requirements for the provided services, and thus for the audio and video capture, have increased substantially. For example, spatial audio capture and rendering is becoming prevalent and accordingly there is need for accurate and efficient capture, distribution, and rendering of spatial audio.
  • In many practical applications, audio is captured by a plurality of microphones at different positions. For example, a plurality of microphones is often used to capture audio in an environment. In particular, when audio of a large environment/scene is to be captured, such as when capturing audio in a sports arena or stadium, it is necessary to deploy a substantial number of widely distributed microphones in order to ensure sufficient capture across the entire scene. Such multiple distributed microphones further provide spatial information for the audio in the scene. The use of multiple microphones allows spatial information of the audio to be captured and applications have been developed that exploit such spatial information in order to allow or support improved and/or new services.
  • One frequently used approach is to try to separate audio sources by applying audio beamforming to form beams in the direction of arrival of audio from specific audio sources. However, although this may provide advantageous performance in many scenarios, it is not optimal in all cases. For example, it may not provide optimal source separation in many cases, and indeed in some applications such a spatial beamforming may not provide audio properties that are ideal for further processing to achieve a given effect. In particular, audio beamforming tends to only be efficient or useful when the microphone array is small and with the inter-microphone distance being low (typically with the microphone array extending over less than 50cm and with inter-microphone distances being around 5-10cm). Beamforming in general is not suitable for large microphone distances or for audio sources that are too close to the microphones, and specifically is not suitable for audio sources that are close to the microphones relative to the size of the microphone distribution size. Specifically, whereas beamforming in many scenarios may be suitable for small microphone arrays capturing audio sources in the far field, it tends to not perform well for large microphone arrangements and/or for audio sources in the near field.
  • Accordingly, there is a desire to improve audio capture and processing for large and distributed microphone arrangements and/or for flexibly capturing audio relatively close to microphone arrays.
  • Hence, an improved audio capture/processing approach would be advantageous, and in particular an approach allowing reduced complexity, increased flexibility, facilitated implementation, reduced cost, improved audio capture, improved spatial perception/differentiation of audio sources, improved audio source separation, improved spatial audio/processing, improved adaptation for audio sources in the near field, flexibility and customization to different audio environments and scenarios, improved capture of large areas using large microphone arrangements, improved potential for emphasizing and/or de-emphasizing specific audio sources, an improved trade-off between performance and complexity/ resource usage, and/or improved performance would be advantageous.
  • SUMMARY OF THE INVENTION
  • Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.
  • According to an aspect of the invention there is provided an apparatus for generating an audio mix signal, the apparatus comprising: a video signal receiver arranged to receive video signals from one or more video cameras capturing a scene, each video camera having a video capture pose; an audio signal receiver arranged to receive audio signals from a plurality of microphones capturing audio of the scene, each microphone of the plurality of microphones having an audio capture pose; an object selector arranged to select a first object of the scene by selecting an image object in at least one of the video signals; an object position determiner arranged to determine a first object position for the first object in the scene from detections of the first object in images of the video signals and the video capture poses, the first object position being in a part of the scene for which a distance to a closest of the audio capture does not exceed five times a largest distance between the audio capture positions; and an audio mixer arranged to generate the audio mix signal by mixing the audio signals, at least some mixing weights for at least some audio signals of the audio signals being determined as a function of the first object position relative to audio capture poses for the at least some audio signals.
  • The invention may provide improved audio capture in many embodiments. In particular, it may provide an improved audio representation of a large scene-area using a large microphone/audio capture arrangement. The approach may allow efficient audio capture and processing for audio sources in the near field of a large audio capture/microphone arrangement. The approach may in particular allow efficient emphasis or de-emphasis of individual audio sources for visual objects in a scene even when using very large distributed audio capture arrangement.
  • The approach may provide an improved trade-off allowing audio capture across a large area by using a large distributed microphone/capture arrangement while still allowing adaptation to emphasize and/or de-emphasize individual audio sources for specific scene objects in the near field. The approach allows e.g. a user to select a specific visual object and to emphasize or de-emphasize the audio associated with the object in the audio mix. The large distributed capture area results in a large near-field which typically tends to prevent audio beamforming to effectively select or de-select audio sources. The described approach allows adaptation of near-field audio thereby providing a flexible approach that may support both large scene areas and individual adaptation.
  • In some embodiments, a distance between at least two of the capture poses is no less than 5, 10, 20, or 50 meters. In many embodiments, a distance between nearest audio capture positions may be no less than 1, 2, 3, or 5 meters. In some embodiments, the first object position is in a part of the scene for which a distance to a closest microphone/audio capture position does not exceed three, two or even one times a largest distance between the audio capture positions
  • In some embodiments, the audio mixing may be a non-coherent mixing/combination/summation of the weighted microphone audio signals. In many embodiments, the mixing weights may be scalar values. In many embodiments, mixing weights may be gains for the audio signals. The mixing may be a weighted combination/summation.
  • A pose may be a position and/or direction/orientation.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to determine a mixing weight for a first audio signal of the plurality of audio signals depending on a distance from the first object position to an audio capture position for the first audio signal.
  • This may provide improved performance and/or operation in many embodiments and may in particular provide a highly advantageous adaptation of individual audio sources linked with a visual scene object.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to determine a mixing weight for a first audio signal depending on a direction from the first object position to an audio capture position for the first audio signal and a directional sensitivity pattern for a microphone capturing the first audio signal.
  • This may provide improved performance and/or operation in many embodiments and may in particular provide a highly advantageous adaptation of individual audio sources linked with a visual scene object.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to generate sound level attenuation estimates for acoustic propagation from the first object position to the audio capture poses, and to determine the at least some mixing weights in dependence on the acoustic energy loss estimates and a predetermined constraint on the weights.
  • This may provide improved performance and/or operation in many embodiments.
  • In some embodiments, the audio mixer is arranged to generate acoustic propagation loss estimates/acoustic energy loss estimates/intensity reduction estimates, geometric acoustic energy loss estimates for acoustic propagation from the first object position to the audio capture poses, and to determine the at least some mixing weights in dependence on the estimates and a predetermined constraint on the weights.
  • In accordance with an optional feature of the invention, the constraint is a constraint on a combination of the weights.
  • This may provide improved performance and/or operation in many embodiments.
  • In some embodiments, the constraint is a constraint on a summation of the weights.
  • In accordance with an optional feature of the invention, the video signal receiver is arranged to determine at least one video capture pose from the video signals.
  • This may provide improved performance and/or operation in many embodiments.
  • In accordance with an optional feature of the invention, the audio signal receiver is arranged to determine at least one audio capture pose from at least one video capture pose and a predetermined relationship between the at least one audio capture pose and the at least one video capture pose.
  • This may provide improved performance and/or operation in many embodiments.
  • In accordance with an optional feature of the invention, the apparatus further comprises a user interface arranged to select the first object in response to a user input.
  • This may provide improved performance and/or operation in many embodiments.
  • In accordance with an optional feature of the invention, the user interface comprises a display for displaying an image of the scene generated from at least one of the video signals and a selection input for selecting the first object in response to the user selecting an image position on the image.
  • This may provide improved performance and/or operation in many embodiments.
  • In some embodiments, the user interface is arranged to determine a scene position corresponding to the image position and to determine the first object as an object being proximal to the scene position.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to apply a variable delay to at least one microphone audio signal prior to the audio mixing, the delay being dependent on a distance between the first object position and the audio capture pose for the at least one microphone audio signal.
  • This may provide improved performance and/or operation in many embodiments.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to determine at least one gain for a first microphone signal in dependence on an observer position for an observer in the scene.
  • This may provide improved performance and/or operation in many embodiments.
  • In accordance with an optional feature of the invention, the audio mixer is arranged to determine the gains in dependence on estimated acoustic propagations from the object pose to the audio capture poses, the estimated acoustic propagations being determined based on an acoustic propagation model, the audio mixer being arranged to switch between a spherical wave acoustic propagation model for a distance between the first object position and the audio capture pose being below a first threshold distance and between a plane wave acoustic propagation model for the distance between the first object position and the audio capture pose being above a second threshold distance being larger than the first threshold distance.
  • This may provide improved performance and/or operation in many embodiments. The first and second thresholds may specifically be the same threshold. The first threshold may be used to correspond to/define an area/maximum distance for what may be considered to be the near-field. For the near-field, a spherical wave acoustic propagation model provides an advantageous and typically accurate modelling of the acoustic propagation. The second threshold may be used to correspond to/define an area/minimum distance for what may be considered to be the far-field. For the far-field, a plane wave acoustic propagation model provides an advantageous and typically accurate modelling of the acoustic propagation.
  • In some embodiments, the audio receiver is arranged to receive a first audio object linked with the scene audio object and the audio mixer is arranged to include the first audio object in the audio mix signal.
  • In accordance with an optional feature of the invention, at least one microphone of the plurality of microphones is a beamforming directional microphone array.
  • This may provide improved performance and/or operation in many embodiments.
  • In some embodiments, at least one microphone of the plurality of microphones may be a spherical microphone array or high-order ambisonic array, from which the audio signal may be derived and the directionality of the virtual microphone controlled / steered.
  • In accordance with an optional feature of the invention, the object position determiner arranged to determine a second object position for a second object in the scene from detections of the second object in the video signals and the video capture poses, the second object being in a part of the scene for which a distance to a closest of the audio capture positions exceeds the largest distance between the audio capture positions by no less than ten times; and wherein the audio mixer arranged is arranged to determine the at least some mixing weights as a function of the second object position relative to audio capture poses for the at least some microphone signals.
  • This may provide improved performance and/or operation in many embodiments.
  • According to another aspect of the invention, there is provided a method of generating an audio mix signal, the method comprising: receiving video signals from one or more video cameras capturing a scene, each video camera having a video capture pose; receiving audio signals from a plurality of microphones capturing audio of the scene, each microphone of the plurality of microphones having an audio capture pose; selecting a first object of the scene by selecting an image object in at least one of the video signals; determining a first object position for the first object in the scene from detections of the first object in images of the video signals and the video capture poses, the first object position being in a part of the scene for which a distance to a closest of the audio capture does not exceed five times a largest distance between the audio capture positions; generating the audio mix signal by mixing the audio signals, at least some mixing weights for at least some audio signals of the audio signals being determined as a function of the first object position relative to audio capture poses for the at least some audio signals.
  • These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.
  • BRIEF DESCRIPTION OF THE DRAWINGS
  • Embodiments of the invention will be described, by way of example only, with reference to the drawings, in which
    • FIG. 1 illustrates an example of an audio and video capture of a real world scene;
    • FIG. 2 illustrates an example of elements of an audio apparatus in accordance with some embodiments of the invention; and
    • FIG. 3 illustrates some elements of a possible arrangement of a processor for implementing elements of an audio apparatus in accordance with some embodiments of the invention.
    DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION
  • An example of an audio and video capture setup for capturing audio and video of a real-world scene is illustrated in FIG. 1. In the example, video of the scene is captured from three different video capture positions/poses and from four different audio capture positions/poses. The figure shows an example where the audio capture is by four different microphones 101 and three different video cameras 103 but it will be appreciated that this is merely an illustration and that substantially more audio and/or video capture positions/poses (and microphones and/or video cameras) may be applied in many scenarios.
  • In the field, the terms placement and pose are used as a common term for position and/or direction/ orientation. The combination of the position and direction/ orientation of e.g. an object, a camera, a microphone, a head, or a view may be referred to as a pose or placement. Thus, a placement or pose indication may comprise six values/ components/ degrees of freedom with each value/ component typically describing an individual property of the position/ location or the orientation/ direction of the corresponding object. Of course, in many situations, a placement or pose may be considered or represented with fewer components, for example if one or more components is considered fixed or irrelevant (e.g. if all objects are considered to be at the same height and have a horizontal orientation, four components may provide a full representation of the pose of an object). In the following, the term pose is used to refer to a position and/or orientation which may be represented by one to six values (corresponding to the maximum possible degrees of freedom). A pose may be a position and/or orientation.
  • A pose having the maximum degrees of freedom, i.e. three degrees of freedom of each of the position and the orientation resulting in a total of six degrees of freedom. A pose may thus be represented by a set or vector of six values representing the six degrees of freedom and thus a pose vector may provide a three-dimensional position and/or a three-dimensional direction indication. However, it will be appreciated that in other embodiments, the pose may be represented by fewer values. Similarly, in some embodiments, more than six values may be used to represent a pose. For example, rotation/orientation may be given in quaternions which means a pose may be represented by seven values
  • (three for position and four for rotation)
  • A pose may be at least one of an orientation and a position. A pose value may be indicative of at least one of an orientation value and a position value. A pose may be a position and/or orientation.
  • A system or entity based on providing the maximum degree of freedom for the viewer is typically referred to as having 6 Degrees of Freedom (6DoF). Many systems and entities provide only an orientation or position and these are typically known as having 3 Degrees of Freedom (3DoF).
  • In the example, a number of the audio and video capture poses are co-located but this is merely an option and in many embodiments some or all of the audio and video capture poses may be different from each other. As in the example of FIG. 1, the number of video capture and audio capture poses may be different, or may indeed be the same in some embodiments.
  • The capture arrangement may specifically be arranged to capture audio and video for a large scene, such as a concert venue or a stadium (e.g. during a sports event). For example, the audio and video capture arrangement may seek to capture video and audio for a scene area in excess of 20m2, 50m2, 100m2, or 1000m2. In order to capture such a large scene, the capture should preferably be performed over a large area and the audio capture poses and video cameras are accordingly preferably distributed over a large area. The distance between the audio capture poses is accordingly large and is typically substantially larger than what is suitable for performing audio beamforming. In many embodiments, inter-capture distances between nearest neighbor microphones/audio capture positions may be no less than 1m, 2m, 5, 10m, or 50m for at least some of the audio capture poses. The largest distance between audio capture poses may be no less than 2m, 5m, 10m, 50m, or 100m in many embodiments.
  • In the example, the video cameras 103 may be angled towards a common focal point 105. Further, the microphones 101 may have a directional sensitivity and may in the specific example also be directed towards the focal point 105. However, it will be appreciated that this is merely exemplary, and that different microphones and capture configurations may be used in other embodiments.
  • FIG. 2 illustrates example elements of an audio processing apparatus which is arranged to process audio signals captured by a plurality of microphones. The processing of the apparatus is particularly suitable for providing a flexible and adaptive processing of the captured audio to generate an audio signal representing audio in a large scene while maintaining low complexity and resource usage.
  • The audio apparatus of FIG. 2 comprises a video signal receiver 201 which is arranged to receive video signals from one or more video cameras capturing a scene, where each video camera has a video capture pose (position and/or orientation). Thus, the video signals represent/capture views of the scene from the specific capture poses. The following description will focus on a scenario in which a plurality of video cameras and video capture poses are employed but it will be appreciated that in some embodiments only a single video camera/pose may be employed. For example, deep learning neural networks (typically so called vision transformers) that allow the per pixels estimation of depth for a single camera have been developed and may be applied in some embodiments.
  • The audio apparatus further comprises an audio signal receiver 203 arranged to receive audio signals from a plurality of microphones capturing audio of the scene. Each of the microphones has a capture pose corresponding to the position/orientation of the microphone in the scene. Thus, each of the microphone signals is linked with a capture pose and represents a capture of audio of the scene from that capture pose. In some embodiments, only two microphones/capture positions may be employed but in many embodiments the plurality of microphones may include substantially more microphones, and indeed in some embodiments 5,10, or even more microphones may be considered.
  • In many embodiments, the distance between at least two of the capture positions is no less than 1 meter, 2 meter, 5 meters, 10 meters, or even 50 meters. Thus, the microphone signals provide a large scale capture of the audio of the scene.
  • The audio apparatus is arranged to generate an audio mix signal by mixing the microphone audio signals based on the video signals. The audio apparatus allows a flexible generation of the audio output signal allowing this to adapt to the scene, including e.g. emphasizing or de-emphasizing audio from specific objects in the scene. The approach in particular allows such advantageous operations to be applied to objects that are in the near field of the microphone arrangement.
  • The audio apparatus comprises an object selector 205 which is arranged to select an object in the scene from an analysis/processing of the video signals. The object selector 205 is coupled to an object position determiner 207 which is arranged to determine an object position for the selected object from detections of the selected object in the video signals.
  • In some embodiments, the audio apparatus may include a user interface 211 coupled to the object selector 205 and allowing a user to specifically select an object, such as for example by manually selecting an image object in a view image of the scene generated from the view signals and presented on a suitable display. As another example, the selection of the object may be an automated process such as by detecting an object that has properties sufficiently matching a set of predetermined qualities (for example particular color and/or size).
  • The object position determiner 207 may proceed to determine the position of the selected object in the scene. Specifically, the coordinates of the object in a scene coordinate system may be determined. It will be appreciated that many different approaches and algorithms for determining a scene position based on images from one or multiple different viewpoints are known to the skilled person and that any suitable approach may be used.
  • As a low complexity example, image objects corresponding to an image object selected by a user may be determined in the images from different cameras (different video signals) and based on the relative positions of the image object in the view images and knowledge of the positions of the video cameras, relatively simple geometric calculations may be performed to determine a 3D position of the object in the scene.
  • In particular, techniques for determining a depth for scene objects based on parallax between images from different capture positions can be used to estimate a distance from the capture positions to the object in the world/scene coordinate system. The direction from given capture positions can be determined by comparing the positions of the image object in the images and considering the orientation of the camera. The object position may be estimated based on different images and the object position in the scene/world space may be determined by e.g. averaging multiple estimates from different images (or image pairs).
  • In some situations where audio objects are in the far field and the microphone arrangements are sufficiently small, audio beamforming may be suitable for adapting an audio signal representing the scene, such as for example by using audio beamforming to focus on a particular audio source (or to attenuate a particular audio source). However, for objects that are close to the microphone arrangement relative to the size of the arrangement, and in particular which are in the near field, audio beamforming is not a feasible approach and therefore audio representations for audio objects in the near field may be achieved simply by selecting the audio of the nearest microphone.
  • The audio apparatus of FIG. 2 is arranged to allow a flexible adaptation of the generation of an audio mix signal representing the audio in a scene which in particular provides an efficient adaptation of the representation of an audio object in the near field, and in particular allow audio for such an object and audio source to be emphasized (increased relative level) or de-emphasized (reduced relative level). The audio apparatus is in particular arranged to perform an advantageous mixing that can provide emphasis/de-emphasis for audio sources/selected scene objects that are in a part of the scene for which the distance to the nearest audio capture position does not exceed five times, or in many cases even three, two, or one times, a largest distance between the audio capture positions. Thus, the distance from the nearest audio capture position/microphone is less than 5, 3, 2, 1 times the maximum distance between capture positions for the audio signals.
  • The audio apparatus comprises an audio mixer 209 which is arranged to mix the microphone audio signals to generate an audio mix signal. The mixing weights/gains for at least some of the microphone signals are determined as a function of the position of the object relative to the audio capture poses. For example, in scenarios where the audio apparatus is seeking to emphasize audio generated by a selected object, the mixing of the audio signals may apply a weight to the individual microphone signals which is monotonically decreasing with an increasing distance from the audio capture position for the microphone signal to the object position. Conversely, in scenarios where the audio apparatus is seeking to de-emphasize audio generated by a selected object, the mixing of the audio signals may apply a weight to the individual microphone signals which is monotonically increasing with an increasing distance from the audio capture position for the microphone signal to the object position. Thus, in the former scenario, the signals that are more likely to include substantial audio power/energy from the selected object are weighted relatively higher than the signals more likely to include less audio power, and in the latter scenario, the signals that are more likely to include substantial audio power/energy from the selected object are weighted relatively lower than the signals more likely to include less audio power. The weights may typically be due to various constraints, such as a minimum value to ensure some contribution/energy being included for each microphone signal to ensure a broad coverage of audio in the scene and/or a constraint that the weights are adapted to ensure that that power/energy level does not change.
  • The audio mixing may typically be a weighted combination/summation where each microphone signal is weighted by weights with at least one weight being adapted based on the selected object position relative to the audio capture positions. The weights may in many embodiments be scalar values and thus may in many embodiments be gains that are applied to the microphone signals.
  • If we weigh the microphone signals pi with weights wi then for uncorrelated signals we get: p mix 2 = i = 1 N mics w i p i 2 where the summation may be over all microphone signals and the weights wi may specifically be scalar (and typically positive) values.
  • In many embodiments, the gain for a given microphone signal may depend on the distance from the object position to the audio capture position for the microphone signal.
  • For example, if it is desired to emphasize audio from the selected object in the generated audio mix, the gain/weight for a given microphone signal may be a monotonically decreasing function of the distance from the object position to the capture position for the microphone capturing that microphone signal. Thus, in such a case, the relative power (of the captured signal) in the mix of audio capture close to the object will be increased with respect to the relative power of audio capture further from the object. As a result, the object will tend to be perceived as louder and more emphasized in the audio mix than if all gains/relative intensities were identical.
  • Conversely, if it is desired to de-emphasize audio from the selected object in the generated audio mix, the gain/weight for a given microphone signal may be monotonically increasing function of the distance from the object position to the capture position for the microphone capturing that microphone signal. Thus, in such a case, the relative power in the mix of audio capture far from the object will be increased with respect to the relative power of audio capture closer to the object. As a result, the object will tend to be perceived as quieter and less emphasized in the audio mix than if all gains/relative powers were identical.
  • In some embodiments, one or more of the microphones may be directional microphones and thus may have a given directional sensitivity pattern. In such cases, the gain for a given microphone signal may be dependent on the direction from the capture position to the object position and the directional sensitivity pattern.
  • For example, the relative gain for the directional microphone in the direction from the capture position to the object position may be determined. In cases where the audio from the object is desired to be emphasized, the gains may then be increased for microphone signals for which the gain in the direction towards the object position is high, and may be decreased for microphone signals for which the gain in the direction towards the object position is low. Conversely, if the desire is to de-emphasize the audio from the selected object, the gains may be decreased for microphone signals for which the gain in the direction towards the object position is high and may be increased for microphone signals for which the gain in the direction towards the object position is low.
  • As a specific example, one or more of the microphones may in some embodiments be spherical arrays where the directional response of the microphone may also be adapted to provide the best sensitivity/ attenuation of a desired source.
  • In some cases, both the directional sensitivity pattern and the direction and distance from the capture position to the object position may be considered. For example, a sound level estimate such as a propagation attenuation from the object position to the audio capture position may be determined taking into consideration both the distance and the microphone sensitivity in the given direction. The mixing gains for the microphone signals may then be set dependent on the propagation attenuations. In order to emphasize the object audio, the gains may be increased for a microphone signal with low propagation attenuation relative to gains for microphone signal with high propagation attenuation. In order to de-emphasize the object audio, the gains may be decreased for microphone signal with low propagation attenuation relative to gains for microphone signal with high propagation attenuation.
  • Thus, in many embodiments, the audio mixer may be arranged to generate sound level attenuation estimates (or acoustic propagation loss estimates/acoustic energy loss estimates/ intensity reduction estimates, geometric acoustic energy loss estimates) for the acoustic propagation from the object position to the audio capture poses may be determined and the gains may then set dependent on these estimates.
  • The mixing weights may typically be determined subject to a predetermined constraint on the weights. In many cases, the constraint may for example be that each mixing gain is subject to a minimum value. Another constraint may be that all mixing weights/gains combine to a given value, such as e.g. to maintain a constant amplitude and/or energy.
  • Thus, in many embodiments, the constraint may be a constraint on the combination, and specifically on the summation, of the weights.
  • As a specific example, given all the objects in a scene, it may be desired to increase/decrease the relative sound level of a single object. To do this, the audio mixer 209 may proceed to take all objects into account by first calculating the contribution of each object/audio source to each microphone intensity assuming free space sound propagation for each source. For a given audio capture position/microphone i, we can model the measured sound pressure received from source j as: p i , j 2 = g d i , x j x i x j x i 2 p j 2 where we assume that to a first approximation, the received squared sound pressure drops with the inverse of the squared distance between the object/sound source position x j and the receiving microphone x i . The directional microphone has a pickup sensitivity pattern that results in a gain g that depends on the microphone direction vector d i and on the vector x j - x i that points from the microphone to the sound source.
    • As a specific example, it may be assumed that the microphone has a cardioid pickup pattern of the form: r φ = 2 a 1 cos π φ where a defines the maximum pickup which is a microphone characteristic. The pickup function r(φ) may be defined on a logarithmic scale (in dB) and the intensity ratio on a linear scale is therefore: Intensity ratio = 10 r φ 10
    • The pickup on a linear intensity scale can therefore be written as: r linear φ = 10 2 a 1 cos π φ 10
    • Since the term cos(π - φ) can be written as a function of microphone direction vector d i , the microphone position x i and the sound source position x j , the pickup/sensitivity function and hence the gain can also be expressed as a function of these variables: g d i , x j x i = 10 a 5 1 d i x j x i d i x j x i
    • Finally, we have: p i , j 2 = G i , j p j 2 where Gi,j is the geometric contribution/gain/weight to the audio mix for object/audio source j to microphone signal/microphone i : G i , j = 10 a 5 1 d i x j x i d i x j x i x j x i 2
    • As another example, for a cardioid pickup pattern r φ = 2 a 1 cos π φ that is already on a linear scale, the following result is achieved: g d i , x j x i = 2 a 1 d i x j x i d i x j x i
    • The cardioid pattern is related to the pressure, so we have: p i , j = G i , j p j and the geometric contribution Gi,j in the audio mix becomes: G i , j = g x j x i = 2 a 1 d i x j x i d i x j x i x j x i = 2 a 1 x j x i d i x j x i d i x j x i 2
    • As another example, if it may be desired to minimize the sound for object k without reducing the sound for other objects. This can be done by creating an audio mix of the microphone signals pi. p mix 2 = i = 1 N mics w i p i 2 = i = 1 N mics w i j = 1 N objects p i , j 2
    • A-priori the sound source pressure p j 2 at source j is typically not known. Assuming all sources have the same constant source level we can drop p j 2 from the above equations for the purpose of finding the weights for the audio mix. We want to find the coefficients wi = w 1 ... w N mics that minimizes the contribution from object k: w mix = arg min w i = 1 N mics w i G i , k subject to i = 1 N mics w i > T
    • Requiring that the sum of weights remain larger than a threshold (e.g. T > 0.5) guarantees that the overall audio level is not reduced too much.
  • To maximize the sound level for object k we can replace arg min w by arg max w subject to i = 1 N mics w i < T Requiring that the sum of weights remains smaller than a threshold (e.g. T < 0.5) guarantees that the overall audio level is not increased too much.
  • Thus, the audio apparatus allows the sound for a given selected object to be controlled in general by modifying the mixing weights such that the geometric contribution Gi,k is changed (specifically increased or decreases) for a given object k.
  • The above described detailed approaches do not consider the effects on other audio objects by setting constraints on e.g. the weight sum. However, the individual positions of the other objects relative to the position and direction of microphones determines the individual effects. Instead of a global constraint on the weights, individual constraints can be set directly based on contributions of other objects jk in the resulting audio mix relative to a reference situation (e.g. when all microphones are mixed with equal weight).
  • The approach is useful e.g. to emphasize an individual object versus other objects, or versus a diffuse field. The total level can e.g. be controlled using an Automatic Gain Control (AGC). The approach specifically provides such effects to objects that are in the near-field of the capture environment and for objects that are close to the capture arrangement with respect to the size of the capture arrangement. The scenario and suitable processing requirements for audio sources that are close to the capture arrangement with respect to the size of the arrangement is fundamentally different from the characteristics for audio sources that are in the far field and is at a substantial distance to the capture arrangement. Indeed, for such far field audio sources, the difference in propagation loss/attenuation is relatively small and audio tends to be received at the same sound level at the different microphones. This may result in the conditions being suitable for audio beamforming. However, for close audio sources in the near-field, the sound level in the different microphone signals may be substantially different and much less correlated, and this may render the signals unsuitable for beamforming. However, the Inventors have realized that an adaptive audio mixing approach as described may allow individual objects to be selected in the scene based on the captured video with the sound of such individual objects then being adapted (typically emphasized or de-emphasized).
  • In particular, audio beamforming (such as filter-and-sum or delay-and-sum beamformers) are based on considering that the sound field can be approximated by plane waves. However, this is not a suitable consideration for audio sources in the near field. In the described examples, the capture arrangement may be very large, and specifically may extend over 10s or even 100s of meters, e.g. to cover an entire stadium or concert venue. As a consequence, the near field is also very large, and the arrangement is typically not suitable for the use of beamforming to emphasize or de-emphasize specific audio sources.
  • In some embodiments, the video signal receiver my be arranged to determine one or more of the video captures pose from the captured video signals. For example, the center camera may be considered to be at a nominal position and nominal direction, e.g. position coordinates may be set to (0,0,0) and orientation coordinates to (1,0,0). The other cameras may then be detected in the captured video signals and based on their image position, the direction to the other cameras may be estimated. Similarly, in the video signals of other cameras, the directions to other captured cameras may be estimated. Based on the estimated directions, the coordinates and orientations of the cameras may then be estimated.
  • As a particular example, a Structure from Motion (SfM) algorithm may be used to calculate the poses (typically position and orientation) of each camera in world/scene space. An example of such an approach may be found in the well-known software package COLMAP (https://colmap.github.io/) which is the result of years of research into automatic 3D calibration of multi camera systems based on image features detected.
  • In many embodiments, one, more, or all of the audio capture poses may be determined from one (or more) video capture poses and a predetermined relationship between this and the audio capture pose(s). In many embodiments, one, more, or all of the microphones may be attached with a known relative pose to each camera. For example, in many embodiments, a microphone may be implemented as an integral or fixedly attached microphone for the individual camera. The offset between the (optical) pose of the camera and the microphone may accordingly be fixed and known. For example, a user may prior to the audio apparatus being used to generate the audio mix provide a user input which for each camera indicates the offset between the microphone capture pose and the camera capture pose. The information may for example be provided to the user as part of the documentation for the video camera. During operation, the audio capture pose may then be determined by offsetting the corresponding video camera pose by the indicated offset.
  • As previously described, in some embodiments, the audio apparatus may include a user interface 211. The user interface 211 may specifically include a user input function with the object selector 205 being arranged to select the object based on the user input.
  • For example, in some embodiments, the user interface may allow the user to input a number of characteristics of the desired object, e.g. by entering keywords or allowing a selection between a number of predetermined options. This may for example allow a user to define a color, texture, pattern, size etc. of the desired object. The object selector 206 may then analyze the captured images to detect an image object having matching properties.
  • In many embodiments, the user interface 211 may include a display which may render views of the scene based on the received video signals. In some embodiments, the user interface 211 may be arranged to present one of the video signals and specifically a fixed pose view may be presented. In other embodiments, the user interface 211 may be arranged to generate and present view images from different viewpoints/poses, and indeed in many embodiments the user may control the view pose for which a view image is generated. It will be appreciated that many different algorithms and techniques for rendering view images for desired view poses based on captured views for other poses are known to the skilled person.
  • In some embodiments, the user may be arranged to identify/select an object by selecting a given image object in the rendered images. For example, a mouse, joystick, or touch display input may be used to select a specific image object. The object position detector 207 may then proceed to determine the position of the corresponding object in the scene space.
  • In many embodiments, the audio mixer may be arranged to generate the audio mix to only include a scaling (gain/attenuation) of the microphone signals prior to a combination, and specifically summation, of the scaled microphone signal.
  • However, in some embodiments, the audio mixer 209 may also be arranged to apply a variable delay to one or more of the microphone audio signals prior to the audio mixing/combination. The variable delay is made dependent on the distance between the object position and the audio capture pose for the microphone audio signal(s).
  • In particular, a variable delay may be applied to the microphone signals to compensate for the differences in the path length, and thus delay, between the object and the different microphones. The delays may then be adapted to account for the travel time differences between the sound source and microphones. In particular, as the position of the microphones and the sound source for the selected object are known, the audio signals can be delayed/time shifted depending on the Euclidian distances between the microphones and object.
  • In some embodiments, the audio mixer is arranged to determine one or more of the mixing gains/ for a first microphone signal in dependence on an observer position for an observer in the scene. The observer position may be a position for which a view image is generated/rendered based on the video signals and/or may be a desired listening position.
  • The audio apparatus may for example include a separate geometric observer weight in all above equations where the observer weight depends on the distance from the observer to each of the candidate sound sources. This observer weight enters the above equations as a multiplier. i.e. the weight that emphasizes or de-emphasizes a given candidate source sources is multiplied with an observer weight that increases monotonically with decreasing distance to that particular candidate source. Note that in the simplest case, this observer weight is directionally invariant but that when the observer wears a XR glasses and/or headphones or in-ear devices, the head pose of the observer needs to be accounted for in the observer weight. The modeling here is similar to the directionality of the microphones but now relates to the head pose relative to the sound sources.
  • Such an approach may provide advantageous operation and services. For example, it may allow a user to virtually change observer position with the presented video and audio automatically adapting to the changes in the observer position. For example, a user viewing a capture of a football match in a stadium may switch observer position (e.g. between center field and close to a goal) with not just the presented video but also the audio adapting in response.
  • In some embodiments, the audio apparatus may also include an audio object source 213 which may provide one or more audio objects which may be fed to the audio mixer 209 and included in the audio mix signal.
  • In some embodiments, the audio object may be a virtual audio object which can be added to the audio mix. In other examples, the audio object may for example represent a specific audio source of the scene, such as for example a dedicated microphone positioned close to a specific audio source (e.g. a reference wearing a microphone).
  • The audio mixer 209 may be arranged to include such audio objects in the generated audio mix. For example, it may simply include such audio object(s) as another term in the combination/summation/mixing. The gain/weight for audio objects may for example be predetermined, or may e.g. be dependent on the current signal level. For example, the gain/weight may be set to correspond to the intensity of the audio object in the audio mix being a predetermined percentage of the total intensity of the audio mix. Such an approach may for example allow e.g. the speech of a referee to be captured and included in the audio mix at a level that ensures that it is clearly perceptible.
  • In many embodiments, each of the microphones may be a single microphone that captures audio with a specific directional pattern. For example, cardioid and/or omnidirectional microphones may be used.
  • In some embodiments, one or more of the microphones may specifically be a beamforming directional microphone array. Thus, in some embodiments, one or more of the individual microphones may be implemented by an array of a plurality of microphone elements that each capture audio. The distance between nearest microphone elements may typically be relatively small, such as not exceeding 10cm, 20cm, or 50cm respectively. This may allow for accurate beamforming for areas relatively close to the array and may provide an efficient approach for achieving directional sensitivity.
  • In many cases, the beamforming may be made adaptive, and it may for example be arranged to adapt the directional sensitivity depending on the selected object position and the capture position for the array. Thus, the adaptive beamforming microphone array may be arranged steer the maximum sensitivity towards (or away from) the object thereby providing an additional means of emphasizing/de-emphasizing the individual audio source of the selected object.
  • Thus, in many approaches, far-field beamforming may be applied for individual capture positions in combination with near-field audio mixing based on a large audio capture arrangement. Such an arrangement may thus allow both far-field and near-field processing of the same audio to be applied to provide improved operation and allowing both efficient emphasis/de-emphasis of individual audio for a selected object and capture of a large area.
  • In some embodiments, a plurality of microphones may for example be provided for each camera. Adding multiple microphones per camera unit can help to improve sound level control quality. For example, by mounting three microphones on each camera unit, one in the direction of the camera, one at an azimuth angle of -90 degree and one at an azimuth angle of +90 degree. This may make it possible to better suppress sound coming from a certain position in the scene space.
  • In some embodiments, the audio mixer is arranged to determine the gains in dependence on estimated acoustic propagations from the object pose to the audio capture poses where the estimated acoustic propagations are determined based on an acoustic propagation model. For example, a given propagation model may be used to estimate the propagation loss from the selected object position to the capture position. However, in addition, the audio mixer 209 may be arranged to use different propagation models depending on whether the distance from the capture position to the selected object position is above or below a given threshold. The threshold may typically be given as n times the largest distance between audio capture positions of the audio capture environment, where n is between 2 and 5.
  • In particular, in many embodiments, the audio mixer may be arranged to switch between using a spherical wave propagation model when the distance is less than the threshold and using a planar wave propagation model when the distance is above the threshold.
  • Such an approach may reflect that for sources close to the array, the spherical wave model is appropriate as it may accurately reflect that the levels at different microphones depend on the position of the source relative to the microphones (as e.g. indicated by the previously described formulas).
  • It further reflects that for sources that are sufficiently far away, the spherical wave model can suitably be approximated by the simpler plane wave model where the level from to source on the different microphones is substantially the same but with a different phase due to the difference in the travel distance. This phase difference depends on the direction (angle) and the microphone distance.
  • For many uncorrelated sources far away (the opposite tribune in a stadium for example), the average level will be equal for all microphones and the average phase differences will be zero. Such a sound field is called a diffuse sound field. Except for the low frequencies (dependent on the ratio of the distance between the microphones and the wavelength) there is no correlation between the microphone signals for a diffuse field. As a result, if we weight the microphones signals with wi the mixed output signal will have the same "ambient" noise as long as i = 1 N mics w i 2 is constant, since p mix 2 = i = 1 N mics w i 2 p i 2 = p a 2 i = 1 N mics w i 2 , with p a 2 the ambient diffuse noise on each microphone.
  • In addition to the generation of the audio mix dependent on the selected object, the audio apparatus may also be arranged to adapt the processing dependent on a second object. The second object may be one that is in the far field, and specifically may be an object in a part of the scene for which the distance to the closest of the audio capture positions exceeds the largest distance between the audio capture positions by no less than ten times.
  • In addition to the mixing gains/weights being dependent on the selected object position, the gains/weights may also be dependent on the position of the second object.
  • For example, suppose we have a source S1 in the nearfield with S1 much closer to microphone i than to microphone j, such that p S 1 , i 2 p S 1 , j 2 . Suppose we have another source S2 in the far-field, that can be approximated at the microphones by a plane wave. Then the microphone signals for S2 have equal amplitudes and, possibly a delay w.r.t. each other. Suppose we can write p S 2 , i t = p S 2 , j t Δ Further, assume we want to suppress the contribution of S2 in the output. This can be done by delaying the signal of microphone j by Δ and subtract the microphone signal, i.e. apply the same weight but with a different sign.
  • We then can write for the output of the mixer: p mix = w 1 p s 1 , i + p s 2 , i w 1 p s 1 , j Δ + p s 2 , j Δ or p mix = w 1 p s 1 , i w 1 p s 1 , j Δ For the powers we get, using p S 1 , i 2 p S 1 , j 2 : p mix 2 = w 1 2 p S 1 , i 2
  • FIG. 3 is a block diagram illustrating an example processor 300 according to embodiments of the disclosure. Processor 300 may be used to implement one or more processors implementing an apparatus as previously described or elements thereof (including in particular the beamformers as described). Processor 300 may be any suitable processor type including, but not limited to, a microprocessor, a microcontroller, a Digital Signal Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA has been programmed to form a processor, a Graphical Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC) where the ASIC has been designed to form a processor, or a combination thereof.
  • The processor 300 may include one or more cores 302. The core 302 may include one or more Arithmetic Logic Units (ALU) 304. In some embodiments, the core 302 may include a Floating Point Logic Unit (FPLU) 306 and/or a Digital Signal Processing Unit (DSPU) 308 in addition to or instead of the ALU 304.
  • The processor 300 may include one or more registers 312 communicatively coupled to the core 302. The registers 312 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and/or any memory technology. In some embodiments the registers 312 may be implemented using static memory. The register may provide data, instructions and addresses to the core 302.
  • In some embodiments, processor 300 may include one or more levels of cache memory 310 communicatively coupled to the core 302. The cache memory 310 may provide computer-readable instructions to the core 302 for execution. The cache memory 310 may provide data for processing by the core 302. In some embodiments, the computer-readable instructions may have been provided to the cache memory 310 by a local memory, for example, local memory attached to the external bus 316. The cache memory 310 may be implemented with any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS) memory such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), and/or any other suitable memory technology.
  • The processor 300 may include a controller 314, which may control input to the processor 300 from other processors and/or components included in a system and/or outputs from the processor 300 to other processors and/or components included in the system. Controller 314 may control the data paths in the ALU 304, FPLU 306 and/or DSPU 308. Controller 314 may be implemented as one or more state machines, data paths and/or dedicated control logic. The gates of controller 314 may be implemented as standalone gates, FPGA, ASIC or any other suitable technology.
  • The registers 312 and the cache 310 may communicate with controller 314 and core 302 via internal connections 320A, 320B, 320C and 320D. Internal connections may be implemented as a bus, multiplexer, crossbar switch, and/or any other suitable connection technology.
  • Inputs and outputs for the processor 300 may be provided via a bus 316, which may include one or more conductive lines. The bus 316 may be communicatively coupled to one or more components of processor 300, for example the controller 314, cache 310, and/or register 312. The bus 316 may be coupled to one or more components of the system.
  • The bus 316 may be coupled to one or more external memories. The external memories may include Read Only Memory (ROM) 332. ROM 332 may be a masked ROM, Electronically Programmable Read Only Memory (EPROM) or any other suitable technology. The external memory may include Random Access Memory (RAM) 333. RAM 333 may be a static RAM, battery backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology. The external memory may include Electrically Erasable Programmable Read Only Memory (EEPROM) 335. The external memory may include Flash memory 334. The External memory may include a magnetic storage device such as disc 336. In some embodiments, the external memories may be included in a system.
  • It will be appreciated that the above description for clarity has described embodiments of the invention with reference to different functional circuits, units and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units or processors may be used without detracting from the invention. For example, functionality illustrated to be performed by separate processors or controllers may be performed by the same processor or controllers. Hence, references to specific functional units or circuits are only to be seen as references to suitable means for providing the described functionality rather than indicative of a strict logical or physical structure or organization.
  • The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and/or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
  • Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the accompanying claims. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in accordance with the invention. In the claims, the term comprising does not exclude the presence of other elements or steps.
  • Furthermore, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and/or advantageous. Also the inclusion of a feature in one category of claims does not imply a limitation to this category but rather indicates that the feature is equally applicable to other claim categories as appropriate. Furthermore, the order of features in the claims do not imply any specific order in which the features must be worked and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. In addition, singular references do not exclude a plurality. Thus, references to "a", "an", "first", "second" etc. do not preclude a plurality. Reference signs in the claims are provided merely as a clarifying example shall not be construed as limiting the scope of the claims in any way.

Claims (15)

  1. An apparatus for generating an audio mix signal, the apparatus comprising:
    a video signal receiver (201) arranged to receive video signals from at least one video camera capturing a scene, each video camera having a video capture pose;
    an audio signal receiver (203) arranged to receive audio signals from a plurality of microphones capturing audio of the scene, each microphone of the plurality of microphones having an audio capture pose;
    an object selector arranged (205) to select a first object of the scene by selecting an image object in at least one of the video signals;
    an object position determiner (207) arranged to determine a first object position for the first object in the scene from detections of the first object in images of the video signals and the video capture poses, the first object position being in a part of the scene for which a distance to a closest of the audio capture does not exceed five times a largest distance between the audio capture positions; and
    an audio mixer (209) arranged to generate the audio mix signal by mixing the audio signals, at least some mixing weights for at least some audio signals of the audio signals being determined as a function of the first object position relative to audio capture poses for the at least some audio signals.
  2. The apparatus of claim 1 wherein the audio mixer (209) is arranged to determine a mixing weight for a first audio signal of the plurality of audio signals depending on a distance from the first object position to an audio capture position for the first audio signal.
  3. The apparatus of claim 1 or 2 wherein the audio mixer (209) is arranged to determine a mixing weight for a first audio signal depending on a direction from the first object position to an audio capture position for the first audio signal and a directional sensitivity pattern for a microphone capturing the first audio signal.
  4. The apparatus of any previous claim wherein the audio mixer (209) is arranged to generate sound level attenuation estimates for acoustic propagation from the first object position to the audio capture poses, and to determine the at least some mixing weights in dependence on the acoustic energy loss estimates and a predetermined constraint on the weights.
  5. The apparatus of claim 4 wherein the constraint is a constraint on a combination of the weights.
  6. The apparatus of any of the previous claims wherein the video signal receiver (201) is arranged to determine at least one video capture pose from the video signals.
  7. The apparatus of any of the previous claims wherein the audio signal receiver (203) is arranged to determine at least one audio capture pose from at least one video capture pose and a predetermined relationship between the at least one audio capture pose and the at least one video capture pose.
  8. The apparatus of any previous claim further comprising a user interface (211) arranged to select the first object in response to a user input.
  9. The apparatus of claim 8 wherein the user interface (211) comprises a display for displaying an image of the scene generated from at least one of the video signals and a selection input for selecting the first object in response to the user selecting an image position on the image.
  10. The apparatus of any previous claim wherein the audio mixer (209) is arranged to apply a variable delay to at least one microphone audio signal prior to the audio mixing, the delay being dependent on a distance between the first object position and the audio capture pose for the at least one microphone audio signal.
  11. The apparatus of any previous claim wherein the audio mixer (209) is arranged to determine at least one gain for a first microphone signal in dependence on an observer position for an observer in the scene.
  12. The apparatus of any previous claim wherein the audio mixer (209) is arranged to determine the gains in dependence on estimated acoustic propagations from the object pose to the audio capture poses, the estimated acoustic propagations being determined based on an acoustic propagation model, the audio mixer being arranged to switch between a spherical wave acoustic propagation model for a distance between the first object position and the audio capture pose being below a threshold distance and between a plane wave acoustic propagation model for the distance between the first object position and the audio capture pose being above the threshold distance.
  13. The apparatus of any previous claim wherein at least one microphone of the plurality of microphones is a beamforming directional microphone array.
  14. The apparatus of any previous claim wherein the object position determiner (207) arranged to determine a second object position for a second object in the scene from detections of the second object in the video signals and the video capture poses, the second object being in a part of the scene for which a distance to a closest of the audio capture positions exceeds the largest distance between the audio capture positions by no less than ten times; and wherein
    the audio mixer (209) arranged is arranged to determine the at least some mixing weights as a function of the second object position relative to audio capture poses for the at least some microphone signals.
  15. A method of generating an audio mix signal, the method comprising:
    receiving video signals from at least one video camera capturing a scene, each video camera having a video capture pose;
    receiving audio signals from a plurality of microphones capturing audio of the scene, each microphone of the plurality of microphones having an audio capture pose;
    selecting a first object of the scene by selecting an image object in at least one of the video signals;
    determining a first object position for the first object in the scene from detections of the first object in images of the video signals and the video capture poses, the first object position being in a part of the scene for which a distance to a closest of the audio capture does not exceed five times a largest distance between the audio capture positions;
    generating the audio mix signal by mixing the audio signals, at least some mixing weights for at least some audio signals of the audio signals being determined as a function of the first object position relative to audio capture poses for the at least some audio signals.
EP24205998.8A 2024-10-10 2024-10-10 Audio apparatus and method of operation therefor Pending EP4727164A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP24205998.8A EP4727164A1 (en) 2024-10-10 2024-10-10 Audio apparatus and method of operation therefor
PCT/EP2025/077735 WO2026077738A1 (en) 2024-10-10 2025-09-29 Audio apparatus and method of operation therefor

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP24205998.8A EP4727164A1 (en) 2024-10-10 2024-10-10 Audio apparatus and method of operation therefor

Publications (1)

Publication Number Publication Date
EP4727164A1 true EP4727164A1 (en) 2026-04-15

Family

ID=93100326

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24205998.8A Pending EP4727164A1 (en) 2024-10-10 2024-10-10 Audio apparatus and method of operation therefor

Country Status (2)

Country Link
EP (1) EP4727164A1 (en)
WO (1) WO2026077738A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9749738B1 (en) * 2016-06-20 2017-08-29 Gopro, Inc. Synthesizing audio corresponding to a virtual microphone location
US20170366896A1 (en) * 2016-06-20 2017-12-21 Gopro, Inc. Associating Audio with Three-Dimensional Objects in Videos

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9749738B1 (en) * 2016-06-20 2017-08-29 Gopro, Inc. Synthesizing audio corresponding to a virtual microphone location
US20170366896A1 (en) * 2016-06-20 2017-12-21 Gopro, Inc. Associating Audio with Three-Dimensional Objects in Videos

Also Published As

Publication number Publication date
WO2026077738A1 (en) 2026-04-16

Similar Documents

Publication Publication Date Title
EP3721187B1 (en) An apparatus and method for processing volumetric audio
US10820097B2 (en) Method, systems and apparatus for determining audio representation(s) of one or more audio sources
US10034113B2 (en) Immersive audio rendering system
EP2800402B1 (en) Sound field analysis system
US9338544B2 (en) Determination, display, and adjustment of best sound source placement region relative to microphone
Farmani et al. Informed sound source localization using relative transfer functions for hearing aid applications
US20160119734A1 (en) Mixing Desk, Sound Signal Generator, Method and Computer Program for Providing a Sound Signal
US11012774B2 (en) Spatially biased sound pickup for binaural video recording
CN112735461B (en) Pickup method, and related device and equipment
Khan et al. Video-aided model-based source separation in real reverberant rooms
US10715914B2 (en) Signal processing apparatus, signal processing method, and storage medium
EP3275213B1 (en) Method and apparatus for driving an array of loudspeakers with drive signals
EP2362238B1 (en) Estimating the distance from a sensor to a sound source
EP4727164A1 (en) Audio apparatus and method of operation therefor
CN115396783B (en) Microphone array-based adaptive beam width audio acquisition method and device
US20220377456A1 (en) Determination of Sound Source Direction
CN113450769B (en) Voice extraction method, device, equipment and storage medium
JP2019054340A (en) Signal processing apparatus and control method thereof
EP4462769A1 (en) Generation of an audiovisual signal
RU2859877C2 (en) Method and system for controlling audio source directivity in virtual reality environment
US12621572B2 (en) Systems and methods for talker tracking and camera positioning in the presence of acoustic reflections
RU2793625C1 (en) Device, method or computer program for processing sound field representation in spatial transformation area
EP4443901A1 (en) Generation of an audio stereo signal
US20250030947A1 (en) Systems and methods for talker tracking and camera positioning in the presence of acoustic reflections
Palacino et al. Spatial sound pick-up with a low number of microphones

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN PUBLISHED

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR