WO2025190810A1 - Systems and methods for spatial fidelity improving dialogue estimation - Google Patents
Systems and methods for spatial fidelity improving dialogue estimationInfo
- Publication number
- WO2025190810A1 WO2025190810A1 PCT/EP2025/056307 EP2025056307W WO2025190810A1 WO 2025190810 A1 WO2025190810 A1 WO 2025190810A1 EP 2025056307 W EP2025056307 W EP 2025056307W WO 2025190810 A1 WO2025190810 A1 WO 2025190810A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- dialogue
- estimate
- signal
- input
- mix signal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
- G10L21/028—Voice signal separating using properties of sound source
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
Definitions
- Dialogue estimation is the process of extracting dialogue from an original signal where dialogue and non-dialogue sounds are mixed.
- DE exploits temporal, spectral and spatial properties of dialogue and non-dialogue sounds to extract dialogue from a mix.
- DE for monoaural audio exploits temporal and spectral properties of dialogue.
- NN neural network
- NN based systems for multi-channel DE are still in their infancy.
- a dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of ⁇ ⁇ 1 channels and/or ⁇ ⁇ 1 audio objects, wherein ⁇ + ⁇ ⁇ 2.
- the system comprises an input dialogue estimate processor (first dialogue estimate processor) configured to determine a first dialogue estimate of the input mix signal and a spatial analyzer configured to determine spatial parameters (e.g. panning parameters) based on the first dialogue estimate of the input mix signal.
- the system further comprises an adaptive downmixer configured to downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal and an output dialogue estimate processor (second dialogue estimate processor) configured to determine a second dialogue estimate of the downmixed signal.
- the system further comprises an adaptive upmixer configured to upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
- an adaptive upmixer configured to upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
- dialogue estimation refers to methods that extract a multi-channel dialogue estimate from a multi-channel input mix signal.
- the presented methods and systems may also operate on input mix signals comprising audio objects rather than audio channels, or input mix signals comprising a combination of audio objects and channels.
- element is used to mean either a channel or an object, and we use the terms “channel”, “object”, and “element” interchangeably in the following.
- the system comprising a first dialogue estimate processor configured to determine a first dialogue estimate of the input mix signal and a spatial analyzer configured to determine spatial parameters based on the first dialogue estimate of the input mix signal.
- the system further comprises an adaptive downmixer configured to downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal and an adaptive upmixer configured to upmix the downmixed signal based on the spatial parameters to generate an upmixed signal.
- the system further comprises a second dialogue estimate processor configured to determine a second dialogue estimate of the upmixed signal to generate an output dialogue signal. [013]
- the output dialogue estimate processor is applied to the upmixed signal.
- a dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of ⁇ ⁇ 1 channels and/or ⁇ ⁇ 1 audio objects, wherein ⁇ + ⁇ ⁇ 2.
- the method comprises determining a first dialogue estimate of the input mix signal and determining spatial parameters based on the first dialogue estimate of the input mix signal.
- the method further comprises downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal, determining a second dialogue estimate of the downmixed signal and upmixing the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
- a dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of ⁇ ⁇ 1 channels and/or ⁇ ⁇ 1 audio objects, wherein ⁇ + ⁇ ⁇ 2.
- the method comprising determining a first dialogue estimate of the input mix signal, determining spatial parameters based on the first dialogue estimate of the input mix signal and downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal.
- the method further comprises upmixing the downmixed signal based on the spatial parameters to generate an upmixed signal and determining a second dialogue estimate of the upmixed signal to generate an output dialogue signal.
- an apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to the third or fourth aspect of the disclosure.
- a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to the third or fourth aspect of the disclosure.
- a computer- readable storage medium storing the program according to the sixth aspect of the disclosure.
- Figure 1 is a block diagram illustrating a dialogue estimation system with a single dialogue estimation processor.
- Figure 2 is a block diagram illustrating an improved dialogue estimation system with a single dialogue estimation processor.
- Figure 3A is a block diagram illustrating a dialogue estimation system with two dialogue estimation processors, wherein one dialogue estimation processor is located upstream of the adaptive upmixer according to some implementations.
- Figure 3B is a flowchart describing a dialogue estimation method according to some implementations.
- Figure 4A is a block diagram illustrating an alternative dialogue estimation system with two dialogue estimation processors, wherein one dialogue estimation processor is located downstream of the adaptive upmixer according to some implementations.
- Figure 4B is a flowchart describing an alternative dialogue estimation method according to some implementations.
- Figure 5 is a block diagram illustrating a two-element dialogue estimation processor, according to some implementations.
- Figure 6 is a block diagram illustrating a multi-element dialogue estimation processor, according to some implementations.
- Figure 7 is a block diagram illustrating a two-element dialogue estimation processor with single gain mask computation, according to some implementations.
- Figure 8 is a block diagram illustrating a multi-element dialogue estimation processor with single gain mask computation, according to some implementations.
- Figure 9 is a block diagram illustrating a three-element dialogue estimation processor with dual gain mask computation, according to some implementations.
- Figure 10 is a block diagram illustrating another exemplary three-element dialogue estimation processor with dual gain mask computation.
- Figure 11 is a block diagram illustrating a first exemplary dialogue estimation system for processing three audio channels, according to some implementations.
- Figure 12 is a block diagram illustrating a second exemplary dialogue estimation system for processing three audio channels, according to some implementations.
- Figure 13 is a block diagram illustrating a third exemplary dialogue estimation system for processing three audio channels, according to some implementations.
- Figure 14 is a block diagram illustrating a fourth exemplary dialogue estimation system for processing three audio channels, according to some implementations.
- Figure 15 is a block diagram illustrating a first exemplary dialogue estimation system for processing two audio channels, according to some implementations.
- Figure 16 is a block diagram illustrating a second exemplary dialogue estimation system for processing two audio channels, according to some implementations.
- Figure 17 is a block diagram illustrating an exemplary dialogue estimation system for processing nine audio channels, according to some implementations.
- Figure 18 is a block diagram illustrating an exemplary dialogue estimation system for processing a mix of five audio channels and four audio elements, according to some implementations.
- Figure 19 is a block diagram illustrating a first exemplary dialogue estimation system with dialogue detection, according to some implementations.
- Figure 20 is a block diagram illustrating a second exemplary dialogue estimation system with dialogue detection, according to some implementations.
- Figure 21 is a block diagram illustrating a third exemplary dialogue estimation system with dialogue detection, according to some implementations.
- Figure 22 is a block diagram illustrating a fourth exemplary dialogue estimation system with dialogue detection, according to some implementations.
- Figure 23 is a block diagram illustrating a fifth exemplary dialogue estimation system with dialogue detection, according to some implementations.
- Figure 24 is a block diagram illustrating a first exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations.
- Figure 25 is a block diagram illustrating a second exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations.
- Figure 26 is a block diagram illustrating a third exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations.
- Figure 27 is a block diagram illustrating a coding system for encoding the dialogue estimate and the mix input signal, according to some implementations.
- FIG. 28 is a block diagram illustrating another coding system for encoding the dialogue estimate and an extracted non-dialogue signal according to some implementations.
- Figure 29 is a block diagram illustrating an apparatus for performing or implementing methods and techniques described throughout the present disclosure.
- DETAILED DESCRIPTION OF CURRENTLY PREFERRED EMBODIMENTS [052]
- FIG.1 is a block diagram showing a dialogue estimation system 10A bearing some resemblance to the solutions presented in the international patent application with publication number WO2021/252823 and in International Application No. PCT/US23/63717, each of which are hereby incorporated by reference in its entirety.
- the dialogue estimation system 10A is sometimes referred to as a “DeepSpace” system.
- the dialogue estimation system 10A obtains an input stereo audio signal SIN comprising two channels wherein each of the channels carries a mix of speech content and non-speech content.
- the input stereo audio signal S IN is provided to a spatial analyzer 1 which extracts spatial parameters Pspatial (e.g. panning parameters, inter-channel time difference, inter-channel level difference and magnitude) of a target audio source (e.g. dialogue) present in the input stereo audio signal S IN.
- the spatial parameters are adaptive and may vary over time and may vary for different frequency bands.
- the spatial parameters are then provided to an adaptive downmixer 2 which forms a downmix signal SD based on the adaptive spatial parameters and the input stereo audio signal S IN .
- the downmix signal S D may be a two-channel signal (e.g.
- the downmix signal S D is provided to a spatial filtering module 3 which performs spatial filtering.
- the spatial filtering module 3 is optional and may be omitted or located at other positions in the processing chain.
- the spatial filtering may involve setting the adaptive side channel of the downmix signal to zero.
- the spatial filtering is equivalent to the Spatio-Level Filtering (SLF) presented in A. S. Master, et. al "Dialog Enhancement via Spatio-Level Filtering and Classification," in 149 th Audio Engineering Society Convention, New York, 2020, which is hereby incorporated by reference in its entirety.
- SPF Spatio-Level Filtering
- the output spatially filtered downmix signal S DF is provided to an adaptive upmixer 4 which performs adaptive upmixing using the spatial parameters Pspatial. If the spatial filtering module 3 is omitted, the downmix signal SD is provided directly to the adaptive upmixer 4.
- the adaptive upmixer 4 performs upmixing, based on the spatial parameters P spatial to obtain a stereo output signal S OUT .
- the adaptive upmixer 4 may e.g. convert a downmix mid channel and side channel to pair of stereo channels or upmix a single mono channel to pair of stereo channels, based on the spatial parameters to retain e.g. the panning and/or inter-channel phase difference of the stereo input signal S IN .
- the stereo output signal S OUT is subsequently provided to a dialogue estimation processor 5.
- the dialogue estimation processor 5 may be a low spatial fidelity (LOSPAFI) dialogue estimator, as described below.
- LOSPAFI low spatial fidelity
- dialogue is used to refer to any type of speech in general. While the term “dialogue” in some contexts is interpreted to mean a conversation between two or more persons the term “dialogue” is in the present disclosure used to refer to any type of human speech sounds, regardless of the speech sounds being uttered by a single person or multiple persons.
- the dialogue estimate processor 5 performs dialogue separation (e.g.
- the output of the dialogue estimate processor 5 is accordingly a dialogue estimate stereo output signal S OUTDE .
- a weakness of the dialogue estimation system 10A in FIG.1 is that the spatial analyzer 1 extracts the spatial parameters based on the input stereo signal SIN, and the input stereo signal SIN may be degraded by the presence of the interfering non-dialogue sounds (e.g. noise, music and/or effects). This degradation may manifest itself as spatial instabilities in the stereo output signal SOUT and also in the output dialogue estimate SOUTDE.
- FIG. 2 An improvement to the dialogue estimation system 10A in FIG.1 is shown in FIG. 2.
- the dialogue estimation system 10B of FIG.2 is similar to the system of FIG.1, however a difference is that the dialogue estimate processor 5 has been moved upstream such that the dialogue estimation is performed upstream of the adaptive downmixer 2, upstream of the spatial filtering module 3 and upstream of the adaptive upmixer 4.
- the spatial analyzer 1 may now determine the spatial parameters Pspatial based on the dialogue estimate stereo input signal SDE, instead of the input stereo signal S IN , meaning that the spatial analysis becomes more robust and generally achieves improved dialogue estimation even when the stereo input signal SIN contains a large portion of non-dialogue audio content. Additionally, since both the input stereo signal SIN and dialogue estimate stereo input signal S DE are available it is possible to extract the spatial parameters P spatial based on a combination of both the input stereo signal S IN and the dialogue estimate stereo input signal SDE. [062]
- the dialogue estimate stereo input signal S DE is provided to the adaptive downmixer 2 which downmixes the stereo speech estimate S DE based on the spatial parameters P spatial to form a downmix dialogue estimate SDED.
- the downmix dialogue estimate SDED is then optionally subject to spatial filtering using the spatial filtering module 3 forming a filtered downmix dialogue estimate S DEDF .
- the filtered downmix dialogue estimate S DEDF is provided to the adaptive upmixer 4 which upmixes the filtered downmix dialogue based on the spatial parameters Pspatial estimate SDEDF to form the dialogue estimate stereo output signal SOUTDE.
- the adaptive downmixer 2 in FIG.2 is operating in a suboptimal way (the nature of the suboptimality depends on how the dialogue estimate processor 5 is configured).
- the dialogue estimation system 10B additionally allows more freedom in what signals the spatial analysis is based upon since the stereo input signal S IN and the dialogue estimate stereo input signal S DE are both available and may be used in combination.
- FIG.3A shows a dialogue estimation system 10C according to some implementations of the present disclosure.
- at least two dialogue estimate processors 5a, 5b are used in the processing chain, an input dialogue estimate processor 5a (or “first” dialogue estimate processor) and an output dialogue estimate processor 5b (or “second” dialogue estimate processor).
- the input dialogue estimate processor 5a makes the spatial analysis, and the resulting spatial parameters Pspatial, more robust to non- dialogue interference in the input mix signal.
- the input dialogue estimate processor 5a does not affect the main signal path (the main signal path comprising the adaptive downmixer 2, the spatial filtering module 3 and the adaptive upmixer 4) and the adaptive downmixer 2 is allowed to operate on the input mix signal MIN directly, which is beneficial as explained below.
- a second dialogue estimate processor 5b is introduced in the main signal path downstream of the adaptive downmixer 2. [066] In general, however, the second dialogue estimate processor 5b may be located at any point downstream of the adaptive downmixer 2 and in FIG.4 an alternative dialogue estimation system 10D is shown where the second dialogue estimate processor 5b is located downstream of the adaptive upmixer 4, as described below.
- an input mix signal MIN is obtained.
- the input mix signal comprises at least two audio elements wherein each audio element is either an audio object or a channel as described below.
- a first dialogue estimate DE1 is extracted from the input mix signal MIN by the first (input) dialogue estimate processor 5a.
- the method then goes to step S2A involving determining spatial parameters based on the first dialogue estimate DE1. This is achieved using the spatial analyzer 1.
- the method then goes to step S3A comprising downmixing the input mix signal MIN based on the input mix signal MIN and the spatial parameters ⁇ ⁇ to generate the downmixed signal M D .
- the downmixing is performed by the adaptive downmixer 2 which is controlled by the spatial parameters ⁇ ⁇ .
- the method then goes to step S4A comprising determining a second dialogue estimate DE2D based on the downmixed signal MD assuming no spatial filtering is done.
- the second dialogue estimate DE2D is extracted by the second (output) dialogue estimate processor 5b.
- the second dialogue estimate DE2 D is a downmixed representation and at step S5a the second dialogue estimate DE2D is upmixed by adaptive upmixer 4 to yield an output dialogue estimate DE2 as an upmixed version of DE2D.
- the adaptive upmixer 4 is controlled by the spatial parameters.
- a dialogue estimation system 10C is in some implementations used together with a dialogue detector (see FIGS.19-23) and/or with a mixer for mixing the input mix signal back to the dialogue estimate DE (see FIGS.24-26). Accordingly, the method may optionally comprise step S6 comprising applying a soft gate gain to the dialogue estimate DE, wherein the soft gate gain is based on a dialogue confidence value extracted by a dialogue detector. Additionally or alternatively, the method may comprise optional step S7 comprising mixing the dialogue estimate DE with the input mix signal M IN to obtain a final dialogue estimate as described below.
- Steps S1B and S2B are generally similar to steps S1A and S1B described above. That is, at step S1B a first dialogue estimate DE1 is extracted from the input mix signal MIN by the first (input) dialogue estimate processor 5a. The method then goes to step S2B involving determining spatial parameters ⁇ ⁇ based on the first dialogue estimate DE1.
- the spatial parameters ⁇ ⁇ may e.g. comprise panning parameters ⁇ ⁇ as described below, or other types of spatial parameters.
- the spatial parameters ⁇ ⁇ are determined using the spatial analyzer 1.
- step S3B comprising downmixing the input mix signal M IN based on the input mix signal MIN and the spatial parameters ⁇ ⁇ to generate the downmixed signal M D .
- the downmixing is performed by the adaptive downmixer 2 which is controlled by the spatial parameters ⁇ ⁇ .
- step S4B the downmixed input mix signal M D is upmixed using adaptive upmixer 4 and at step S5B the second (output) dialogue estimate processor 5b extracts the second dialogue estimate DE2 based on the upmixed signal M UF (compare to step S4A where the second dialogue estimate DE2D is extracted from the downmixed signal MD).
- the second dialogue estimate DE2 is used as the dialogue estimate DE output of the dialogue estimate system 10C.
- the dialogue estimator systems 10C, 10D operate on an input mix signal MIN and outputs a dialogue estimate DE.
- the input mix signal MIN is not limited to a stereo audio signal comprising two stereo channels and may for the purposes of this disclosure be any type of audio signal comprising (A) at least two channels, (B) at least two audio objects or (C) at least one audio channel and at least one audio object.
- the term audio “element” will be used to refer generally to an audio channel, or an audio object whereby input mix signal M IN and the dialogue estimate DE may be said to comprise at least two audio elements.
- An audio channel may be an audio channel associated with a specific loudspeaker of an audio presentation format.
- An audio object is an audio signal which has an associated spatial position which may vary with time.
- An object based signal may comprise one or more audio objects wherein the spatial position of each object may vary with time as signaled with metadata.
- the input mix signal M IN comprises N number of channels and/or M number of audio objects wherein ⁇ + ⁇ ⁇ 2. Accordingly, the total number of audio elements are at least two.
- the number and type of elements i.e. audio objects and/or channels) define the spatial configuration of the input mix signal M IN .
- the dialogue estimate signal DE also consists of N spatial channels and/or M audio objects corresponding to the number N of channels and/or number M of audio objects of the input mix signal M IN .
- the dialogue estimate systems described herein may be used to process any type of multi-channel (multi-element) audio signal wherein the output dialogue estimate DE2 has corresponding audio elements to the input mix signal MIN.
- the dialogue estimation systems envisaged in the present disclosure need not operate on all elements of a spatial configuration, i.e. the input mix signal MIN is not necessarily all elements of an original mix signal.
- dialogue estimation systems may operate on a predetermined subset elements of a spatial configuration whereby the remaining elements (referred to as residual elements) remain unprocessed and are combined with the processed elements to yield a hybrid output dialogue estimate spatial configuration with at least two processed elements and one or more unprocessed elements.
- residual elements the mix input signal M IN is only the LRC triplet whereas the surround channels and the height channels (as well as the LFE channel) remain unprocessed.
- the processing is typically done in time-frequency (TF) tiles with a suitable time- frequency resolution, although broadband (but still time varying) processing may be used in some applications. There are plenty of options when it comes to suitable TF transforms.
- ST-DFT short-term discrete Fourier transform
- a concrete example is using a ST- DFT analysis block size of 4096 samples at a sampling frequency of 48000 Hz.
- Each dialogue estimate processor 5a, 5b may implement any type of processing that processes a single- or multi-element mix signal (e.g.
- the dialogue estimate processors 5a, 5b envisaged are of a “low spatial fidelity” (abbreviated LOSPAFI) which indicates that the spatial fidelity of the output of each estimator may leave room for improvement, and such improvement may be achieved by implementing the proposed systems 10C, 10D of FIGS.3A and 4A. Below, multiple concrete examples of dialogue estimation systems 10C, 10D and dialogue estimate processors 5a, 5b are presented.
- LOSPAFI low spatial fidelity
- All exemplary dialogue estimation systems 10C, 10D and dialogue estimate processors 5a, 5b use single-element dialogue estimators or single-element mask computation modules as the basic building block but the single-element processed by each dialogue estimator or mask computation module varies and may be either an audio channel or an audio object, or a mix, e.g., the sum, of certain audio channels and/or audio objects.
- any multi-channel dialogue estimate processor or multi-element dialogue estimate system
- single-channel dialogue estimate processors i.e., single-element dialogue estimators
- single-element dialogue estimators may be categorized into filtering/mask based estimators and direct mapped estimators.
- a time and frequency varying filter/mask is computed (e.g., via a mask generator) and then applied to the signal input to the processor.
- direct-mapped estimators there is no explicit filter or mask, and the output of the processor is computed as a direct mapping from the signal input to the processor.
- Both categories may advantageously employ neural networks (NN).
- NN neural networks
- An example of a mask based NN DE system is LensNet. LensNet is described in U.S. Patent Publication No.20230368807 titled “Deep learning based speech enhancement”, hereby incorporated by reference in its entirety.
- the LensNet model takes as input a mono audio signal and predicts a gain mask for estimating the dialogue of the input mono signal.
- Examples of direct mapped systems include so called generative NNs, for example UNIVERSE. UNIVERSE is described in PCT Publication No. WO 2023/052523 A1 titled “Universal speech enhancement using generative neural networks”, hereby incorporated by reference in its entirety.
- the UNIVERSE model includes a neural network based system for speech (dialogue) enhancement, consisting of a generative network for generating an enhanced audio signal and a conditioning network for generating conditioning information for the generative network.
- Masks and/or filters may in some cases be derived from direct mapped systems.
- Obtaining a gain mask from a direct-mapped dialogue estimator may be accomplished by generating the (mono) direct mapped dialogue estimate, and deriving a mono gain mask based on the mono input and mono output of the direct mapped system, e.g., by letting the mask be the ratio of the energy of the output to the energy of the input (done in appropriate TF tiles to obtain a frequency selective mask).
- This mono mask may then be used analogously to a gain mask obtained with a e.g. filter/mask based dialogue estimator.
- the feasibility of deriving masks in a direct mapped system depends on the nature of the dialogue estimation processing.
- the output of NN based generative dialogue estimators may deviate so much from the dialogue in the input signal that the derived mask “misses” the original dialogue, yielding distorted dialogue with high levels of residual background.
- some dialogue estimators presented below rely on a mask being computed. However it is understood that a mask may also be derived from directed mapped dialogue estimators meaning that the same general processing may be applied with masks derived from direct mapped dialogue estimators.
- the mask computation module 52 implemented in FIGS.5-10 below may accordingly generate a mask using any of the methods described, including both a masked-based system and a direct-mapped system. Other types of dialogue estimation are also within the scope of this disclosure.
- FIGS.5-10 show various examples of first (input) dialogue estimate processors 5a.
- the second (output) dialogue estimate processors 5b may be similar or even identical to the first dialogue estimate processor 5a.
- the second dialogue estimate processor 5b may be identical to the first dialogue estimate processor 5a.
- the adaptive dowmixer 2 downmixes from e.g.
- FIG.5 is a block diagram illustrating a two element dialogue estimate processor 501 using two single-element dialogue estimators 51:1, 51:2.
- the dialogue estimate processor 501 is configured for processing a stereo channel input mix signal MIN and output a stereo channel first dialogue estimate DE1.
- the dialogue estimation processing is here done independently in the left channel and in the right channel by using two independent instances of a single-channel dialogue estimator 51:1, 51:2 (i.e., two instances of a single-element dialogue estimator).
- Each dialogue estimator 51:1, 51:2 may be a direct-mapped dialogue estimator or mask-based dialogue estimator.
- Weaknesses of this type of dialogue estimation processing include suboptimal dialogue-to-non-dialogue separation when the dialogue of the input mix signal MIN is panned near center. Another weakness is that the estimation errors in the respective single-channel estimators 51:1, 51:2 is independent.
- Each single channel estimator 51:1, 51:2 may comprise a gain mask calculator and gain mask applicator.
- the dialogue estimate processor 501 of FIG.5 may also operate in the same way on an input mix signal M IN comprising two audio objects or on an input mix signal M IN comprising one audio object and one audio channel.
- each audio element of the input mix signal is processed with a respective single-channel dialogue estimator 51:1, 51:2.
- the K-element dialogue estimate processor 502 using K-single-element dialogue estimators 51:1, 51:2, ..., 51:K is shown.
- the dialogue estimate processor 502 comprises K instances of a single-element dialogue estimator 51:1, 51:2, ... 51:K and each single-element dialogue estimator 51:1, 51:2, ... 51:K obtains as an input a respective input channel/object of the input mix signal M IN and outputs an associated dialogue estimate channel/object.
- FIG.7 a block diagram illustrating a first dialogue estimate processor 511 according to some implementation is shown.
- the first dialogue estimate processor 511 of FIG.7 is configured to obtain an input mix signal M IN comprising two elements (i.e.
- the dialogue estimate processor 511 accomplishes this using a single-element mask computation module 52.
- the first dialogue estimate processor 511 downmixes the two channels of the input mix signal MIN to a mono signal MONO1 which is provided to a single-channel (single element) mask computation module 52.
- the single element mask computation module 52 generates a gain mask for removing, or at least attenuating, non-dialogue content in the mono signal MONO 1 .
- the gain mask is subsequently provided to two respective mask application modules 56a, 56b which apply the (same) computed gain mask to each of the two channels of the input mix signal M IN so as to generate two dialogue estimate channels that form the first dialogue estimate signal DE1.
- the first dialogue estimate processor 512 is configured to obtain an input mix signal MIN comprising N channels (or equivalently e.g., M audio objects, or N channels and M objects). Each of the N channels are provided to a multi-channel downmixer 53 which downmixes the N channels to a single mono signal MONO 1 which in turn is provided to a single- channel (single element) mask computation module 52.
- the single-channel mask computation module 52 generates a gain mask for removing or at least attenuating non-dialogue content in the mono signal MONO 1 .
- the gain mask is subsequently provided to a multi-channel mask application module 54 which applies the (same) computed gain mask to each of the N channels of the input mix signal MIN to obtain corresponding dialogue estimate channels that form the first dialogue estimate signal DE1.
- Benefits of the first dialogue estimate processors 511, 512 of FIG.7 and FIG.8, compared to the channel (or element) independent dialogue estimate processors 501, 502 of FIG. 5 and FIG.6, include lower computational complexity, assuming the downmixing and mask application is of much lower complexity than the mask computation, and better spatial stability.
- FIG.9 Another first dialogue estimate processor 513 is depicted. This dialogue estimate processor 513 obtains three input elements and extracts a dialogue estimate of each element using only two gain mask computation modules 52 extracting a respective gain mask.
- one of the input channels is split and input into two instances 511a, 511b of the dialogue estimate processor 511 from FIG.7. While each of the dialogue estimate processors instances 511a, 511b takes two channels as an input, and outputs two channels, each dialogue estimate processor instance 511a, 511b utilizes a respective single- element gain mask computation module 52.
- first dialogue estimate processor 513 of FIG.9 operates on a left, right and center (LRC) channel triplet.
- the left channel and the center channel are input to a first instance 511a of a two channel estimator which mixes the left and center channel to a mono signal MONO 1 , computes a first gain mask for dialogue estimation for the mono signal MONO1 with a gain mask computation module 52 and applies the first gain mask to each of the left channel and the center channel with respective gain applicators 56a, 56b.
- the center channel is split and also input to the second instance 511b of the two channel dialogue estimation process alongside the right channel.
- the second instance 511b of the two channel estimator mixes the center and right channel to a second mono signal MONO 2 , computes a second gain mask for dialogue estimation for the mono signal MONO2 with a gain mask computation module 52 and applies the second gain mask to each of the left channel and the center channel with respective gain applicators 56a, 56b.
- the output from the first and second instance 511a, 511b of the two channel dialogue estimators comprises two instances of a dialogue estimate center channel, a dialogue estimate left channel and a dialogue estimate right channel.
- the dialogue estimate left and right channels are output by the first dialogue estimate processor 513.
- the two instances of the dialogue estimate center channel are provided to a mixer 57 which combines the two dialogue estimate center channels to a combined center channel and, to avoid an unnatural boost in signal energy, the combined dialogue estimate center channel is provided to a gain modification module 58 with a gain of 1 ⁇ 2 which reduces the signal energy by a factor 4 (i.e. attenuates the signal energy with 6 dB).
- the gain modified dialogue estimate center channel is then output alongside the dialogue estimate left channel and right channel to form the first dialogue estimate DE1.
- FIG.10 shows another first dialogue estimate processor 521 which obtains as an input an input mix signal MIN comprising a left channel, a right channel and a center channel.
- the first dialogue estimate processor 521 comprises a two-channel dialogue estimate processor estimator, similar to the dialogue estimate processor 511 estimator from FIG.7, which forms dialogue estimates of the left and right channel using a mixer 55, a gain mask computation module 52 and gain mask applicators 56a, 56b.
- the center channel is processed by a single- channel dialogue estimator 51. While the first dialogue estimate processor 521 of FIG.10 may be used for any triplet of channels (or elements) it is especially useful for the LRC front triplet of a 5.1 presentation.
- the center channel may be processed individually.
- the various first dialogue estimate processors 501, 502, 511, 512, 513, 521 described above are exemplary and that other versions are possible.
- the first dialogue estimate processors may be combined with each other and/or modified to suit the number and type of elements present in the input mix signal MIN.
- the first dialogue estimate processor 5a is used as pre-processing before the spatial analyzer 1, and as such the first dialogue estimate processor 5a maps the N-channel input mix signal M IN to an N- channel first dialogue estimate DE1 (or in general an ⁇ + ⁇ -audio element input mix to an ⁇ + ⁇ -audio element first dialogue estimate DE1).
- the second (output) dialogue estimate processor 5b is a single-channel (or element) system if it is performed upstream of the adaptive upmixer 4 (see FIG.3A).
- FIGS.11-18 illustrate examples of how the proposed dialogue estimation systems 10C, 10D can make use of the different dialogue estimate processor architectures 501, 502, 511, 512, 513, 521 presented in connection with FIGS.5-10.
- the exemplary systems of FIGS.11-17 all operate on channels, but it is understood that one or more (or all) channels may be replaced with audio objects.
- all exemplary dialogue estimation systems 101-108 of FIGS.11-18 are of generally the same type as dialogue estimation system 10C from FIG.3A, wherein the second (output) dialogue estimate processor 5b is located upstream from the adaptive upmixer 4. However, it is understood that the dialogue estimation systems 101-108 may also be used with the second (output) dialogue estimate processor 5b located downstream of the adaptive upmixer 4.
- FIG.11 shows an exemplary dialogue estimation system 101 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel.
- the channels are provided to a first dialogue estimate processor 502 having three single-channel dialogue estimation modules 51:1, 51:2, 51:3 that extract an individual dialogue estimate of each of the three input channels.
- the dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1.
- the spatial analyzer 1 extracts spatial parameters P spatial which are provided to the optional spatial filtering module 3.
- FIG.12 shows another exemplary dialogue estimation system 102 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel.
- the channels are provided to a first dialogue estimate processor 521 having a two-channel dialogue estimate processor 511 (see FIG.7) for processing the left and right channel and a single-element dialogue estimate processor 51 for processing the center channel.
- dialogue estimate processor 521 operates as the general dialogue estimate processor from FIG.6.
- the dialogue estimate channels output by the dialogue estimate processor 521 form the first dialogue estimate DE1 which is provided to the spatial analyzer 1 for extraction of spatial parameters P spatial .
- the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11.
- FIG.13 shows yet another exemplary dialogue estimation system 103 obtaining an input mix signal M IN with three channels, a left channel, a right channel and a center channel.
- the channels are provided to a first dialogue estimate processor 513 having two instances 511a, 511b of two-channel dialogue estimate processors for processing the left and center channel and processing the right and center channel respectively.
- dialogue estimate processor 513 operates as the dialogue estimate processor from FIG.9.
- the dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1 tasked with extracting the spatial parameters Pspatial.
- the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11.
- FIG.14 shows yet another exemplary dialogue estimation system 104 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel.
- the channels are provided to a first dialogue estimate processor 512 having a multi-channel downmixer and a single-element gain mask calculator 52.
- the first dialogue estimate processor 512 operates as the dialogue estimate processor from FIG.8 and downmixes all channels to a mono signal MONO1 for which a gain mask is calculated by gain mask calculator 52 whereafter the same gain mask is applied to each of the three channels by the mask applicator 54 yielding the corresponding dialogue estimate channels.
- the dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1.
- FIG.15 shows an exemplary dialogue estimation system 105 obtaining an input mix signal M IN with two channels, a left channel and a right channel.
- the two channels are provided to a first dialogue estimate processor 501 having two single-element dialogue estimators 51:1, 51:2 that form a dialogue estimate for each channel and outputs these dialogue estimate channels as the first dialogue estimate DE1.
- the first dialogue estimate DE1 is provided to the spatial analyzer 1 which extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11.
- FIG.16 shows another exemplary dialogue estimation system 106 obtaining an input mix signal MIN with two channels, a left channel and a right channel.
- the two channels are provided to a two-channel dialogue estimate processor 511 identical to the dialogue estimate processor described in connection with FIG.7.
- the two channels are combined into a mono signal MONO 1 and provided to a gain mask calculation module 52 which calculates a gain mask which is applied to each of the two channels, to yield dialogue estimate channels.
- the dialogue estimate channels form first dialogue estimate DE1 which is provided to the spatial analyzer 1.
- the spatial analyzer extracts spatial parameters P spatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11.
- FIG.17 shows an exemplary speech estimation system 107 obtaining an input mix signal M IN with nine channels.
- the nine channels are in this implementation the nine non-LFE channels of a 5.1.4 presentation. That is, the nine channels comprising the left channel (L), the right channel (R), the center channel (C), the left surround channel (Ls) and the right surround channel (Rs) of the horizonal listening plane. The nine channels further comprising the top front left (Tfl) channel, top front right (Tfr) channel, top back left (Tbl) channel and top back right (Tbr) channel forming the four height channels of the 5.1.4 presentation. All nine channels are provided to dialogue estimate processor 514 implementing three instances of multi-channel dialogue estimate processors 512a, 512b, 512c, each being equivalent to the multi-channel dialogue estimate processor of FIG.8 and receiving a unique subset of the nine channels.
- the first multi-channel dialogue estimate processor 512a obtains the four height channels, mixes them to a same mono signal MONO 1 , generates a gain mask for the mono signal MONO1 and applies the same gain mask to all four height channels to obtain the associated dialogue estimate channels.
- the second multi-channel dialogue estimate processor 512b obtains the two surround channels (Ls, Rs), mixes them to a mono signal MONO 2 , generates a gain mask for the mono signal MONO2 and applies the same gain mask to each surround channel to obtain the associated dialogue estimate channels.
- the third multi-channel dialogue estimate processor 512c obtains the LRC-triplet channels, mixes them to a mono channel MONO3, generates a gain mask for the mono signal MONO3 and applies the same gain mask to each of the L, R and C channels to obtain the associated dialogue estimate channels.
- the dialogue estimate channels together form the first dialogue estimate DE1 which is provided to the spatial analyzer 1.
- the spatial analyzer 1 extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [127]
- FIG.18 shows an exemplary dialogue estimation system 108 obtaining an input mix signal MIN with five channels (in this case the L, R, C, Rs and Ls channels of a 5.1 presentation) and four audio objects (objects 1 - 4).
- the dialogue estimate processor 514 is however the same dialogue estimator as in speech estimate system 107, with the only difference being that instead of the first multi-channel dialogue estimate processor 512a obtaining the (four) height channels it now obtains the (four) audio objects and extracts the dialogue estimate of each audio object using a same gain mask for each object wherein the gain mask is extracted from a mono downmix of all four audio objects.
- the dialogue estimate channels and objects forming the first dialogue estimate DE1 are provided to the spatial analyzer 1 which extracts spatial parameters P spatial , and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11.
- the first dialogue estimate DE1 is to be mapped to spatial parameters P spatial which are used to control the adaptive downmixing, adaptive upmixing and optionally the spatial filtering.
- P spatial spatial parameters
- inter-channel phase difference such as one or more of inter-channel phase difference, inter-channel time difference, inter-channel level difference and magnitude
- the adaptive downmixing, upmixing and spatial filtering described in International Application No. PCT/US23/63717 uses the inter-channel phase difference, a panning parameter and the magnitude as well as indicators of the average value and spread of the panning and inter-channel phase difference labelled thetaMid, thetaWidth, phiMid and phiWidth.
- the working example will be an application where the input mix signal MIN is the LRC triplet of a 5.1 presentation.
- the presented methods generalize straightforwardly to signals with a higher channel count, higher object count and signals that are a combination of channels and objects. Both methods yield a set of panning parameters in the form of one scalar value per channel/object, and these may be used in the adaptive downmixing and upmixing.
- the processing described below is done on a group of frequency indices that are grouped together in a processing band forming a TF tile.
- a static mono downmix M is calculated for the first dialogue estimate DE1 as for all ⁇ in a current processing band.
- a panning parameter ⁇ ⁇ may now be calculated for each of the N channels (or elements) as and the panning parameter ⁇ ⁇ is, as mentioned above, one example of a spatial parameter P spatial .
- the input mix signal M IN labeled Y
- Energy preserving panning means that
- a benefit with correlation based spatial analysis is that it is robust against “energy leakage” between elements of the dialogue estimate channels and/or objects X.
- the input mix signal MIN labeled Y
- Y X' + W
- X' uncorrelated with the non-dialogue component W
- ⁇ the panning coefficient for element n and that the panning is energy preserving
- the adaptive downmix signal ⁇ ⁇ ( ⁇ ) may be computed as where ⁇ ⁇ is the adaptive downmix coefficient for element (a channel or an object) n.
- Each of the proposed dialogue estimation systems 10C-D, 101-108 constitutes a self-contained audio processing system which obtains a multi-element input M IN and outputs a dialogue estimate DE2 having the corresponding elements with dialogue extracted.
- FIG.19 is a block diagram illustrating a dialogue estimation system 201 provided with a dialogue detector 6.
- the dialogue detector 6 produces a sequence of dialogue confidence values, typically one value per processing frame (which typically is 20-100 ms long).
- the confidence values are then post-processed by post processor 7, and a temporal soft gate gain is computed by soft gate computation module 8.
- the soft gate gain is then applied by soft gate applicator 9 to the second dialogue estimate DE2 to form a gated second dialogue estimate DE2G.
- the soft gate gain consists of one gain per audio element (and per processing frame), the gains vary over time as a function of the output from the dialogue detector 6.
- the dialogue detector 6 may produce a binary valued output (e.g.0 or 1) or a continuous valued output, typically in the interval [0,1], indicating the probability of dialogue being present in the current time segment.
- the output may also be referred to as the dialogue confidence.
- the post-processing performed by the post processor 7 may comprise smoothing of the dialogue confidence so as to e.g. bridge “gaps” (short segments of confidence 0 or very low confidence surrounded by segments of high confidence) or remove “spikes” (short segments of high confidence surrounded by low confidence).
- the soft gate computation performed by the soft gate computation module 8 may consist of a non-decreasing mapping of the post processed dialogue confidence to a soft gate gain.
- the input to the dialogue detector may be the signal from the dialogue estimation system 10C-D, 101-108 at any point of the main signal chain.
- the block diagram in FIG.19 illustrates the flexibility in choosing from where to get the input to the dialogue detector 6 (see the dashed lines).
- the input mix signal MIN, the adaptive downmix signal MD, MDF before or after spatial filtering, the downmixed dialogue estimate DE2 D or the second dialogue estimate DE2 after upmixing may be provided as input to the dialogue detector 6.
- the soft gate applicator 9 may be done at any point of the main signal chain (although it advantageously occurs after the input to the dialogue detector).
- the soft gate applicator 9 is arranged downstream of the adaptive upmixer 4 and an alternative dialogue estimate system 202 is shown in FIG.20, where the soft gate applicator 9 is arranged upstream of the adaptive upmixer 4, downstream of the second dialogue estimator 5b.
- Applying the soft gate to a mono signal may save complexity compared to applying the gate to a multi-channel/object signal (as in FIG.19).
- the dialogue detector 6 may handle multi-channel (or multi-element) input in a way similar to the dialogue estimate processors 5a, 5b.
- the dialogue detector 6 may rely on downmixing and/or separate processing of the individual channels/objects.
- the dialogue detector 6 is operating on a mono signal, and the processing performed by the post processor 7 and the soft gate computation module 8 results in a single time-varying soft gate gain. This same soft gate gain is applied to all the channels/objects by the gate applicator 9.
- the mono signal input to the dialogue detector 6 may be computed as a downmix of the multi- channel/object signal from some place in the main signal path.
- the channels/objects may e.g. be obtained upstream of the adaptive downmixer 2.
- input mix signal MIN is the non-LFE channels of a 5.1 signal
- Ls and Rs of the 5.1 input mix signal MIN it is understood that dialogue in the surround channels Ls and Rs of the 5.1 input mix signal MIN is uncommon but occurs sometimes.
- the attenuation gain with which Ls and Rs are attenuated may be a predetermined attenuation gain such as 6 dB.
- this attenuation gain is merely exemplary, and other predetermined attenuation gains may be used.
- Yet another option is to configure post processor 7 to output one unique dialogue confidence value per audio element, and to configure soft gate computation module 8 to output one unique gain per element.
- the soft gate application module 9 may then apply an individual soft gate gain to each element.
- Yet another option is to combine the single-element dialogue detection with the downmix based dialogue detection discussed above. As an example, a dialogue estimation system 203 with dialogue detection is shown in FIG.21. [154] The dialogue estimation system 203 obtains as an input a 5.1 input mix signal MIN.
- the dialogue detection is split into three individual paths, one path for the center channel resulting in a soft gate gain being applied to the center channel in soft gate applicator sub- module 91, a second path for the left and right channel pair resulting in a same soft gate gain being applied to the left and right channels in soft gate applicator sub-module 92 and a third path for the left and right surround channel pair resulting in a same soft gate gain being applied to the left and right surround channels in soft gate applicator sub-module 93.
- the dialogue detector 6 comprises a plurality of dialogue detector sub- modules 61, 62, 63 and downmixers 64, 65.
- Dialogue detector sub-module 63 obtains the center channel and produces a dialogue confidence value for each frame of the center channel alone.
- the L and R channels are downmixed with downmixer 65 to a mono channel which is provided to dialogue detector sub-module 62 which produces a dialogue confidence value for each frame of the mono channel downmix of the L and R channels.
- the surround channels Ls and Rs are downmixed with downmixer 64 to a mono channel which is provided to dialogue detector sub-module 61 which produces a dialogue confidence value for each frame of the mono channel downmix of the Ls and Rs channels.
- the dialogue confidence values output by each respective dialogue detector sub- module 61, 62, 63 is provided to the post-processor 7 which e.g.
- the soft gate gain applicator module 9 comprises soft gate gain applicator sub-modules 91, 92, 93 for applying the soft gate gains output by the soft gain computation module to the respective channels.
- a unique soft gate gain is applied to the center channel in soft gate gain applicator sub-module 93 and the soft gate gain calculated from the L and R channel downmix mono signal is applied to both the L and R channel in the soft gate gain applicator sub-module 92.
- the soft gate gain calculated from the Ls and Rs channel downmix mono signal is applied to both the Ls and Rs channel in the soft gate gain applicator sub-module 91.
- FIG.22 is a block diagram showing yet another dialogue estimation system 204 with dialogue detection.
- Dialogue detection may be enhanced by combining the dialogue confidence extracted directly from the input mix signal MIN using dialogue detector 6a with a so- called alternative dialogue confidence, ⁇ ⁇ , derived from e.g. the dialogue estimation processing performed in the second dialogue estimation processor 5b using dialogue detector 6b. Examples of how to determine the alternative dialogue confidence may be found in International Application No. PCT/US24/14225, hereby incorporated by reference in its entirety.
- the alternative dialogue confidence ⁇ ⁇ may in FIG.22 be computed as, for example, the ratio of the energy of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy of the downmix signal MDF input to the dialogue estimator 5b.
- E(•) denotes the norm of a signal and we will use the (squared) L 2 -norm (i.e. the energy) in the following.
- any norm could be used in the computation of the alternative dialogue confidence, e.g., the 1-norm (absolute value norm), the infinity-norm (max norm) to mention two additional examples well-known to those skilled in the art.
- E(DE2 D ) is the energy of the downmix signal DE2D.
- the computation of the energy of a signal S may be done using various methods depending e.g. on the domain in which the signal S is represented. For example, if the signal S is in time domain, its energy E(S) is proportional to the sum of the squared signal S (i.e. the sum of each squared time domain sample). As another example, if a signal S is represented in a transform domain (e.g.
- E(S) may be found by summation of transform samples (e.g., after taking the absolute value and squaring) across frequency or across both time and frequency (e.g. summation over an entire frame).
- ⁇ ⁇ is in FIG.22 based on the energy ratio E(DE2 D )/E(M DF ) which is typically a value in the interval [0,1]. It is also noted that in some implementations this value may fall outside the interval [0,1] and bounding it to be within this interval may be advantageous.
- This computation of ⁇ ⁇ may be further refined by applying frequency weighting prior to computing the energies, and by applying a non-decreasing mapping to the energy ratio E(DE2 D )/E(M DF ).
- the alternative dialogue confidence values ⁇ ⁇ from dialogue detector 6b are combined with the dialogue confidence values from dialogue detector 6a into a single dialogue confidence value by the dialogue confidence combiner module 67. Examples of how to determine the single, combined, dialogue confidence value are also found in International Application No. PCT/US24/14225.
- the alternative dialogue confidence ⁇ ⁇ may be computed as the ratio of the energy E(DE2D) of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy E(MD) of the downmix signal MD input to the spatial filtering module 3.
- the alternative dialogue confidence ⁇ ⁇ may be computed as the ratio of the energy E(DE2D) of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy (M IN ) of the input mix signal M IN. That is, ⁇ ⁇ may be determined by the dialogue detector 6b based on the energy ratios E(DE2 D )/E(M D ) or E(DE2 D )/E(M IN ).
- the alternative dialogue confidence ⁇ ⁇ may be based on one or more of the energy ratios E(DE2 D )/E(M DF ), E(DE2 D )/E(M D ) or E(DE2 D )/E(M IN ).
- a single alternative dialogue confidence ⁇ ⁇ may be determined based on both the energy ratio E(DE2 D )/E(M DF ) and the energy ratio E(DE2 D )/E(M D ).
- DE2D, MD and MDF are all mono signals meaning that computation of the energy ratios E(DE2 D )/E(M D ) and E(DE2 D )/E(M DF ) is straightforward.
- M IN is a multi-element signal meaning when the alternative dialogue confidence ⁇ ⁇ is to be based energy ratio E(DE2 D )/E(M IN ) the energy of the input mix signal M IN may be defined by combining the energy of each individual element in the input mix signal MIN using e.g. an average-, min- or max-function.
- up to three alternative dialogue confidence values ⁇ ⁇ may be computed from signals between the adaptive downmixer 2 and the adaptive upmixer 4 (based on the energy ratios E(DE2D)/E(MDF), E(DE2D)/E(MD) or E(DE2D)/E(MIN)) when the second dialogue estimate processor 5b is located upstream of the adaptive upmixer 4.
- up to ⁇ + ⁇ additional dialogue confidence values ⁇ may be determined based on the energies of signals input to, and output from, the first dialogue estimate processor 5a. For example, a ratio between the energy of each respective element output from, and input to, the first dialogue estimate processor 5a may be determined.
- up to 3 + ⁇ + ⁇ alternative dialogue confidence values ⁇ ⁇ may be determined in the dialogue estimation system 204 of FIG.22.
- the second dialogue estimate processor 5b is placed downstream of the adaptive upmixer 4 (see e.g. FIG.4A) and in such implementations 3( ⁇ + ⁇ ) alternative dialogue confidence values ⁇ ⁇ , i.e. three per channel/object, may be computed.
- up to ⁇ + ⁇ dialogue confidence values ⁇ may be determined based on the energy of the elements output from, and input to, the first dialogue estimate processor 5a and, similarly, up to ⁇ + ⁇ additional alternative dialogue confidence values ⁇ may be determined based on the energy of each element output from, and input to, the second dialogue estimate processor 5b. Additionally, a further ⁇ + ⁇ alternative dialogue confidence values ⁇ ⁇ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the mix input signal MIN making for a total of 3( ⁇ + ⁇ ) alternative dialogue confidence values.
- an additional ⁇ + ⁇ alternative dialogue confidence values ⁇ ⁇ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the downmix signal MD.
- another ⁇ + ⁇ alternative dialogue confidence values ⁇ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the downmix signal M DF , output by the spatial filtering module 3.
- an alternative dialogue confidence value ⁇ ⁇ may be determined based on the energy of each element output by the second dialogue estimate processor 5b and the energy of any corresponding element upstream of the second dialogue estimate processor 5b.
- an alternative dialogue confidence value ⁇ ⁇ may be determined based on the energy of each element output by the first dialogue estimate processor 5a and the energy of any corresponding element upstream of the first dialogue estimate processor 5a.
- the one or more computed alternative dialogue confidence ⁇ ⁇ may be combined (e.g. using an average function, max function or min function) in various ways with each other and/or with the dialogue confidence from the dialogue detector 6a.
- the ⁇ + ⁇ alternative dialogue confidence values extracted from the first dialogue estimate processor 5a inputs and outputs are labeled ⁇ ⁇ ⁇ + ⁇
- the ⁇ + ⁇ alternative dialogue confidence values based on the inputs and outputs of the second dialogue estimate processor 5b are labeled ⁇ ⁇ ⁇ ⁇ + ⁇ for a total of 2( ⁇ + ⁇ ) alternative , , ... , ⁇ ( ⁇ ) ⁇ , ⁇ ( 1 , , ... , ⁇ .
- At least one alternative dialogue confidence value may be computed for each element for each of the two dialogue estimate processors 5a, 5b it is envisaged that two or more elements are grouped together for the purpose of computing a single alternative dialogue confidence value.
- an input/output signal pair associated with one channel/object e.g. the left channel
- an input/output pair associated with another channel/object e.g. the right channel
- a mono downmix is computed for all the input channels/objects belonging to the group, and similarly a mono downmix is computed for all output channels/objects belonging to the group, and an alternative dialogue confidence value for the group is computed based on the energy of the mono downmix of the input and energy of the mono downmix of the output.
- an alternative dialogue confidence value for the group is computed based on the energy of the mono downmix of the input and energy of the mono downmix of the output.
- the input and output of the dialogue estimator 51:1 are provided to alternative dialogue confidence computation module 6b which extracts a first alternative dialogue confidence value by calculating the energy of the input and output and computing the ratio E(output)/E(input).
- the second group is associated with the first dialogue estimator 51:2 processing the L and R channel, and the L and R input channels are downmixed in downmixer 68 and the L and R output channels are downmixed in downmixer 69.
- a second alternative dialogue confidence value is then computed based on the energy of the mono downmix from downmixer 68 and the energy of the mono downmix from downmixer 69, by alternative dialogue confidence computation module 6c.
- FIG.24 shows one example of a dialogue enhancement system 301 comprising a dialogue estimation system 10C, dialogue estimate gain module 11, an input mix signal gain module 12 and a mixer 13.
- the dialogue estimation system 10C obtains the input mix signal M IN as an input, processes the input mix signal M IN in accordance with the above and outputs a dialogue estimate DE2 (or possibly a soft gated dialogue estimate DE2G).
- the dialogue estimate DE2 may be referred to as a dialogue signal in the context of dialogue enhancement systems.
- the dialogue estimate DE2 is provided to the dialogue estimate gain module 11 which applies a dialogue gain to dialogue estimate DE2 forming a gain modified dialogue estimate DE2 Gain .
- the dialogue gain may be any value and may be an attenuating gain or a boosting gain.
- the input mix signal MIN is also provided to the input mix signal gain module 12 which applies a mix gain to the input mix signal MIN yielding input mix signal MGain.
- the mix gain may also be any value and may be a boosting gain or an attenuating gain.
- the mix gain is constant and equal to 1, meaning the input MIN to the input mix signal gain module 12 is equal to the output MGain.
- the gain modified dialogue estimate DE2Gain and input mix signal M Gain are provided to mixer 13 which mixes the two signals to form a final mix output signal DE2 Final containing the gained dialogue estimate DE2 Gain mixed with the gained input mix signal MGain.
- the final mix output signal DE2Final is the output of the dialogue enhancement system 301.
- a dialogue gain of -1 may lead to complete silencing of dialogue in DE2Final if DE2 is equal to the dialogue component of M IN .
- FIG.25 shows another exemplary dialogue enhancement system 302, similar to the dialogue enhancement system 301 of FIG.24, which also forms an output dialogue estimate DE2 Final based on the input mix signal M IN .
- a difference between the dialogue enhancement system 302 and the dialogue enhancement system 301 is that dialogue enhancement system 302 of FIG.25 comprises a background estimator 14 upstream of the input mix signal gain module 12.
- the background signal MB is provided to the input mix signal gain module 12 which applies the mix gain to form a gain modified background signal M BGain .
- a dialogue gain of 1 yields the input mix signal MIN at the output, a dialogue gain between 0 and 1 results in dialogue attenuation, and a gain larger than 1 results in a dialogue boost.
- One advantage of this implementation is that processing of the dialogue estimate DE2 and non-dialogue (i.e.
- the background MB may be done separately, i.e., the processing of the dialogue estimate DE2 will not affect the background content and vice versa.
- the dialogue estimate DE2 Gain is mixed with the non-dialogue estimate MB,Gain resulting in a (dialogue enhanced) dialogue estimate signal DE2Final comprising a mix of the gained dialogue estimate DE2Gain and the gained non-dialogue MB,Gain.
- the background gain and dialogue gain applied to DE2 and MB respectively to form DE2Gain and MB,Gain are configured to boost the dialogue estimate DE2 or attenuate the non-dialogue estimate MB.
- a parametric coding block 17 is used to parametrically encode the dialogue estimate DE2 with respect to the input mix signal MIN to yield a dialogue estimate parameter bitstream.
- the input mix signal MIN as such is encoded (e.g. waveform encoded) using encoding block 16 to yield an input mix signal bitstream.
- the dialogue estimate parameter bitstream and the input mix signal bitstream are provided to a bitstream multiplexer 18 which multiplexes the two bitstreams into a complete bitstream B.
- the dialogue estimate system 10C is used with an audio codec that encodes the dialogue waveform and the background waveform into two separate bitstreams using encoding modules 19a, 19b.
- the separate bitstreams are subsequently multiplexed to form a complete bitstream B by bitstream multiplexing module 18.
- the audio codec could be AC4.
- the present disclosure likewise relates to an apparatus (e.g., computer-implemented apparatus or apparatus having processing capability in general) for performing or implementing methods and techniques described throughout the present disclosure.
- this apparatus may relate to a dialogue estimation system (with or without dialogue detection) or a dialogue enhancement system incorporating the dialogue estimation system (with or without dialogue detection).
- FIG.29 shows an example of such apparatus 1100.
- apparatus 1100 comprises a processor 1110 and a memory 1120 coupled to the processor 1110.
- the memory 1120 may store instructions for the processor 1110.
- the processor 1110 may also receive, among others, suitable input data 1130 (e.g., input frames of time-frequency coefficients, etc.), depending on use cases and/or implementations.
- suitable input data 1130 e.g., input frames of time-frequency coefficients, etc.
- the processor 1110 may be adapted to carry out or implement the methods/techniques described throughout the present disclosure and to generate corresponding output data 1140, depending on use cases and/or implementations.
- the present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products.
- aspects of the methods and apparatus/systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files.
- Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers.
- Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
- WAN Wide Area Network
- LAN Local Area Network
- One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system.
- the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and/or application specific integrated circuits (“ASICs”).
- ASICs application specific integrated circuits
- the apparatus e.g., encoders
- the apparatus may include one or more electronic processors, one or more computer-readable medium modules, one or more input/output interfaces, and various connections (e.g., a system bus) connecting the various components.
- various connections e.g., a system bus
- a dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of N ⁇ 1 channels and/or M ⁇ 1 audio objects, wherein M + N ⁇ 2, the system comprising: an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal; a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal; an adaptive downmixer configured to: downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; an adaptive upmixer configured to: upmix the downmixed signal based on the spatial parameters to generate an upmixed signal; and an output dialogue estimate processor configured to: determine a second dialogue estimate of the upmixed signal to generate an output dialogue signal.
- an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal
- a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal
- an adaptive downmixer configured to: downmix the input mix signal based on the input mix
- EEE6 The dialogue estimation system of EEE5, wherein the mask-based system comprises one or more neural networks.
- the dialogue estimation system of EEE8 or EEE9, wherein the direct-mapped system comprises one or more neural networks.
- EEE11 The dialogue estimation system of any one of EEE8-EEE10, wherein the direct- mapped system comprises at least one of a generative model or a UNIVERSE model.
- EEE12. The dialogue estimation system of any one of EEE1-EEE11, wherein the input dialogue estimate processor comprises: K single element dialogue estimate processors, wherein K ⁇ 1, and wherein each single element dialogue estimate processor of the K single element dialogue estimate processors is configured to: receive, as input, a single channel input mix signal or a single object input mix signal; and determine a dialogue estimate of the single channel input mix signal or the single object input mix signal.
- EEE14 The dialogue estimation system of any one of EEE1-EEE11, wherein the input dialogue estimate processor comprises: a single element mask generator; and wherein the input dialogue estimate processor is configured to: downmix the input mix signal into a mono mix signal; determine a mask of the mono mix signal via the single element mask generator; and apply the mask to each channel of the input mix signal and/or each audio object of the input mix signal to determine a dialogue estimate of each channel and/or each audio object.
- EEE16 The dialogue estimation system of EEE15, wherein the input dialogue estimate processor is further configured to: apply the first mask to the second channel or object to determine a first masked version of second channel or object; apply the second mask to the second channel or object to determine a second masked version of second channel or object; and combine the first masked version and the second masked version of the second channel or object to determine a dialogue estimate of the second channel or object of the input mix signal.
- EEE18 The dialogue estimation system of any one of EEE1 and EEE3-EEE17, wherein the output dialogue estimate processor comprises a mask-based system configured to determine the second dialogue estimate by: determining a time and frequency varying mask based on the downmixed signal; and applying the mask to the downmixed signal to determine the second dialogue estimate of the downmixed signal.
- EEE19 The dialogue estimation system of EEE18, wherein the mask-based system comprises one or more neural networks.
- EEE20 The dialogue estimation system of EEE18 or EEE19, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model.
- EEE21 The dialogue estimation system of EEE18 or EEE19, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model.
- the dialogue estimation system of EEE21, wherein the direct-mapped system comprises one or more neural networks.
- the dialogue estimation system of EEE21 or EEE22, wherein the direct-mapped system comprises at least one of a generative model or a UNIVERSE model.
- the dialogue estimation system of any one of EEE1-EEE23, wherein the spatial parameters comprise at least one of adaptive downmixing parameters and adaptive upmixing parameters.
- the dialogue estimation system of any one of EEE1-EEE24 wherein the spatial analyzer is configured to determine the spatial parameters based on the first dialogue estimate of the input mix signal via correlation based spatial analysis and/or amplitude ratio based spatial analysis.
- EEE26 The dialogue estimation system of any one of EEE1-EEE25, wherein the spatial parameters comprise one scalar value per channel of the input mix signal or one scalar value per audio object of the input mix signal.
- EEE27 The dialogue estimation system of any one of EEE1-EEE26, further comprising a dialogue detector configured to determine one or more dialogue confidence values.
- EEE28 The dialogue estimation system according to any one of the preceding EEEs, wherein the spatial parameters are panning parameters.
- EEE29 The dialogue estimation system according to any one of the preceding EEEs, wherein the spatial parameters are panning parameters.
- the input mix signal comprises a subset of the channels and/or audio objects of an original mix signal, the original mix signal comprising at least one residual channel and/or audio object in addition to the at least one of N ⁇ 1 channels and/or M ⁇ 1 audio objects of the mix input signal.
- the dialogue estimation system is further configured to: combine the output dialogue signal with the at least one residual channel and/or audio object to form a hybrid dialogue estimate signal.
- a dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: apply a dialogue gain to the output dialogue signal of the dialogue estimation system to compute a dialogue estimation; and determine a dialogue enhanced mix signal based on a combination of the dialogue estimation and the input mix signal.
- a dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; apply a dialogue gain to the output dialogue signal to determine a first signal; apply a background gain to the background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
- a dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; process the background estimate and the output dialogue signal; apply a dialogue gain to the processed dialogue signal to determine a first signal; apply a background gain to the processed background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
- the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; process the background estimate and the output dialogue signal; apply a dialogue gain to the processed dialogue signal to determine a first signal; apply a background gain to the processed background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
- a dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: apply parametric coding to the output dialogue signal of the dialogue estimation system to determine a dialogue parameter bitstream; and encode the input mix signal to determine a mix waveform bitstream.
- EEE35 A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; encode the output dialogue signal to determine a dialogue bitstream; and encode the background estimate to determine a background bitstream.
- EEE36
- EEE38 The dialogue estimation method according to EEE36 or EEE37, wherein the spatial parameters are panning parameters.
- EEE39 An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of EEE36- EEE38.
- EEE40 A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE36-EEE38.
- EEE41 A computer-readable storage medium storing the program according to EEE40.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Stereophonic System (AREA)
Abstract
An aspect of the present disclosure relates to a dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of audio channels and/or audio objects. The system comprising a first dialogue estimate processor configured to determine a first dialogue estimate of the input mix signal and a spatial analyzer configured to determine spatial parameters based on the first dialogue estimate of the input mix signal. The system further comprises an adaptive downmixer configured to downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal and a second dialogue estimate processor configured to determine a second dialogue estimate of the downmixed signal. The system further comprises an adaptive upmixer configured to upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
Description
SYSTEMS AND METHODS FOR SPATIAL FIDELITY IMPROVING DIALOGUE ESTIMATION CROSS-REFERENCE TO RELATED APPLICATIONS [001] This application claims priority of the following priority application: U.S. Provisional Patent Application No.63/563,651, filed on 11 March 2024, and U.S. Provisional Patent Application No.63/678,508, filed on 1 August 2024, each of which is hereby incorporated by reference in its entirety. TECHNICAL FIELD [002] The present disclosure relates to a method and system for dialogue estimation, and a dialogue enhancement system utilizing the dialogue estimation system. BACKGROUND [003] Dialogue estimation (DE) is the process of extracting dialogue from an original signal where dialogue and non-dialogue sounds are mixed. In general, DE exploits temporal, spectral and spatial properties of dialogue and non-dialogue sounds to extract dialogue from a mix. [004] DE for monoaural audio exploits temporal and spectral properties of dialogue. Recent advancements in neural network (NN) training and in NN model structures have greatly improved the DE quality in single channel audio applications. However, NN based systems for multi-channel DE are still in their infancy. [005] Existing solutions for multi-channel DE involves downmixing the multiple channels into a mono signal which is spatially stable, processing the mono signal to obtain a mono dialogue estimate and upmixing the mono dialogue estimate to the desired multi-channel configuration (stereo, 3.0, 5.1 etc.). In some implementations, to downmix a stereo input signal to a mono signal the stereo signal is converted to a mid-side signal pair whereby the side signal is set to zero, which effectively yields a downmix from stereo to mono. [006] To improve the signal to noise ratio the downmixing and upmixing is adaptive, wherein weights are assigned to the input channels according to where the spatial analysis indicates dialog is present. For example, if a stereo input signal has dialog only in the left channel and interfering background in the right channel, adaptive downmixing would only choose contributions from the left channel.
SUMMARY [007] A weakness of existing systems for DE is that the spatial analysis is based on the input mix signal, which may be degraded by the presence of the interfering non-dialogue sounds. This degradation may manifest itself as spatial instabilities, and in severe cases loss of energy in the dialogue estimate (because the adaptive downmix is “misled” by the spatial analysis). Since the processing is frequency selective, the loss of energy may be perceived as timbre changes in the dialogue estimate. [008] It is a purpose of the present disclosure to present an improved dialogue estimation system and method for performing dialogue estimation which overcomes at least some of the shortcomings of existing solutions. [009] According to a first aspect there is provided a dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of ^^^^ ≥ 1 channels and/or ^^^^ ≥ 1 audio objects, wherein ^^^^ + ^^^^ ≥ 2. The system comprises an input dialogue estimate processor (first dialogue estimate processor) configured to determine a first dialogue estimate of the input mix signal and a spatial analyzer configured to determine spatial parameters (e.g. panning parameters) based on the first dialogue estimate of the input mix signal. The system further comprises an adaptive downmixer configured to downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal and an output dialogue estimate processor (second dialogue estimate processor) configured to determine a second dialogue estimate of the downmixed signal. The system further comprises an adaptive upmixer configured to upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal. [010] By utilization of a first dialogue estimate processor that determines a first dialogue estimate, whereby the spatial parameters are determined based on the first dialogue estimate, the spatial analysis of this system becomes more robust even when the input mix signal is dominated by non-dialogue content. To further facilitate dialogue estimation the second dialogue estimate processor is applied in the main signal processing chain. Here, the second dialogue estimate processor is applied on the downmixed signal. [011] The dialogue estimation system may be used with any type of multi-channel input mix signal and the presented methods and systems may produce a multi-channel dialogue estimate. In the following, dialogue estimation refers to methods that extract a multi-channel dialogue estimate from a multi-channel input mix signal. The presented methods and systems may also operate on input mix signals comprising audio objects rather than audio channels, or input mix signals comprising a combination of audio objects and channels. Sometimes the term
“element” is used to mean either a channel or an object, and we use the terms “channel”, “object”, and “element” interchangeably in the following. [012] According to a second aspect of the disclosure there is provided a dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of ^^^^ ≥ 1 channels and/or ^^^^ ≥ 1 audio objects, wherein ^^^^ + ^^^^ ≥ 2. The system comprising a first dialogue estimate processor configured to determine a first dialogue estimate of the input mix signal and a spatial analyzer configured to determine spatial parameters based on the first dialogue estimate of the input mix signal. The system further comprises an adaptive downmixer configured to downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal and an adaptive upmixer configured to upmix the downmixed signal based on the spatial parameters to generate an upmixed signal. The system further comprises a second dialogue estimate processor configured to determine a second dialogue estimate of the upmixed signal to generate an output dialogue signal. [013] Hereby, in the second aspect the output dialogue estimate processor is applied to the upmixed signal. The benefits of improved robustness even when the input mix signal is dominated by non-dialogue content are also applicable to the second aspect. [014] According to a third aspect of the disclosure there is provided a dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of ^^^^ ≥ 1 channels and/or ^^^^ ≥ 1 audio objects, wherein ^^^^ + ^^^^ ≥ 2. The method comprises determining a first dialogue estimate of the input mix signal and determining spatial parameters based on the first dialogue estimate of the input mix signal. The method further comprises downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal, determining a second dialogue estimate of the downmixed signal and upmixing the second dialogue estimate based on the spatial parameters to generate an output dialogue signal. [015] According to a fourth aspect of the disclosure there is provided a dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of ^^^^ ≥ 1 channels and/or ^^^^ ≥ 1 audio objects, wherein ^^^^ + ^^^^ ≥ 2. The method comprising determining a first dialogue estimate of the input mix signal, determining spatial parameters based on the first dialogue estimate of the input mix signal and downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal. The method further comprises upmixing the downmixed signal based on the spatial parameters to generate an upmixed signal and determining a second dialogue estimate of the upmixed signal to generate an output dialogue signal.
[016] According to a fifth aspect of the disclosure there is provided an apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to the third or fourth aspect of the disclosure. [017] According to a sixth aspect of the disclosure there is provided a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to the third or fourth aspect of the disclosure. [018] According to a seventh aspect of the disclosure there is provided a computer- readable storage medium storing the program according to the sixth aspect of the disclosure. [019] Any features and benefits described in connection with a system are applicable also for the associated method, and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS [020] Embodiments of the invention will be described in more detail with reference to the appended drawings. [021] Figure 1 is a block diagram illustrating a dialogue estimation system with a single dialogue estimation processor. [022] Figure 2 is a block diagram illustrating an improved dialogue estimation system with a single dialogue estimation processor. [023] Figure 3A is a block diagram illustrating a dialogue estimation system with two dialogue estimation processors, wherein one dialogue estimation processor is located upstream of the adaptive upmixer according to some implementations. [024] Figure 3B is a flowchart describing a dialogue estimation method according to some implementations. [025] Figure 4A is a block diagram illustrating an alternative dialogue estimation system with two dialogue estimation processors, wherein one dialogue estimation processor is located downstream of the adaptive upmixer according to some implementations. [026] Figure 4B is a flowchart describing an alternative dialogue estimation method according to some implementations. [027] Figure 5 is a block diagram illustrating a two-element dialogue estimation processor, according to some implementations. [028] Figure 6 is a block diagram illustrating a multi-element dialogue estimation processor, according to some implementations. [029] Figure 7 is a block diagram illustrating a two-element dialogue estimation processor with single gain mask computation, according to some implementations.
[030] Figure 8 is a block diagram illustrating a multi-element dialogue estimation processor with single gain mask computation, according to some implementations. [031] Figure 9 is a block diagram illustrating a three-element dialogue estimation processor with dual gain mask computation, according to some implementations. [032] Figure 10 is a block diagram illustrating another exemplary three-element dialogue estimation processor with dual gain mask computation. [033] Figure 11 is a block diagram illustrating a first exemplary dialogue estimation system for processing three audio channels, according to some implementations. [034] Figure 12 is a block diagram illustrating a second exemplary dialogue estimation system for processing three audio channels, according to some implementations. [035] Figure 13 is a block diagram illustrating a third exemplary dialogue estimation system for processing three audio channels, according to some implementations. [036] Figure 14 is a block diagram illustrating a fourth exemplary dialogue estimation system for processing three audio channels, according to some implementations. [037] Figure 15 is a block diagram illustrating a first exemplary dialogue estimation system for processing two audio channels, according to some implementations. [038] Figure 16 is a block diagram illustrating a second exemplary dialogue estimation system for processing two audio channels, according to some implementations. [039] Figure 17 is a block diagram illustrating an exemplary dialogue estimation system for processing nine audio channels, according to some implementations. [040] Figure 18 is a block diagram illustrating an exemplary dialogue estimation system for processing a mix of five audio channels and four audio elements, according to some implementations. [041] Figure 19 is a block diagram illustrating a first exemplary dialogue estimation system with dialogue detection, according to some implementations. [042] Figure 20 is a block diagram illustrating a second exemplary dialogue estimation system with dialogue detection, according to some implementations. [043] Figure 21 is a block diagram illustrating a third exemplary dialogue estimation system with dialogue detection, according to some implementations. [044] Figure 22 is a block diagram illustrating a fourth exemplary dialogue estimation system with dialogue detection, according to some implementations. [045] Figure 23 is a block diagram illustrating a fifth exemplary dialogue estimation system with dialogue detection, according to some implementations. [046] Figure 24 is a block diagram illustrating a first exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations.
[047] Figure 25 is a block diagram illustrating a second exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations. [048] Figure 26 is a block diagram illustrating a third exemplary dialogue enhancement system incorporating a dialogue estimation system according to some implementations. [049] Figure 27 is a block diagram illustrating a coding system for encoding the dialogue estimate and the mix input signal, according to some implementations. [050] Figure 28 is a block diagram illustrating another coding system for encoding the dialogue estimate and an extracted non-dialogue signal according to some implementations. [051] Figure 29 is a block diagram illustrating an apparatus for performing or implementing methods and techniques described throughout the present disclosure. DETAILED DESCRIPTION OF CURRENTLY PREFERRED EMBODIMENTS [052] FIG.1 is a block diagram showing a dialogue estimation system 10A bearing some resemblance to the solutions presented in the international patent application with publication number WO2021/252823 and in International Application No. PCT/US23/63717, each of which are hereby incorporated by reference in its entirety. [053] The dialogue estimation system 10A is sometimes referred to as a “DeepSpace” system. The dialogue estimation system 10A obtains an input stereo audio signal SIN comprising two channels wherein each of the channels carries a mix of speech content and non-speech content. The input stereo audio signal SIN is provided to a spatial analyzer 1 which extracts spatial parameters Pspatial (e.g. panning parameters, inter-channel time difference, inter-channel level difference and magnitude) of a target audio source (e.g. dialogue) present in the input stereo audio signal SIN. The spatial parameters are adaptive and may vary over time and may vary for different frequency bands. The spatial parameters are then provided to an adaptive downmixer 2 which forms a downmix signal SD based on the adaptive spatial parameters and the input stereo audio signal SIN. The downmix signal SD may be a two-channel signal (e.g. comprising an adaptive mid and side channel) or a mono channel signal (e.g. comprising only an adaptive mid channel). [054] The downmix signal SD is provided to a spatial filtering module 3 which performs spatial filtering. The spatial filtering module 3 is optional and may be omitted or located at other positions in the processing chain. The spatial filtering may involve setting the adaptive side channel of the downmix signal to zero. In another example, the spatial filtering is equivalent to the Spatio-Level Filtering (SLF) presented in A. S. Master, et. al "Dialog Enhancement via Spatio-Level Filtering and Classification," in 149th Audio Engineering Society Convention, New York, 2020, which is hereby incorporated by reference in its entirety.
[055] If the spatial filtering module 3 is used, the output spatially filtered downmix signal SDF is provided to an adaptive upmixer 4 which performs adaptive upmixing using the spatial parameters Pspatial. If the spatial filtering module 3 is omitted, the downmix signal SD is provided directly to the adaptive upmixer 4. The adaptive upmixer 4 performs upmixing, based on the spatial parameters Pspatial to obtain a stereo output signal SOUT. The adaptive upmixer 4 may e.g. convert a downmix mid channel and side channel to pair of stereo channels or upmix a single mono channel to pair of stereo channels, based on the spatial parameters to retain e.g. the panning and/or inter-channel phase difference of the stereo input signal SIN. [056] The stereo output signal SOUT is subsequently provided to a dialogue estimation processor 5. The dialogue estimation processor 5 may be a low spatial fidelity (LOSPAFI) dialogue estimator, as described below. [057] Herein the term “dialogue” is used to refer to any type of speech in general. While the term “dialogue” in some contexts is interpreted to mean a conversation between two or more persons the term “dialogue” is in the present disclosure used to refer to any type of human speech sounds, regardless of the speech sounds being uttered by a single person or multiple persons. [058] The dialogue estimate processor 5 performs dialogue separation (e.g. using filtering and/or using a neural network) to isolate dialogue content and remove, or at least attenuate, non- dialogue content such as noise, music and/or effects. [059] The output of the dialogue estimate processor 5 is accordingly a dialogue estimate stereo output signal SOUTDE. [060] However, a weakness of the dialogue estimation system 10A in FIG.1 is that the spatial analyzer 1 extracts the spatial parameters based on the input stereo signal SIN, and the input stereo signal SIN may be degraded by the presence of the interfering non-dialogue sounds (e.g. noise, music and/or effects). This degradation may manifest itself as spatial instabilities in the stereo output signal SOUT and also in the output dialogue estimate SOUTDE. In severe cases this could lead to loss of energy in the dialogue estimate SOUTDE (because the adaptive downmix, spatial filtering and/or adaptive upmixing is “misled” by the spatial parameters). In cases when the processing is frequency selective, the loss of dialogue energy may be perceived as disruptive timbre changes in the dialogue estimate stereo output signal SOUTDE. [061] An improvement to the dialogue estimation system 10A in FIG.1 is shown in FIG. 2. The dialogue estimation system 10B of FIG.2 is similar to the system of FIG.1, however a difference is that the dialogue estimate processor 5 has been moved upstream such that the dialogue estimation is performed upstream of the adaptive downmixer 2, upstream of the spatial filtering module 3 and upstream of the adaptive upmixer 4. The spatial analyzer 1 may now
determine the spatial parameters Pspatial based on the dialogue estimate stereo input signal SDE, instead of the input stereo signal SIN, meaning that the spatial analysis becomes more robust and generally achieves improved dialogue estimation even when the stereo input signal SIN contains a large portion of non-dialogue audio content. Additionally, since both the input stereo signal SIN and dialogue estimate stereo input signal SDE are available it is possible to extract the spatial parameters Pspatial based on a combination of both the input stereo signal SIN and the dialogue estimate stereo input signal SDE. [062] The dialogue estimate stereo input signal SDE is provided to the adaptive downmixer 2 which downmixes the stereo speech estimate SDE based on the spatial parameters Pspatial to form a downmix dialogue estimate SDED. The downmix dialogue estimate SDED is then optionally subject to spatial filtering using the spatial filtering module 3 forming a filtered downmix dialogue estimate SDEDF. Lastly, the filtered downmix dialogue estimate SDEDF is provided to the adaptive upmixer 4 which upmixes the filtered downmix dialogue based on the spatial parameters Pspatial estimate SDEDF to form the dialogue estimate stereo output signal SOUTDE. [063] However, the adaptive downmixer 2 in FIG.2 is operating in a suboptimal way (the nature of the suboptimality depends on how the dialogue estimate processor 5 is configured). The dialogue estimation system 10B additionally allows more freedom in what signals the spatial analysis is based upon since the stereo input signal SIN and the dialogue estimate stereo input signal SDE are both available and may be used in combination. [064] FIG.3A shows a dialogue estimation system 10C according to some implementations of the present disclosure. [065] To address the disadvantages of the dialogue estimation system 10A from FIG.1, at least two dialogue estimate processors 5a, 5b are used in the processing chain, an input dialogue estimate processor 5a (or “first” dialogue estimate processor) and an output dialogue estimate processor 5b (or “second” dialogue estimate processor). The input dialogue estimate processor 5a makes the spatial analysis, and the resulting spatial parameters Pspatial, more robust to non- dialogue interference in the input mix signal. Furthermore, the input dialogue estimate processor 5a does not affect the main signal path (the main signal path comprising the adaptive downmixer 2, the spatial filtering module 3 and the adaptive upmixer 4) and the adaptive downmixer 2 is allowed to operate on the input mix signal MIN directly, which is beneficial as explained below. To further leverage the dialogue estimation of dialogue estimate processors, a second dialogue estimate processor 5b is introduced in the main signal path downstream of the adaptive downmixer 2. [066] In general, however, the second dialogue estimate processor 5b may be located at any point downstream of the adaptive downmixer 2 and in FIG.4 an alternative dialogue
estimation system 10D is shown where the second dialogue estimate processor 5b is located downstream of the adaptive upmixer 4, as described below. [067] With further reference to FIG.3A a dialogue estimation method for extracting dialogue from the input mix signal MIN will now be described. First, an input mix signal MIN is obtained. The input mix signal comprises at least two audio elements wherein each audio element is either an audio object or a channel as described below. At step S1A a first dialogue estimate DE1 is extracted from the input mix signal MIN by the first (input) dialogue estimate processor 5a. The method then goes to step S2A involving determining spatial parameters based on the first dialogue estimate DE1. This is achieved using the spatial analyzer 1. [068] The method then goes to step S3A comprising downmixing the input mix signal MIN based on the input mix signal MIN and the spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ to generate the downmixed signal MD. The downmixing is performed by the adaptive downmixer 2 which is controlled by the spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^. [069] The method then goes to step S4A comprising determining a second dialogue estimate DE2D based on the downmixed signal MD assuming no spatial filtering is done. The second dialogue estimate DE2D is extracted by the second (output) dialogue estimate processor 5b. The second dialogue estimate DE2D is a downmixed representation and at step S5a the second dialogue estimate DE2D is upmixed by adaptive upmixer 4 to yield an output dialogue estimate DE2 as an upmixed version of DE2D. The adaptive upmixer 4 is controlled by the spatial parameters. [070] A dialogue estimation system 10C is in some implementations used together with a dialogue detector (see FIGS.19-23) and/or with a mixer for mixing the input mix signal back to the dialogue estimate DE (see FIGS.24-26). Accordingly, the method may optionally comprise step S6 comprising applying a soft gate gain to the dialogue estimate DE, wherein the soft gate gain is based on a dialogue confidence value extracted by a dialogue detector. Additionally or alternatively, the method may comprise optional step S7 comprising mixing the dialogue estimate DE with the input mix signal MIN to obtain a final dialogue estimate as described below. [071] Turning to FIGS.4A and 4B, an alternative dialogue estimation system 10D is shown alongside a flowchart illustrating an alternative dialogue estimation method for extracting dialogue from the input mix signal MIN. [072] Steps S1B and S2B are generally similar to steps S1A and S1B described above. That is, at step S1B a first dialogue estimate DE1 is extracted from the input mix signal MIN by the first (input) dialogue estimate processor 5a. The method then goes to step S2B involving determining spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ based on the first dialogue estimate DE1. The spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ may e.g. comprise panning parameters ^^^^^^^^ as described below, or other types
of spatial parameters. The spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ are determined using the spatial analyzer 1. The method then goes to step S3B comprising downmixing the input mix signal MIN based on the input mix signal MIN and the spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ to generate the downmixed signal MD. The downmixing is performed by the adaptive downmixer 2 which is controlled by the spatial parameters ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^. At step S4B, the downmixed input mix signal MD is upmixed using adaptive upmixer 4 and at step S5B the second (output) dialogue estimate processor 5b extracts the second dialogue estimate DE2 based on the upmixed signal MUF (compare to step S4A where the second dialogue estimate DE2D is extracted from the downmixed signal MD). The second dialogue estimate DE2 is used as the dialogue estimate DE output of the dialogue estimate system 10C. [073] The dialogue estimator systems 10C, 10D operate on an input mix signal MIN and outputs a dialogue estimate DE. The input mix signal MIN is not limited to a stereo audio signal comprising two stereo channels and may for the purposes of this disclosure be any type of audio signal comprising (A) at least two channels, (B) at least two audio objects or (C) at least one audio channel and at least one audio object. The term audio “element” will be used to refer generally to an audio channel, or an audio object whereby input mix signal MIN and the dialogue estimate DE may be said to comprise at least two audio elements. [074] An audio channel may be an audio channel associated with a specific loudspeaker of an audio presentation format. For example, in a stereo audio signal there are two channels, a left channel associated with a left loudspeaker and a right channel associated with a right loudspeaker. In a 5.1 audio signal there are five channels associated with the left, right, center, left surround, right surround loudspeaker respectively and one low frequency effects, LFE, channel associated with the subwoofer. [075] An audio object is an audio signal which has an associated spatial position which may vary with time. An object based signal may comprise one or more audio objects wherein the spatial position of each object may vary with time as signaled with metadata. [076] The input mix signal MIN comprises N number of channels and/or M number of audio objects wherein ^^^^ + ^^^^ ≥ 2. Accordingly, the total number of audio elements are at least two. For example, the input mix signal MIN comprises two channels and no audio objects meaning ^^^^ = 2, ^^^^ = 0 and ^^^^ + ^^^^ = 2, or the input mix signal MIN comprises one audio channel and audio object meaning ^^^^ = 1, ^^^^ = 1 and ^^^^ + ^^^^ = 2. [077] Further examples of an input mix signal MIN include a stereo signal (N = 2), a left, right and center (LRC) channel triplet of a 5.1 channel presentation format (N = 3), and a complete 5.1 channel set excluding the LFE channel (N=5).
[078] The number and type of elements (i.e. audio objects and/or channels) define the spatial configuration of the input mix signal MIN. For example, if the spatial configuration is “stereo”, the MIN comprises two channels, a left channel and a right channel and if the spatial configuration is “5.1” the input mix signal MIN comprises the center, left, right, surround left, surround right and LFE channels typically associated with a 5.1 presentation. [079] The dialogue estimate signal DE also consists of N spatial channels and/or M audio objects corresponding to the number N of channels and/or number M of audio objects of the input mix signal MIN. Hereby, the dialogue estimate systems described herein may be used to process any type of multi-channel (multi-element) audio signal wherein the output dialogue estimate DE2 has corresponding audio elements to the input mix signal MIN. [080] It is also understood that the dialogue estimation systems envisaged in the present disclosure need not operate on all elements of a spatial configuration, i.e. the input mix signal MIN is not necessarily all elements of an original mix signal. For example, dialogue estimation systems may operate on a predetermined subset elements of a spatial configuration whereby the remaining elements (referred to as residual elements) remain unprocessed and are combined with the processed elements to yield a hybrid output dialogue estimate spatial configuration with at least two processed elements and one or more unprocessed elements. For example, for an original 5.1 signal or an original 5.1.4 signal the mix input signal MIN is only the LRC triplet whereas the surround channels and the height channels (as well as the LFE channel) remain unprocessed. Since the surround channels and the height channels (and the LFE channel) rarely contain dialogue, these may be excluded from the processing. Processing fewer channels reduces the computational complexity and including additional channels which are not expected to contain dialogue only increases interference from background audio content. [081] The processing is typically done in time-frequency (TF) tiles with a suitable time- frequency resolution, although broadband (but still time varying) processing may be used in some applications. There are plenty of options when it comes to suitable TF transforms. [082] In some implementations the short-term discrete Fourier transform (ST-DFT) is used, operating on overlapping blocks of sizes 20-100 ms. A concrete example is using a ST- DFT analysis block size of 4096 samples at a sampling frequency of 48000 Hz. For each 4096 samples block (85.3 ms), the DFT yields 2048 frequency domain coefficients representing frequencies from 0-24000 Hz. Those coefficients may be grouped into processing bands, and any modifications of the signal are done by modifying all coefficients in a band by the same gain. These basic concepts of signal adaptive frequency domain processing are considered well-known and further details are left to those skilled in the art. In the following, this type of processing in
TF tiles is assumed even if this is not explicitly stated. Furthermore, for brevity and increased readability, time and frequency indices are often omitted from equations and formulas. [083] Each dialogue estimate processor 5a, 5b may implement any type of processing that processes a single- or multi-element mix signal (e.g. a single or multi audio object/channel mix signal, or a mix signal with a combination of one or more channels and one or more audio objects) and outputs a dialogue estimate DE1, DE2, DE2D. [084] The dialogue estimate processors 5a, 5b envisaged are of a “low spatial fidelity” (abbreviated LOSPAFI) which indicates that the spatial fidelity of the output of each estimator may leave room for improvement, and such improvement may be achieved by implementing the proposed systems 10C, 10D of FIGS.3A and 4A. Below, multiple concrete examples of dialogue estimation systems 10C, 10D and dialogue estimate processors 5a, 5b are presented. All exemplary dialogue estimation systems 10C, 10D and dialogue estimate processors 5a, 5b use single-element dialogue estimators or single-element mask computation modules as the basic building block but the single-element processed by each dialogue estimator or mask computation module varies and may be either an audio channel or an audio object, or a mix, e.g., the sum, of certain audio channels and/or audio objects. [085] However, it is envisaged that any multi-channel dialogue estimate processor (or multi-element dialogue estimate system) may be used in the proposed systems 10C, 10D. [086] In general, single-channel dialogue estimate processors (i.e., single-element dialogue estimators) may be categorized into filtering/mask based estimators and direct mapped estimators. [087] In filtering/mask based dialogue estimate processors, a time and frequency varying filter/mask is computed (e.g., via a mask generator) and then applied to the signal input to the processor. In direct-mapped estimators, there is no explicit filter or mask, and the output of the processor is computed as a direct mapping from the signal input to the processor. Both categories may advantageously employ neural networks (NN). [088] An example of a mask based NN DE system is LensNet. LensNet is described in U.S. Patent Publication No.20230368807 titled “Deep learning based speech enhancement”, hereby incorporated by reference in its entirety. The LensNet model takes as input a mono audio signal and predicts a gain mask for estimating the dialogue of the input mono signal. [089] Examples of direct mapped systems include so called generative NNs, for example UNIVERSE. UNIVERSE is described in PCT Publication No. WO 2023/052523 A1 titled “Universal speech enhancement using generative neural networks”, hereby incorporated by reference in its entirety. The UNIVERSE model includes a neural network based system for speech (dialogue) enhancement, consisting of a generative network for generating an enhanced
audio signal and a conditioning network for generating conditioning information for the generative network. [090] Masks and/or filters may in some cases be derived from direct mapped systems. Obtaining a gain mask from a direct-mapped dialogue estimator may be accomplished by generating the (mono) direct mapped dialogue estimate, and deriving a mono gain mask based on the mono input and mono output of the direct mapped system, e.g., by letting the mask be the ratio of the energy of the output to the energy of the input (done in appropriate TF tiles to obtain a frequency selective mask). This mono mask may then be used analogously to a gain mask obtained with a e.g. filter/mask based dialogue estimator. [091] The feasibility of deriving masks in a direct mapped system depends on the nature of the dialogue estimation processing. For example, the output of NN based generative dialogue estimators may deviate so much from the dialogue in the input signal that the derived mask “misses” the original dialogue, yielding distorted dialogue with high levels of residual background. [092] In the following, some dialogue estimators presented below rely on a mask being computed. However it is understood that a mask may also be derived from directed mapped dialogue estimators meaning that the same general processing may be applied with masks derived from direct mapped dialogue estimators. [093] The mask computation module 52 implemented in FIGS.5-10 below may accordingly generate a mask using any of the methods described, including both a masked-based system and a direct-mapped system. Other types of dialogue estimation are also within the scope of this disclosure. [094] FIGS.5-10 show various examples of first (input) dialogue estimate processors 5a. However, it is understood that the second (output) dialogue estimate processors 5b may be similar or even identical to the first dialogue estimate processor 5a. Notably, when the second dialogue estimate processor 5b is placed downstream of the adaptive upmixer 4 the second dialogue estimate processor 5b may be identical to the first dialogue estimate processor 5a. On the other hand, when the adaptive dowmixer 2 downmixes from e.g. a stereo channel input mix signal MIN to a mono channel downmix signal MD and the second dialogue estimate processor 5b is located downstream of the adaptive downmixer 2, but upstream of the adaptive upmixer 4 the second dialogue estimate processor 5b may be configured for processing a single element signal whereas the first dialogue estimate processor 5a is configured for a multi-element input signal matching the spatial configuration of the mix input signal MIN or at least a subset of the elements of the input mix signal MIN (such as the LRC triplet of a 5.1 mix input signal MIN).
[095] FIG.5 is a block diagram illustrating a two element dialogue estimate processor 501 using two single-element dialogue estimators 51:1, 51:2. The dialogue estimate processor 501 is configured for processing a stereo channel input mix signal MIN and output a stereo channel first dialogue estimate DE1. The dialogue estimation processing is here done independently in the left channel and in the right channel by using two independent instances of a single-channel dialogue estimator 51:1, 51:2 (i.e., two instances of a single-element dialogue estimator). Each dialogue estimator 51:1, 51:2 may be a direct-mapped dialogue estimator or mask-based dialogue estimator. [096] Weaknesses of this type of dialogue estimation processing include suboptimal dialogue-to-non-dialogue separation when the dialogue of the input mix signal MIN is panned near center. Another weakness is that the estimation errors in the respective single-channel estimators 51:1, 51:2 is independent. In practice, this may manifest itself as spatial instability of the non-dialogue and of the dialogue. In terms of dialogue-to-non-dialogue separation, the performance is good for dialogue that is panned left or right, and worse for center panned dialogue (the properties of the non-dialogue have to be considered as well when evaluating dialogue-background separation and these statements are true for non-dialogue signals which are uncorrelated across channels, which is a fair approximation in many practical scenarios). [097] Each single channel estimator 51:1, 51:2 may comprise a gain mask calculator and gain mask applicator. [098] The dialogue estimate processor 501 of FIG.5 may also operate in the same way on an input mix signal MIN comprising two audio objects or on an input mix signal MIN comprising one audio object and one audio channel. In general, each audio element of the input mix signal is processed with a respective single-channel dialogue estimator 51:1, 51:2. [099] The dialogue estimator 501 of FIG.5 may be generalized to N channel inputs and outputs (or M audio object input and output, or ^^^^ + ^^^^ object and/or channel inputs and outputs) as illustrated in FIG. 6 for ^^^^ input elements, wherein ^^^^ = ^^^^ + ^^^^. [100] The K-element dialogue estimate processor 502 using K-single-element dialogue estimators 51:1, 51:2, …, 51:K is shown. The dialogue estimate processor 502 comprises K instances of a single-element dialogue estimator 51:1, 51:2, … 51:K and each single-element dialogue estimator 51:1, 51:2, … 51:K obtains as an input a respective input channel/object of the input mix signal MIN and outputs an associated dialogue estimate channel/object. [101] Turning to FIG.7, a block diagram illustrating a first dialogue estimate processor 511 according to some implementation is shown. The first dialogue estimate processor 511 of FIG.7 is configured to obtain an input mix signal MIN comprising two elements (i.e. channels and/or objects) and output two dialogue estimation elements as the first dialogue estimate DE1.
The dialogue estimate processor 511 accomplishes this using a single-element mask computation module 52. [102] The first dialogue estimate processor 511 downmixes the two channels of the input mix signal MIN to a mono signal MONO1 which is provided to a single-channel (single element) mask computation module 52. The single element mask computation module 52 generates a gain mask for removing, or at least attenuating, non-dialogue content in the mono signal MONO1. The gain mask is subsequently provided to two respective mask application modules 56a, 56b which apply the (same) computed gain mask to each of the two channels of the input mix signal MIN so as to generate two dialogue estimate channels that form the first dialogue estimate signal DE1. [103] A generalization of the first dialogue estimate processor 511 of FIG.7 is illustrated in FIG.8. Here, the first dialogue estimate processor 512 is configured to obtain an input mix signal MIN comprising N channels (or equivalently e.g., M audio objects, or N channels and M objects). Each of the N channels are provided to a multi-channel downmixer 53 which downmixes the N channels to a single mono signal MONO1 which in turn is provided to a single- channel (single element) mask computation module 52. The single-channel mask computation module 52 generates a gain mask for removing or at least attenuating non-dialogue content in the mono signal MONO1. The gain mask is subsequently provided to a multi-channel mask application module 54 which applies the (same) computed gain mask to each of the N channels of the input mix signal MIN to obtain corresponding dialogue estimate channels that form the first dialogue estimate signal DE1. [104] Benefits of the first dialogue estimate processors 511, 512 of FIG.7 and FIG.8, compared to the channel (or element) independent dialogue estimate processors 501, 502 of FIG. 5 and FIG.6, include lower computational complexity, assuming the downmixing and mask application is of much lower complexity than the mask computation, and better spatial stability. [105] However, regarding spatial fidelity, as for all filter/mask based dialogue estimators, any residuals of the non-dialogue in the output of the estimator are modulated (shaped) in time and frequency by the time-varying mask (the mask will follow the dialogue time and frequency envelopes), creating so called ghost voice artifacts. Ghost voice artifacts are also present in single-channel systems or single-element systems. [106] Ghost voice artifacts may be perceived as a coarse quality to the dialogue estimate. Here, even though the same mask is being applied to all channels (or all elements), the mask will let through independent components of the background, and the ghost voice will be spatially unstable and/or be perceived as wider than the original dialogue (which typically is a point source). Another issue is spatial leakage where, e.g., an input mix signal MIN comprising a stereo
channel pair with dialogue only in the left channel will lead to a ghost voice residual in the right dialogue estimate channel, since the right channel will be shaped by the mask. In terms of dialogue-to-non-dialogue separation of an input mix signal with stereo channels, the performance is good for dialogue in phantom center, but worse for dialogue that is panned left or right, again assuming spatially uncorrelated background. [107] Turning to FIG.9, another first dialogue estimate processor 513 is depicted. This dialogue estimate processor 513 obtains three input elements and extracts a dialogue estimate of each element using only two gain mask computation modules 52 extracting a respective gain mask. To accomplish this, one of the input channels (or elements) is split and input into two instances 511a, 511b of the dialogue estimate processor 511 from FIG.7. While each of the dialogue estimate processors instances 511a, 511b takes two channels as an input, and outputs two channels, each dialogue estimate processor instance 511a, 511b utilizes a respective single- element gain mask computation module 52. In the depicted example, first dialogue estimate processor 513 of FIG.9 operates on a left, right and center (LRC) channel triplet. The left channel and the center channel are input to a first instance 511a of a two channel estimator which mixes the left and center channel to a mono signal MONO1, computes a first gain mask for dialogue estimation for the mono signal MONO1 with a gain mask computation module 52 and applies the first gain mask to each of the left channel and the center channel with respective gain applicators 56a, 56b. [108] The center channel is split and also input to the second instance 511b of the two channel dialogue estimation process alongside the right channel. The second instance 511b of the two channel estimator mixes the center and right channel to a second mono signal MONO2, computes a second gain mask for dialogue estimation for the mono signal MONO2 with a gain mask computation module 52 and applies the second gain mask to each of the left channel and the center channel with respective gain applicators 56a, 56b. [109] The output from the first and second instance 511a, 511b of the two channel dialogue estimators comprises two instances of a dialogue estimate center channel, a dialogue estimate left channel and a dialogue estimate right channel. The dialogue estimate left and right channels are output by the first dialogue estimate processor 513. The two instances of the dialogue estimate center channel are provided to a mixer 57 which combines the two dialogue estimate center channels to a combined center channel and, to avoid an unnatural boost in signal energy, the combined dialogue estimate center channel is provided to a gain modification module 58 with a gain of ½ which reduces the signal energy by a factor 4 (i.e. attenuates the signal energy with 6 dB). The gain modified dialogue estimate center channel is then output
alongside the dialogue estimate left channel and right channel to form the first dialogue estimate DE1. [110] For other types of multi-channel input mix signals the dialogue estimate techniques of international patent application published with publication number WO2023192036, hereby incorporated by reference in its entirety, may be used. [111] FIG.10 shows another first dialogue estimate processor 521 which obtains as an input an input mix signal MIN comprising a left channel, a right channel and a center channel. The first dialogue estimate processor 521 comprises a two-channel dialogue estimate processor estimator, similar to the dialogue estimate processor 511 estimator from FIG.7, which forms dialogue estimates of the left and right channel using a mixer 55, a gain mask computation module 52 and gain mask applicators 56a, 56b. The center channel is processed by a single- channel dialogue estimator 51. While the first dialogue estimate processor 521 of FIG.10 may be used for any triplet of channels (or elements) it is especially useful for the LRC front triplet of a 5.1 presentation. For example, since the center channel often carries the most important dialogue, the center channel may be processed individually. [112] The person skilled in the art will appreciate that the various first dialogue estimate processors 501, 502, 511, 512, 513, 521 described above are exemplary and that other versions are possible. For example the first dialogue estimate processors may be combined with each other and/or modified to suit the number and type of elements present in the input mix signal MIN. [113] There are many considerations when designing the dialogue estimate processor architecture. In addition to the issues of spatial stability and leakage, computational complexity is often a criterion in the design, in particular if the mono signal gain mask computation modules 52 or single-element dialogue estimators 51 are based on neural networks. For systems that are severely complexity constrained, downmix based dialogue estimation that downmixes all input channels (or elements) to a mono channel are attractive options, see e.g. the first dialogue estimate processor 512 of FIG.8. For systems with more relaxed complexity constraints, a completely independent approach may be used, see e.g. the first dialogue estimate processor 502 of FIG.6. [114] In the proposed dialogue estimation systems 10C, 10D of FIGS.3A and 4A, the first dialogue estimate processor 5a is used as pre-processing before the spatial analyzer 1, and as such the first dialogue estimate processor 5a maps the N-channel input mix signal MIN to an N- channel first dialogue estimate DE1 (or in general an ^^^^ + ^^^^-audio element input mix to an ^^^^ + ^^^^-audio element first dialogue estimate DE1). The second (output) dialogue estimate processor 5b is a single-channel (or element) system if it is performed upstream of the adaptive
upmixer 4 (see FIG.3A). In some scenarios it may be beneficial to perform the output dialogue estimation downstream of the adaptive upmixer 6 and in such scenarios it advantageous for the (output) dialogue estimate processor 5b to produce an N-channel second dialogue estimate DE2 (or an ^^^^ + ^^^^-audio element dialogue estimate output signal). The latter case is associated with higher computational complexity. [115] FIGS.11-18 illustrate examples of how the proposed dialogue estimation systems 10C, 10D can make use of the different dialogue estimate processor architectures 501, 502, 511, 512, 513, 521 presented in connection with FIGS.5-10. The exemplary systems of FIGS.11-17 all operate on channels, but it is understood that one or more (or all) channels may be replaced with audio objects. Additionally, in all exemplary systems of FIGS.11-18 a left, right and center triplet is processed, and it is noted that this triplet may be replaced with a two-channel stereo signal where the right and left stereo channels replaces the right and left channel of the LRC triplet and the center channel is set to zero. [116] Additionally, all exemplary dialogue estimation systems 101-108 of FIGS.11-18 are of generally the same type as dialogue estimation system 10C from FIG.3A, wherein the second (output) dialogue estimate processor 5b is located upstream from the adaptive upmixer 4. However, it is understood that the dialogue estimation systems 101-108 may also be used with the second (output) dialogue estimate processor 5b located downstream of the adaptive upmixer 4. [117] FIG.11 shows an exemplary dialogue estimation system 101 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel. The channels are provided to a first dialogue estimate processor 502 having three single-channel dialogue estimation modules 51:1, 51:2, 51:3 that extract an individual dialogue estimate of each of the three input channels. Accordingly, dialogue estimate processor 502 operates similar to the dialogue estimate processor from FIG.6, with K = 3. The dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1. [118] The spatial analyzer 1 extracts spatial parameters Pspatial which are provided to the optional spatial filtering module 3. The spatial filtering module performs spatial filtering of the input mix signal MIN, based on the spatial parameters Pspatial, so as to form a spatial filtered mix signal MF. The spatial filtered mix signal MF is provided to the adaptive downmixer 2 which downmixes the spatial filtered mix signal MF to obtain a downmixed (and spatial filtered) mono mix signal MDF based on the spatial parameters Pspatial. The downmixed mono mix signal is then provided to a single channel second dialogue estimate processor 5b which extracts a downmixed dialogue estimate DE2D of the (spatial filtered and downmixed) mix signal. Finally, the downmixed dialogue estimate DE2D is provided to the adaptive upmixer 4 which upmixes the
dialogue estimate DE2D to yield the second dialogue estimate DE2 which is output by the dialogue estimation system 101. [119] Compared to the dialogue estimation system 10C of FIG.3A the spatial filtering module 3 has been moved in dialogue estimation system 101 upstream of adaptive downmixer 2. The spatial filtering module 3 may operate on one or more input elements. Hereby, the spatial filtering module 3 may be placed anywhere in the signal processing chain, e.g. upstream of the adaptive downmixer 2 or downstream of the adaptive upmixer 4 (where the signals comprise at least two elements) or between the adaptive downmixer 2 and adaptive downmixer 4 (where the signal may be a mono signal with a single element). [120] FIG.12 shows another exemplary dialogue estimation system 102 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel. The channels are provided to a first dialogue estimate processor 521 having a two-channel dialogue estimate processor 511 (see FIG.7) for processing the left and right channel and a single-element dialogue estimate processor 51 for processing the center channel. Accordingly, dialogue estimate processor 521 operates as the general dialogue estimate processor from FIG.6. The dialogue estimate channels output by the dialogue estimate processor 521 form the first dialogue estimate DE1 which is provided to the spatial analyzer 1 for extraction of spatial parameters Pspatial. The rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [121] FIG.13 shows yet another exemplary dialogue estimation system 103 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel. The channels are provided to a first dialogue estimate processor 513 having two instances 511a, 511b of two-channel dialogue estimate processors for processing the left and center channel and processing the right and center channel respectively. Accordingly, dialogue estimate processor 513 operates as the dialogue estimate processor from FIG.9. The dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1 tasked with extracting the spatial parameters Pspatial. The rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [122] FIG.14 shows yet another exemplary dialogue estimation system 104 obtaining an input mix signal MIN with three channels, a left channel, a right channel and a center channel. The channels are provided to a first dialogue estimate processor 512 having a multi-channel downmixer and a single-element gain mask calculator 52. Accordingly, the first dialogue estimate processor 512 operates as the dialogue estimate processor from FIG.8 and downmixes all channels to a mono signal MONO1 for which a gain mask is calculated by gain mask calculator 52 whereafter the same gain mask is applied to each of the three channels by the mask
applicator 54 yielding the corresponding dialogue estimate channels. The dialogue estimate channels form the first dialogue estimate DE1 which is provided to the spatial analyzer 1. The spatial analyzer 1 extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [123] FIG.15 shows an exemplary dialogue estimation system 105 obtaining an input mix signal MIN with two channels, a left channel and a right channel. The two channels are provided to a first dialogue estimate processor 501 having two single-element dialogue estimators 51:1, 51:2 that form a dialogue estimate for each channel and outputs these dialogue estimate channels as the first dialogue estimate DE1. The first dialogue estimate DE1 is provided to the spatial analyzer 1 which extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [124] FIG.16 shows another exemplary dialogue estimation system 106 obtaining an input mix signal MIN with two channels, a left channel and a right channel. The two channels are provided to a two-channel dialogue estimate processor 511 identical to the dialogue estimate processor described in connection with FIG.7. The two channels are combined into a mono signal MONO1 and provided to a gain mask calculation module 52 which calculates a gain mask which is applied to each of the two channels, to yield dialogue estimate channels. The dialogue estimate channels form first dialogue estimate DE1 which is provided to the spatial analyzer 1. The spatial analyzer extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [125] FIG.17 shows an exemplary speech estimation system 107 obtaining an input mix signal MIN with nine channels. The nine channels are in this implementation the nine non-LFE channels of a 5.1.4 presentation. That is, the nine channels comprising the left channel (L), the right channel (R), the center channel (C), the left surround channel (Ls) and the right surround channel (Rs) of the horizonal listening plane. The nine channels further comprising the top front left (Tfl) channel, top front right (Tfr) channel, top back left (Tbl) channel and top back right (Tbr) channel forming the four height channels of the 5.1.4 presentation. All nine channels are provided to dialogue estimate processor 514 implementing three instances of multi-channel dialogue estimate processors 512a, 512b, 512c, each being equivalent to the multi-channel dialogue estimate processor of FIG.8 and receiving a unique subset of the nine channels. [126] Notably, the first multi-channel dialogue estimate processor 512a obtains the four height channels, mixes them to a same mono signal MONO1, generates a gain mask for the mono signal MONO1 and applies the same gain mask to all four height channels to obtain the associated dialogue estimate channels. The second multi-channel dialogue estimate processor 512b obtains the two surround channels (Ls, Rs), mixes them to a mono signal MONO2,
generates a gain mask for the mono signal MONO2 and applies the same gain mask to each surround channel to obtain the associated dialogue estimate channels. The third multi-channel dialogue estimate processor 512c obtains the LRC-triplet channels, mixes them to a mono channel MONO3, generates a gain mask for the mono signal MONO3 and applies the same gain mask to each of the L, R and C channels to obtain the associated dialogue estimate channels. The dialogue estimate channels together form the first dialogue estimate DE1 which is provided to the spatial analyzer 1. The spatial analyzer 1 extracts spatial parameters Pspatial and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [127] FIG.18 shows an exemplary dialogue estimation system 108 obtaining an input mix signal MIN with five channels (in this case the L, R, C, Rs and Ls channels of a 5.1 presentation) and four audio objects (objects 1 - 4). The dialogue estimate processor 514 is however the same dialogue estimator as in speech estimate system 107, with the only difference being that instead of the first multi-channel dialogue estimate processor 512a obtaining the (four) height channels it now obtains the (four) audio objects and extracts the dialogue estimate of each audio object using a same gain mask for each object wherein the gain mask is extracted from a mono downmix of all four audio objects. The dialogue estimate channels and objects forming the first dialogue estimate DE1 are provided to the spatial analyzer 1 which extracts spatial parameters Pspatial, and the rest of the processing is identical to that of the dialogue estimation system 101 from FIG.11. [128] In common for all dialogue estimation systems 10C-D, 101-108 is that the first dialogue estimate DE1 is to be mapped to spatial parameters Pspatial which are used to control the adaptive downmixing, adaptive upmixing and optionally the spatial filtering. [129] There are several options for how to perform spatial analysis and determine the spatial parameters Pspatial from the dialogue estimate signals. Below, two examples of spatial analysis are given: correlation based spatial analysis, and amplitude ratio based spatial analysis. Both examples result in panning parameters ^^^^^^^^ and panning parameters ^^^^^^^^ is one example of spatial parameters Pspatial. It is envisaged that panning parameters may be calculated using other methods as well. [130] Furthermore, it is envisaged that other types of spatial parameters Pspatial, such as one or more of inter-channel phase difference, inter-channel time difference, inter-channel level difference and magnitude may be extracted in addition to, or alternative to, the panning parameters ^^^^^^^^. For example, the adaptive downmixing, upmixing and spatial filtering described in International Application No. PCT/US23/63717 uses the inter-channel phase difference, a panning parameter and the magnitude as well as indicators of the average value and spread of the panning and inter-channel phase difference labelled thetaMid, thetaWidth, phiMid and phiWidth.
It is envisaged that these parameters and the associated processing may be applied in the adaptive upmixing, adaptive downmixing and optional spatial filtering stage of the present disclosure. [131] The working example will be an application where the input mix signal MIN is the LRC triplet of a 5.1 presentation. However, it is understood that the presented methods generalize straightforwardly to signals with a higher channel count, higher object count and signals that are a combination of channels and objects. Both methods yield a set of panning parameters in the form of one scalar value per channel/object, and these may be used in the adaptive downmixing and upmixing. The channels and/or objects of the first dialogue estimate DE1 that is output from the first dialogue estimate processor 5a are denoted X = �^^^^1(f), ^^^^2(f), … , ^^^^N(f)� where e.g., ^^^^1 is a dialogue estimate of the first element (a channel or an object), ^^^^2 is an estimate of the second element (a channel or an object) etc., and ^^^^ is a frequency index (between 0-2048 for a 4096 point DFT). The processing described below is done on a group of frequency indices that are grouped together in a processing band forming a TF tile. [132] In correlation based spatial analysis a static mono downmix M is calculated for the first dialogue estimate DE1 as
for all ^^^^ in a current processing band. Analogous to X =�^^^^1(f), ^^^^2(f), … , ^^^^N(f)� the channels and/or objects of the input mix signal MIN are denoted Y =�^^^^1(f),^^^^2(f), … ,^^^^N(f)� allowing the correlation ^^^^^^^^^^^^^^^^ between M and each channel/object of Y to be calculated as
Alternatively, it is possible to compute ^^^^^^^^^^^^^^^^ instead of ^^^^^^^^^^^^^^^^, and ^^^^^^^^^^^^^^^^ may be determined as
but in the following ^^^^^^^^^^^^^^^^ according to equation 2 is used. [133] In some implementations, it is advantageous to average or smooth ^^^^^^^^^^^^^^^^ over a plurality of neighboring frames. [134] A panning parameter ^^^^^^^^ may now be calculated for each of the N channels (or elements) as
and the panning parameter ^^^^^^^^ is, as mentioned above, one example of a spatial parameter Pspatial.
[135] To illustrate the correlation based spatial analysis a situation is considered where the input mix signal MIN, labeled Y, comprises a mix of a dialogue component X' and a non- dialogue component W such that Y = X' + W. It is also assumed that the dialogue component X' is uncorrelated with the non-dialogue component W. It will further be assumed that X' is a point source created by panning a mono signal d(f) such that ^^^^′ =
2^^^^(^^^^), … ,^^^^^^^^^^^^(^^^^) ) where ^^^^^^^^ is the panning coefficient for element n and that the panning is energy preserving. Energy preserving panning means that
By further assuming that the dialogue estimate X is a perfect capture of the dialogue component, meaning that X = X', it is found that
Furthermore, since d and W are uncorrelated it is found that
By inserting ^^^^^^^^^^^^^^^^ as defined in equation 8 and 9 into the panning parameter definition in equation 4 it is found that ^^^^^^^^ = ^^^^^^^^. [136] A benefit with correlation based spatial analysis is that it is robust against “energy leakage” between elements of the dialogue estimate channels and/or objects X. [137] In amplitude ratio based spatial analysis the panning parameter ^^^^^^^^ for element n is instead calculated as
and it is especially noted that in equation 10 only the dialogue estimate X = ^^^^1(f), ^^^^2(f), ... , ^^^^k(f), … , ^^^^N(f) is used and not the original input mix signal.
[138] Again using the above assumptions (i.e. the input mix signal MIN, labeled Y, fulfills Y = X' + W, wherein X' is uncorrelated with the non-dialogue component W, and wherein X' is a point source created by panning a mono signal d(f) such that ^^^^′ = (^^^^), … ,^^^^^^^^^^^^(^^^^) ) where ^^^^^^^^ is the panning coefficient for element n and that the panning is energy preserving) it is found that the numerator in equation 10 fulfills
And that the denominator in equation 10 fulfills
Accordingly, ^^^^^^^^ = ^^^^^^^^. [139] Turning to the adaptive downmixer 2 and adaptive upmixer 4, some examples of how these modules may operate will now be described. In general there is a lot of freedom in how the adaptive downmix MD may be determined. In the following an example is given that is based on the panning parameters ^^^^^^^^ obtained by the spatial analysis described above (correlation based or amplitude ratio based). The adaptive downmix signal ^^^^^^^^ (^^^^) may be computed as
where ^^^^^^^^ is the adaptive downmix coefficient for element (a channel or an object) n. The upmixing computes the output dialogue estimate for element n, ^^�^^^^^^(^^^^) as ^^�^^^^^^(^^^^) = ^^^^^^^^^^^^^^^^(^^^^) (14) where the upmix coefficient
[140] As an illustrative example, a situation is considered when the input mix signal MIN, labeled Y contains no background and only a panned dialogue source whereby ^^^^ = (^^^^1^^^^(^^^^),^^^^2^^^^(^^^^), … ,^^^^^^^^^^^^(^^^^) ) (16) and thus ^^^^^^^^ = ^^^^^^^^^^^^(^^^^). (17) The adaptive downmix becomes
And the output for element n evaluates to
^^^^ ^^^^ ^^�^^ (^^^^) = ^^^^ ^^^^ (^^^^) ^^^^ ^^^^ ^^^^ ^^^^ = ∑^^^^ ^^^^ ^^^�^^^^^^^^^^^^^^^^^^^^(^^^^) . (19) ^^^^=1 ^^^^ ^^^^^ ^^^^=1 Assuming perfect spatial analysis, i.e. ^^^^^^^^ = ^^^^^^^^, it is found that ^^�^^^^^^(^^^^) = ^^^^^^^^^^^^(^^^^). (20) [141] The choice of ^^^^^^^^ may be made based on the background suppression performance of the adaptive downmix. The downmix coefficients may be chosen like ^^^^^^^^ = ^^^^(^^^^^^^^), where ^^^^ is a non-decreasing mapping, e.g., ^^^^^^^^ = ^^^^^^^^ ^^^^ , ^^^^ > 0. For example for ^^^^ = 1, ^^^^^^^^ = ^^^^^^^^, ^^^^^^^^ = ^^^^^^^^ (assuming energy preserving panning parameters). As another example ^^^^ = 2, ^^^^^^^^ =
^^^^^^^^⁄ ∑^^^^ ^^^^=1 ^^^^3 ^^^^ . Of course, these are only examples, and it is understood that other definitions of ^^^^, ^^^^^^^^ and ^^^^^^^^ are possible. For instance if ^^^^^^^^ = 1 for all ^^^^ it is found that ^^^^^^^^ = ^^^^^^^^⁄ ∑^^^^ ^^^^=1 ^^^^^^^^ . [142] As an alternative to equation 13 it is possible to calculate the adaptive downmix based on the dialogue estimate DE1 of the first dialogue estimate processor 5a as
While this may provide good background suppression, it may lead to over-suppression of the non-dialogue and loss of energy and timbre changes in the dialogue estimate. One reason for this is that in this case, two instances of dialogue estimate processors are used in cascade to produce the output. Another reason is that the first dialogue estimate processor 5a may over-suppress because it “sees” an input mix signal MIN with more non-dialogue content. This should be compared to the case where the adaptive downmix is computed from the input mix signal MIN, here the adaptive downmix will increase the dialogue-to-non-dialogue level, and thus make the second dialogue estimator “see” a signal with less non-dialogue content. This effect may be particularly pronounced if the dialogue estimate processors are based on NNs. [143] Each of the proposed dialogue estimation systems 10C-D, 101-108 constitutes a self-contained audio processing system which obtains a multi-element input MIN and outputs a dialogue estimate DE2 having the corresponding elements with dialogue extracted. However, in some scenarios it may be beneficial to integrate anyone of the dialogue estimation systems 10C- D, 101-108 with a dialogue detector 6 as shown in FIGS.19-23. [144] FIG.19 is a block diagram illustrating a dialogue estimation system 201 provided with a dialogue detector 6. The dialogue detector 6 produces a sequence of dialogue confidence values, typically one value per processing frame (which typically is 20-100 ms long). The confidence values are then post-processed by post processor 7, and a temporal soft gate gain is computed by soft gate computation module 8. The soft gate gain is then applied by soft gate applicator 9 to the second dialogue estimate DE2 to form a gated second dialogue estimate
DE2G. The soft gate gain consists of one gain per audio element (and per processing frame), the gains vary over time as a function of the output from the dialogue detector 6. [145] The dialogue detector 6 may produce a binary valued output (e.g.0 or 1) or a continuous valued output, typically in the interval [0,1], indicating the probability of dialogue being present in the current time segment. The output may also be referred to as the dialogue confidence. The post-processing performed by the post processor 7 may comprise smoothing of the dialogue confidence so as to e.g. bridge “gaps” (short segments of confidence 0 or very low confidence surrounded by segments of high confidence) or remove “spikes” (short segments of high confidence surrounded by low confidence). The soft gate computation performed by the soft gate computation module 8 may consist of a non-decreasing mapping of the post processed dialogue confidence to a soft gate gain. In the soft gate computation module 8 it may also be beneficial to add ramps to the attack and release phase of detected dialogue, in particular if the dialogue confidence is binary (0 or 1). [146] The input to the dialogue detector may be the signal from the dialogue estimation system 10C-D, 101-108 at any point of the main signal chain. The block diagram in FIG.19 illustrates the flexibility in choosing from where to get the input to the dialogue detector 6 (see the dashed lines). For example, the input mix signal MIN, the adaptive downmix signal MD, MDF before or after spatial filtering, the downmixed dialogue estimate DE2D or the second dialogue estimate DE2 after upmixing may be provided as input to the dialogue detector 6. [147] Additionally, it is further noted that the soft gate applicator 9 may be done at any point of the main signal chain (although it advantageously occurs after the input to the dialogue detector). In FIG.19 the soft gate applicator 9 is arranged downstream of the adaptive upmixer 4 and an alternative dialogue estimate system 202 is shown in FIG.20, where the soft gate applicator 9 is arranged upstream of the adaptive upmixer 4, downstream of the second dialogue estimator 5b. Applying the soft gate to a mono signal (as in FIG.20) may save complexity compared to applying the gate to a multi-channel/object signal (as in FIG.19). [148] The dialogue detector 6 may handle multi-channel (or multi-element) input in a way similar to the dialogue estimate processors 5a, 5b. For example, the dialogue detector 6 may rely on downmixing and/or separate processing of the individual channels/objects. [149] In some implementations, often for reasons of computational complexity, the dialogue detector 6 is operating on a mono signal, and the processing performed by the post processor 7 and the soft gate computation module 8 results in a single time-varying soft gate gain. This same soft gate gain is applied to all the channels/objects by the gate applicator 9. The mono signal input to the dialogue detector 6 may be computed as a downmix of the multi-
channel/object signal from some place in the main signal path. The channels/objects may e.g. be obtained upstream of the adaptive downmixer 2. [150] When input mix signal MIN is the non-LFE channels of a 5.1 signal, it is understood that dialogue in the surround channels Ls and Rs of the 5.1 input mix signal MIN is uncommon but occurs sometimes. Hereby, it may be advantageous to attenuate Ls and Rs with an attenuation gain prior to downmixing these channels together with the LRC channel triplet to increase the SNR of the mono signal that is input to the dialogue detector 6, while still being able to capture dialogue in these channels. The attenuation gain with which Ls and Rs are attenuated may be a predetermined attenuation gain such as 6 dB. However, this attenuation gain is merely exemplary, and other predetermined attenuation gains may be used. [151] If computational complexity allows, another option is to run an individual dialogue detector 6 per channel/object. This yields a set of per-channel/object confidence values {^^^^1, ^^^^2, … , ^^^^^^^^}. The set of per-channel/object dialogue confidence values {^^^^1, ^^^^2, … , ^^^^^^^^} may then be combined into a single dialogue confidence value ^^^^ e.g. as ^^^^ = max ^^^^^ . (22) n ^^^ This will typically yield a more “permissive” dialogue detection, and more of the second dialogue estimate DE2 will be let through the soft gate applicator 9 (which in this case applies the same soft gate gain to all output channels/objects). [152] Yet another option, even more computationally complex, is to configure post processor 7 to output one unique dialogue confidence value per audio element, and to configure soft gate computation module 8 to output one unique gain per element. The soft gate application module 9 may then apply an individual soft gate gain to each element. [153] Yet another option is to combine the single-element dialogue detection with the downmix based dialogue detection discussed above. As an example, a dialogue estimation system 203 with dialogue detection is shown in FIG.21. [154] The dialogue estimation system 203 obtains as an input a 5.1 input mix signal MIN. The dialogue detection is split into three individual paths, one path for the center channel resulting in a soft gate gain being applied to the center channel in soft gate applicator sub- module 91, a second path for the left and right channel pair resulting in a same soft gate gain being applied to the left and right channels in soft gate applicator sub-module 92 and a third path for the left and right surround channel pair resulting in a same soft gate gain being applied to the left and right surround channels in soft gate applicator sub-module 93. [155] That is, the dialogue detector 6 comprises a plurality of dialogue detector sub- modules 61, 62, 63 and downmixers 64, 65. Dialogue detector sub-module 63 obtains the center channel and produces a dialogue confidence value for each frame of the center channel alone.
The L and R channels are downmixed with downmixer 65 to a mono channel which is provided to dialogue detector sub-module 62 which produces a dialogue confidence value for each frame of the mono channel downmix of the L and R channels. Similarly, the surround channels Ls and Rs are downmixed with downmixer 64 to a mono channel which is provided to dialogue detector sub-module 61 which produces a dialogue confidence value for each frame of the mono channel downmix of the Ls and Rs channels. [156] The dialogue confidence values output by each respective dialogue detector sub- module 61, 62, 63 is provided to the post-processor 7 which e.g. smooths each respective stream of dialogue confidence values and the soft gain computation module 8 calculates a soft gate gain for each respective stream of dialogue confidence values. The soft gate gain applicator module 9 comprises soft gate gain applicator sub-modules 91, 92, 93 for applying the soft gate gains output by the soft gain computation module to the respective channels. Here, a unique soft gate gain is applied to the center channel in soft gate gain applicator sub-module 93 and the soft gate gain calculated from the L and R channel downmix mono signal is applied to both the L and R channel in the soft gate gain applicator sub-module 92. Similarly, the soft gate gain calculated from the Ls and Rs channel downmix mono signal is applied to both the Ls and Rs channel in the soft gate gain applicator sub-module 91. [157] FIG.22 is a block diagram showing yet another dialogue estimation system 204 with dialogue detection. Dialogue detection may be enhanced by combining the dialogue confidence extracted directly from the input mix signal MIN using dialogue detector 6a with a so- called alternative dialogue confidence, ^^^^^^^^^^^^^^^^, derived from e.g. the dialogue estimation processing performed in the second dialogue estimation processor 5b using dialogue detector 6b. Examples of how to determine the alternative dialogue confidence may be found in International Application No. PCT/US24/14225, hereby incorporated by reference in its entirety. [158] The alternative dialogue confidence ^^^^^^^^^^^^^^^^ may in FIG.22 be computed as, for example, the ratio of the energy of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy of the downmix signal MDF input to the dialogue estimator 5b. In the following, E(•) denotes the norm of a signal and we will use the (squared) L2-norm (i.e. the energy) in the following. However, it should be understood that any norm could be used in the computation of the alternative dialogue confidence, e.g., the 1-norm (absolute value norm), the infinity-norm (max norm) to mention two additional examples well-known to those skilled in the art. Proceeding with the squared L2-norm as an example, E(DE2D) is the energy of the downmix signal DE2D. As appreciated by those skilled in the art, the computation of the energy of a signal S may be done using various methods depending e.g. on the domain in which the signal S is represented. For example, if the signal S is in time domain, its energy E(S) is
proportional to the sum of the squared signal S (i.e. the sum of each squared time domain sample). As another example, if a signal S is represented in a transform domain (e.g. MDCT domain or QMF domain), its energy E(S) may be found by summation of transform samples (e.g., after taking the absolute value and squaring) across frequency or across both time and frequency (e.g. summation over an entire frame). [159] That is, ^^^^^^^^^^^^^^^^ is in FIG.22 based on the energy ratio E(DE2D)/E(MDF) which is typically a value in the interval [0,1]. It is also noted that in some implementations this value may fall outside the interval [0,1] and bounding it to be within this interval may be advantageous. [160] This computation of ^^^^^^^^^^^^^^^^ may be further refined by applying frequency weighting prior to computing the energies, and by applying a non-decreasing mapping to the energy ratio E(DE2D)/E(MDF). The alternative dialogue confidence values ^^^^^^^^^^^^^^^^ from dialogue detector 6b are combined with the dialogue confidence values from dialogue detector 6a into a single dialogue confidence value by the dialogue confidence combiner module 67. Examples of how to determine the single, combined, dialogue confidence value are also found in International Application No. PCT/US24/14225. [161] It is further noted that the alternative dialogue confidence ^^^^^^^^^^^^^^^^ may be computed as the ratio of the energy E(DE2D) of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy E(MD) of the downmix signal MD input to the spatial filtering module 3. As a further example, the alternative dialogue confidence ^^^^^^^^^^^^^^^^ may be computed as the ratio of the energy E(DE2D) of the downmix signal DE2D that is output from the second dialogue estimator 5b to the energy (MIN) of the input mix signal MIN. That is, ^^^^^^^^^^^^^^^^ may be determined by the dialogue detector 6b based on the energy ratios E(DE2D)/E(MD) or E(DE2D)/E(MIN). This is indicated with the dashed lines in FIG.22 providing the downmix signal MD and the input mix signal MIN to dialogue detector 6b, allowing the dialogue detector 6b to compute the signal energies and energy ratios. [162] Hereby, the alternative dialogue confidence ^^^^^^^^^^^^^^^^ may be based on one or more of the energy ratios E(DE2D)/E(MDF), E(DE2D)/E(MD) or E(DE2D)/E(MIN). For example, a single alternative dialogue confidence ^^^^^^^^^^^^^^^^ may be determined based on both the energy ratio E(DE2D)/E(MDF) and the energy ratio E(DE2D)/E(MD). [163] In this example, DE2D, MD and MDF are all mono signals meaning that computation of the energy ratios E(DE2D)/E(MD) and E(DE2D)/E(MDF) is straightforward. However, MIN is a multi-element signal meaning when the alternative dialogue confidence ^^^^^^^^^^^^^^^^ is to be based energy ratio E(DE2D)/E(MIN) the energy of the input mix signal MIN may be defined by combining the energy of each individual element in the input mix signal MIN using e.g. an
average-, min- or max-function. For example, E(MIN) = (E(MIN1) + E(MIN2) + … + E(MINk))/k where MIN1, MIN2, …, MINk are the ^^^^ = ^^^^ + ^^^^ elements of the input mix signal MIN. [164] It is understood that with the proposed dialogue estimate systems 10C-10D, 101- 107, 201-204 there is flexibility in how the alternative dialogue confidence value ^^^^^^^^^^^^^^^^ is computed. In the above, it is described that up to three alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be computed from signals between the adaptive downmixer 2 and the adaptive upmixer 4 (based on the energy ratios E(DE2D)/E(MDF), E(DE2D)/E(MD) or E(DE2D)/E(MIN)) when the second dialogue estimate processor 5b is located upstream of the adaptive upmixer 4. Additionally, or alternatively, up to ^^^^ + ^^^^ additional dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined based on the energies of signals input to, and output from, the first dialogue estimate processor 5a. For example, a ratio between the energy of each respective element output from, and input to, the first dialogue estimate processor 5a may be determined. [165] Accordingly, it is envisaged that up to 3 + ^^^^ + ^^^^ alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined in the dialogue estimation system 204 of FIG.22. [166] In some implementations, the second dialogue estimate processor 5b is placed downstream of the adaptive upmixer 4 (see e.g. FIG.4A) and in such implementations 3(^^^^ + ^^^^) alternative dialogue confidence values ^^^^^^^^^^^^^^^^, i.e. three per channel/object, may be computed. More specifically, up to ^^^^ + ^^^^ dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined based on the energy of the elements output from, and input to, the first dialogue estimate processor 5a and, similarly, up to ^^^^ + ^^^^ additional alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined based on the energy of each element output from, and input to, the second dialogue estimate processor 5b. Additionally, a further ^^^^ + ^^^^ alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the mix input signal MIN making for a total of 3(^^^^ + ^^^^) alternative dialogue confidence values. [167] Furthermore, when the second dialogue estimate processor 5b is located downstream of the adaptive upmixer 4, an additional ^^^^ + ^^^^ alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the downmix signal MD. [168] Similarly, when the second dialogue estimate processor 5b is located downstream of the adaptive upmixer 4, another ^^^^ + ^^^^ alternative dialogue confidence values ^^^^^^^^^^^^^^^^ may be determined, based on the ratio of the energy of each element output from the second dialogue estimate processor 5b to the energy of the corresponding element in the downmix signal MDF, output by the spatial filtering module 3.
[169] In general, it is understood that an alternative dialogue confidence value ^^^^^^^^^^^^^^^^ may be determined based on the energy of each element output by the second dialogue estimate processor 5b and the energy of any corresponding element upstream of the second dialogue estimate processor 5b. The same applies also for each element output by the first dialogue estimate processor 5a, where an alternative dialogue confidence value ^^^^^^^^^^^^^^^^ may be determined based on the energy of each element output by the first dialogue estimate processor 5a and the energy of any corresponding element upstream of the first dialogue estimate processor 5a. [170] The one or more computed alternative dialogue confidence ^^^^^^^^^^^^^^^^ may be combined (e.g. using an average function, max function or min function) in various ways with each other and/or with the dialogue confidence from the dialogue detector 6a. [171] For example, for a dialogue estimation system with the second dialogue estimator 5b located downstream of the adaptive upmixer the ^^^^ + ^^^^ alternative dialogue confidence values extracted from the first dialogue estimate processor 5a inputs and outputs are labeled
^^^^ ≤ ^^^^ + ^^^^, and the ^^^^ + ^^^^ alternative dialogue confidence values based on the inputs and outputs of the second dialogue estimate processor 5b are labeled ^^^^ ≤ ^^^^ + ^^^^ for a total of 2(^^^^ + ^^^^) alternative
, , … , ^^^^(^^^^^^^^) ^^^^ , ^^^^(
1 , , … , }. These confidence values be combined into a single alternative dialogue confidence value, ^^^^^^^^^^^^^^^^, for example by taking the maximum over all input and output confidence values as
[172] Furthermore, as mentioned above, it is possible to compute additional ^^^^ + ^^^^ alternative dialogue confidence values based on the energies of the elements in the input mix signal MIN and the energies of the elements output of the second dialogue estimate processor 5b whereby the 2(^^^^ + ^^^^) alternative dialogue confidence values may be expanded with further (^^^^ + ^^^^) alternative dialogue confidence values, for a total of 3(^^^^ + ^^^^) alternative dialogue confidence values. [173] There is also flexibility in how signals are combined prior to computing the alternative dialogue confidence value(s). While at least one alternative dialogue confidence value may be computed for each element for each of the two dialogue estimate processors 5a, 5b it is envisaged that two or more elements are grouped together for the purpose of computing a single alternative dialogue confidence value. [174] For example, an input/output signal pair associated with one channel/object (e.g. the left channel) may be grouped with an input/output pair associated with another channel/object (e.g. the right channel). For each such group, a mono downmix is computed for all the input
channels/objects belonging to the group, and similarly a mono downmix is computed for all output channels/objects belonging to the group, and an alternative dialogue confidence value for the group is computed based on the energy of the mono downmix of the input and energy of the mono downmix of the output. [175] One example of how this may be realized is shown in the dialogue estimation system 205 of FIG.23, where elements are grouped for the first dialogue estimate processor 5a. Here three groups are defined, a first group is associated with the first dialogue estimator 51:1 processing the C channel. The input and output of the dialogue estimator 51:1 are provided to alternative dialogue confidence computation module 6b which extracts a first alternative dialogue confidence value by calculating the energy of the input and output and computing the ratio E(output)/E(input). The second group is associated with the first dialogue estimator 51:2 processing the L and R channel, and the L and R input channels are downmixed in downmixer 68 and the L and R output channels are downmixed in downmixer 69. A second alternative dialogue confidence value is then computed based on the energy of the mono downmix from downmixer 68 and the energy of the mono downmix from downmixer 69, by alternative dialogue confidence computation module 6c. More specifically, alternative dialogue confidence computation module 6c calculates the ratio between the energy of the mono downmix from downmixer 69 to the energy of the mono downmix from downmixer 68. [176] Finally, a third group is associated with the second dialogue estimate processor 5b which processes the adaptive downmix. The input and output of the second dialogue estimate processor 5b are provided to alternative dialogue confidence computation module 6d which extracts a third alternative dialogue confidence value based on the energy ratio E(DE2D)/E(MDF) as described in FIG.22. The alternative dialogue confidence values from the three groups are combined, for example by a max operation, in the alternative dialogue confidence combiner module 67a and the combined alternative dialogue confidence value is then combined with the dialogue confidence from the dialogue detector 6a in dialogue confidence combiner module 67b. The resulting combined dialogue confidence value is provided to the post processor 7. [177] Of course, similar grouping of elements for the purposes of computing alternative dialogue confidence may be performed for the second dialogue estimate processor 5b, if it is arranged downstream of the adaptive upmixer 4. [178] The output of dialogue estimation systems 10C-D, 101-108, 201-205 is a dialogue estimate DE2. That is, isolated dialogue from the input mix signal MIN, which may also be soft gated based on the dialogue confidence of a dialogue detector. The dialogue estimate output may be used, e.g. provided for playback, as it is, however, in some implementations the input mix
signal MIN is added to the dialogue estimate DE2 for making the dialogue estimate DE2 sound more natural. [179] For example, in cases where the non-dialogue components of the input mix signal MIN contain useful information, for example narrative elements of a movie soundtrack, the input mix signal MIN is mixed at a higher level. In cases where the non-dialogue components are considered interference, for example in some voice communication applications, the input mix signal MIN is mixed at a comparatively lower level to make the dialogue estimate DE2 more natural sounding (compared to using only the dialogue estimate). [180] FIG.24 shows one example of a dialogue enhancement system 301 comprising a dialogue estimation system 10C, dialogue estimate gain module 11, an input mix signal gain module 12 and a mixer 13. In the exemplary dialogue enhancement system 301 of FIG.24 the dialogue estimation system 10C is used however it is envisaged that the dialogue estimate system 10C of FIGS.24-28 may be replaced with any of the dialogue estimation systems 10D, 101-108, 201-205 described herein. [181] The dialogue estimation system 10C obtains the input mix signal MIN as an input, processes the input mix signal MIN in accordance with the above and outputs a dialogue estimate DE2 (or possibly a soft gated dialogue estimate DE2G). The dialogue estimate DE2 may be referred to as a dialogue signal in the context of dialogue enhancement systems. [182] The dialogue estimate DE2 is provided to the dialogue estimate gain module 11 which applies a dialogue gain to dialogue estimate DE2 forming a gain modified dialogue estimate DE2Gain. The dialogue gain may be any value and may be an attenuating gain or a boosting gain. [183] The input mix signal MIN is also provided to the input mix signal gain module 12 which applies a mix gain to the input mix signal MIN yielding input mix signal MGain. The mix gain may also be any value and may be a boosting gain or an attenuating gain. For example, the mix gain is constant and equal to 1, meaning the input MIN to the input mix signal gain module 12 is equal to the output MGain. The gain modified dialogue estimate DE2Gain and input mix signal MGain are provided to mixer 13 which mixes the two signals to form a final mix output signal DE2Final containing the gained dialogue estimate DE2Gain mixed with the gained input mix signal MGain. The final mix output signal DE2Final is the output of the dialogue enhancement system 301. [184] For example, in linear domain and under the assumption that the mix gain is 1, negative values of the dialogue gain larger than or equal to -1 lead to attenuation of dialogue in the output DE2Final from the mixer 13, positive values lead to amplification (boosting) of dialogue in the output DE2Final, and a dialogue gain of zero yields DE2Final = MIN.
[185] A dialogue gain of -1 may lead to complete silencing of dialogue in DE2Final if DE2 is equal to the dialogue component of MIN. However, in practice this occurs only rarely since some residual dialogue may be present in DE2Final due to the dialogue estimation system 10C not being ideal, although the residual dialogue is often inaudible in the presence of other background audio. A dialogue gain exceeding zero will add fractions of the dialogue DE2 to MIN = MGain, resulting in an amplification of the dialogue. [186] FIG.25 shows another exemplary dialogue enhancement system 302, similar to the dialogue enhancement system 301 of FIG.24, which also forms an output dialogue estimate DE2Final based on the input mix signal MIN. A difference between the dialogue enhancement system 302 and the dialogue enhancement system 301 is that dialogue enhancement system 302 of FIG.25 comprises a background estimator 14 upstream of the input mix signal gain module 12. The background estimator 14 extracts a background signal MB (i.e. any non-dialogue content) from the input mix signal MIN by calculating the difference between the input mix signal MIN and the dialogue estimate DE2 as MB = MIN – DE2. The background signal MB is provided to the input mix signal gain module 12 which applies the mix gain to form a gain modified background signal MBGain. [187] Assuming the mix gain is 1, a dialogue gain of 1 yields the input mix signal MIN at the output, a dialogue gain between 0 and 1 results in dialogue attenuation, and a gain larger than 1 results in a dialogue boost. One advantage of this implementation is that processing of the dialogue estimate DE2 and non-dialogue (i.e. background) MB may be done separately, i.e., the processing of the dialogue estimate DE2 will not affect the background content and vice versa. [188] In the mixer 13, the dialogue estimate DE2Gain is mixed with the non-dialogue estimate MB,Gain resulting in a (dialogue enhanced) dialogue estimate signal DE2Final comprising a mix of the gained dialogue estimate DE2Gain and the gained non-dialogue MB,Gain. For example, to make dialogue more intelligible the background gain and dialogue gain applied to DE2 and MB respectively to form DE2Gain and MB,Gain are configured to boost the dialogue estimate DE2 or attenuate the non-dialogue estimate MB. [189] For example, this allows separate EQ filtering processes and/or dynamic range compression (DRC) to be performed for the dialogue estimate DE2 and non-dialogue (background) MB content as shown in the dialogue enhancement system 303 of FIG.26. In the dialogue enhancement system 303 separate audio processing modules 15a, 15b have be inserted to process each of dialogue estimate DE2 and the background signal MB upstream of the dialogue gain application module 11 and mix gain application module 12. Each audio processing module 15a, 15b performs at least one of DRC processing and EQ filter application.
[190] FIG.27 shows a dialogue enhancement system 304 in which the dialogue estimation system 10C is used with an audio codec that encodes the dialogue estimate DE2 parametrically relative to the mix signal. An example of such a parametric encoding of dialogue is the AC4 codec, which are described in ETSI TS 103190-1, section 5.7.8 and the encoding modes "Parametric channel independent enhancement", "Parametric cross-channel enhancement", and "Waveform-parametric hybrid", each of which are hereby incorporated by reference in its entirety. [191] To encode the dialogue estimate DE2 parametrically a parametric coding block 17 is used to parametrically encode the dialogue estimate DE2 with respect to the input mix signal MIN to yield a dialogue estimate parameter bitstream. The input mix signal MIN as such is encoded (e.g. waveform encoded) using encoding block 16 to yield an input mix signal bitstream. The dialogue estimate parameter bitstream and the input mix signal bitstream are provided to a bitstream multiplexer 18 which multiplexes the two bitstreams into a complete bitstream B. [192] In the dialogue enhancement system 305 of FIG.28, the dialogue estimate system 10C is used with an audio codec that encodes the dialogue waveform and the background waveform into two separate bitstreams using encoding modules 19a, 19b. The separate bitstreams are subsequently multiplexed to form a complete bitstream B by bitstream multiplexing module 18. For example, the audio codec could be AC4. Systems and methods to encode the dialogue and background waveforms and multiplex them to form a complete bitstream are described in U.S. Patent No.10,453,467 titled “Transmission-agnostic presentation-based program loudness” and ETSI TS 103190-1, section 4.3.3.3 "AC-4 presentation information", presentation configurations containing dialogue sub streams, each of which is incorporated by reference in its entirety. [193] The present disclosure likewise relates to an apparatus (e.g., computer-implemented apparatus or apparatus having processing capability in general) for performing or implementing methods and techniques described throughout the present disclosure. For example, this apparatus may relate to a dialogue estimation system (with or without dialogue detection) or a dialogue enhancement system incorporating the dialogue estimation system (with or without dialogue detection). [194] FIG.29 shows an example of such apparatus 1100. In particular, apparatus 1100 comprises a processor 1110 and a memory 1120 coupled to the processor 1110. The memory 1120 may store instructions for the processor 1110. The processor 1110 may also receive, among others, suitable input data 1130 (e.g., input frames of time-frequency coefficients, etc.), depending on use cases and/or implementations. The processor 1110 may be adapted to carry out
or implement the methods/techniques described throughout the present disclosure and to generate corresponding output data 1140, depending on use cases and/or implementations. [195] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products. [196] Aspects of the methods and apparatus/systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof. [197] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and/or as data and/or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and/or other characteristics. Computer-readable media in which such formatted data and/or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media. [198] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and/or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the apparatus (e.g., encoders) described above may include one or more electronic processors, one or more computer-readable medium modules, one or more input/output interfaces, and various connections (e.g., a system bus) connecting the various components.
[199] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements. [200] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. [201] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs): EEE1. A dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the system comprising: an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal; a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal; an adaptive downmixer configured to: downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; an output dialogue estimate processor configured to: determine a second dialogue estimate of the downmixed signal; and an adaptive upmixer configured to: upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
EEE2. A dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the system comprising: an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal; a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal; an adaptive downmixer configured to: downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; an adaptive upmixer configured to: upmix the downmixed signal based on the spatial parameters to generate an upmixed signal; and an output dialogue estimate processor configured to: determine a second dialogue estimate of the upmixed signal to generate an output dialogue signal. EEE3. The dialogue estimation system of EEE1 or EEE2, wherein the output dialogue signal comprises N channels, wherein N ≥ 1. EEE4. The dialogue estimation system of any one of EEE1-EEE3, wherein the output dialogue signal comprises M audio objects, wherein M ≥ 1. EEE5. The dialogue estimation system of any one of EEE1-EEE4, wherein the input dialogue estimate processor comprises a mask-based system configured to determine the first dialogue estimate by: determining a time and frequency varying mask based on the input mix signal; and applying the time and frequency varying mask to the input mix signal to determine the first dialogue estimate of the input mix signal. EEE6. The dialogue estimation system of EEE5, wherein the mask-based system comprises one or more neural networks.
EEE7. The dialogue estimation system of EEE5 or EEE6, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model. EEE8. The dialogue estimation system of any one of EEE1-EEE4, wherein the input dialogue estimate processor comprises a direct-mapped system configured to determine the first dialogue estimate by directly mapping the input mix signal to the first dialogue estimate. EEE9. The dialogue estimation system of EEE8, wherein the direct-mapped system is further configured to map the N-channel input mix signal to an N-channel first dialogue estimate signal and/or map the M-audio object input mix signal to an M-audio object input mix signal, wherein the first dialogue estimate comprises the N-channel first dialogue estimate signal and/or the M-audio object first dialogue estimate signal. EEE10. The dialogue estimation system of EEE8 or EEE9, wherein the direct-mapped system comprises one or more neural networks. EEE11. The dialogue estimation system of any one of EEE8-EEE10, wherein the direct- mapped system comprises at least one of a generative model or a UNIVERSE model. EEE12. The dialogue estimation system of any one of EEE1-EEE11, wherein the input dialogue estimate processor comprises: K single element dialogue estimate processors, wherein K ≥ 1, and wherein each single element dialogue estimate processor of the K single element dialogue estimate processors is configured to: receive, as input, a single channel input mix signal or a single object input mix signal; and determine a dialogue estimate of the single channel input mix signal or the single object input mix signal. EEE13. The dialogue estimation system of EEE12, wherein K = M + N . EEE14. The dialogue estimation system of any one of EEE1-EEE11, wherein the input dialogue estimate processor comprises: a single element mask generator; and wherein the input dialogue estimate processor is configured to:
downmix the input mix signal into a mono mix signal; determine a mask of the mono mix signal via the single element mask generator; and apply the mask to each channel of the input mix signal and/or each audio object of the input mix signal to determine a dialogue estimate of each channel and/or each audio object. EEE15. The dialogue estimation system of any one of EEE1-EEE11, wherein M + N ≥ 3, wherein the input dialogue estimate processor comprises: a first single element mask generator; a second single element mask generator; and wherein the input dialogue estimate processor is configured to: downmix a first channel or object and a second channel or object of the input mix signal into a first mono mix signal; downmix the second channel or object and a third channel or object of the input mix signal into a second mono mix signal; determine a first mask of the first mono mix signal via the first single element mask generator; determine a second mask of the second mono mix signal via the second single element mask generator; apply the first mask to the first channel or object to determine a dialogue estimate of the first channel or object of the input mix signal; apply the second mask to the third channel or object to determine a dialogue estimate of the third channel or object of the input mix signal. EEE16. The dialogue estimation system of EEE15, wherein the input dialogue estimate processor is further configured to: apply the first mask to the second channel or object to determine a first masked version of second channel or object; apply the second mask to the second channel or object to determine a second masked version of second channel or object; and combine the first masked version and the second masked version of the second channel or object to determine a dialogue estimate of the second channel or object of the input mix signal.
EEE17. The dialogue estimation system of any one of EEE1-EEE11, wherein M + N ≥ 3, wherein the input dialogue estimate processor comprises: a first single element mask generator; a second element mask generator; and wherein the input dialogue estimate processor is configured to: downmix a first channel or object and a second channel or object of the input mix signal into a mono mix signal or; determine a first mask of the mono mix signal via the first single element mask generator; determine a second mask of a third channel or object of the input mix signal or via the second single element mask generator; apply the first mask to the first channel or object to determine a dialogue estimate of the first channel or object of the input mix signal; apply the first mask to the second channel or object of the input mix signal to determine a dialogue estimate of the second channel or object of the input mix signal; and apply the second mask to the third channel or object of the input mix signal to determine a dialogue estimate of the third channel or object of the input mix signal. EEE18. The dialogue estimation system of any one of EEE1 and EEE3-EEE17, wherein the output dialogue estimate processor comprises a mask-based system configured to determine the second dialogue estimate by: determining a time and frequency varying mask based on the downmixed signal; and applying the mask to the downmixed signal to determine the second dialogue estimate of the downmixed signal. EEE19. The dialogue estimation system of EEE18, wherein the mask-based system comprises one or more neural networks. EEE20. The dialogue estimation system of EEE18 or EEE19, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model.
EEE21. The dialogue estimation system of any one of EEE1 and EEE3-EEE17, wherein the output dialogue estimate processor comprises a direct-mapped system configured to determine the second dialogue estimate by directly mapping the downmixed signal to the second dialogue estimate. EEE22. The dialogue estimation system of EEE21, wherein the direct-mapped system comprises one or more neural networks. EEE23. The dialogue estimation system of EEE21 or EEE22, wherein the direct-mapped system comprises at least one of a generative model or a UNIVERSE model. EEE24. The dialogue estimation system of any one of EEE1-EEE23, wherein the spatial parameters comprise at least one of adaptive downmixing parameters and adaptive upmixing parameters. EEE25. The dialogue estimation system of any one of EEE1-EEE24, wherein the spatial analyzer is configured to determine the spatial parameters based on the first dialogue estimate of the input mix signal via correlation based spatial analysis and/or amplitude ratio based spatial analysis. EEE26. The dialogue estimation system of any one of EEE1-EEE25, wherein the spatial parameters comprise one scalar value per channel of the input mix signal or one scalar value per audio object of the input mix signal. EEE27. The dialogue estimation system of any one of EEE1-EEE26, further comprising a dialogue detector configured to determine one or more dialogue confidence values. EEE28. The dialogue estimation system according to any one of the preceding EEEs, wherein the spatial parameters are panning parameters. EEE29. The dialogue estimation system according to any one of the preceding claims, wherein the input mix signal comprises a subset of the channels and/or audio objects of an original mix signal, the original mix signal comprising at least one residual channel and/or audio object in addition to the at least one of N ≥ 1 channels and/or M ≥ 1 audio objects of the mix input signal.
EEE30. The dialogue estimation system according to EEE29, wherein the dialogue estimation system is further configured to: combine the output dialogue signal with the at least one residual channel and/or audio object to form a hybrid dialogue estimate signal. EEE31. A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: apply a dialogue gain to the output dialogue signal of the dialogue estimation system to compute a dialogue estimation; and determine a dialogue enhanced mix signal based on a combination of the dialogue estimation and the input mix signal. EEE32. A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; apply a dialogue gain to the output dialogue signal to determine a first signal; apply a background gain to the background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals. EEE33. A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; process the background estimate and the output dialogue signal; apply a dialogue gain to the processed dialogue signal to determine a first signal; apply a background gain to the processed background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
EEE34. A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: apply parametric coding to the output dialogue signal of the dialogue estimation system to determine a dialogue parameter bitstream; and encode the input mix signal to determine a mix waveform bitstream. EEE35. A dialogue enhancement system comprising the dialogue estimation system of any one of EEE1-EEE30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; encode the output dialogue signal to determine a dialogue bitstream; and encode the background estimate to determine a background bitstream. EEE36. A dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the method comprising: determining a first dialogue estimate of the input mix signal; determining spatial parameters based on the first dialogue estimate of the input mix signal; downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; determining a second dialogue estimate of the downmixed signal; and upmixing the second dialogue estimate based on the spatial parameters to generate an output dialogue signal. EEE37. A dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the method comprising: determining a first dialogue estimate of the input mix signal; determining spatial parameters based on the first dialogue estimate of the input mix signal; downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal;
upmixing the downmixed signal based on the spatial parameters to generate an upmixed signal; and determining a second dialogue estimate of the upmixed signal to generate an output dialogue signal. EEE38. The dialogue estimation method according to EEE36 or EEE37, wherein the spatial parameters are panning parameters. EEE39 An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of EEE36- EEE38. EEE40. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE36-EEE38. EEE41. A computer-readable storage medium storing the program according to EEE40.
Claims
CLAIMS 1. A dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the system comprising: an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal; a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal; an adaptive downmixer configured to: downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; an output dialogue estimate processor configured to: determine a second dialogue estimate of the downmixed signal; and an adaptive upmixer configured to: upmix the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
2. A dialogue estimation system for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the system comprising: an input dialogue estimate processor configured to: determine a first dialogue estimate of the input mix signal; a spatial analyzer configured to: determine spatial parameters based on the first dialogue estimate of the input mix signal; an adaptive downmixer configured to: downmix the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; an adaptive upmixer configured to: upmix the downmixed signal based on the spatial parameters to generate an upmixed signal; and an output dialogue estimate processor configured to:
determine a second dialogue estimate of the upmixed signal to generate an output dialogue signal.
3. The dialogue estimation system of claim 1 or 2, wherein the output dialogue signal comprises N channels, wherein N ≥ 1.
4. The dialogue estimation system of any one of claims 1-3, wherein the output dialogue signal comprises M audio objects, wherein M ≥ 1.
5. The dialogue estimation system of any one of claims 1-4, wherein the input dialogue estimate processor comprises a mask-based system configured to determine the first dialogue estimate by: determining a time and frequency varying mask based on the input mix signal; and applying the time and frequency varying mask to the input mix signal to determine the first dialogue estimate of the input mix signal.
6. The dialogue estimation system of claim 5, wherein the mask-based system comprises one or more neural networks.
7. The dialogue estimation system of claim 5 or 6, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model.
8. The dialogue estimation system of any one of claims 1-4, wherein the input dialogue estimate processor comprises a direct-mapped system configured to determine the first dialogue estimate by directly mapping the input mix signal to the first dialogue estimate.
9. The dialogue estimation system of claim 8, wherein the direct-mapped system is further configured to map the N-channel input mix signal to an N-channel first dialogue estimate signal and/or map the M-audio object input mix signal to an M-audio object input mix signal, wherein the first dialogue estimate comprises the N-channel first dialogue estimate signal and/or the M-audio object first dialogue estimate signal.
10. The dialogue estimation system of claim 8 or 9, wherein the direct-mapped system comprises one or more neural networks.
11. The dialogue estimation system of any one of claims 8-10, wherein the direct-mapped system comprises at least one of a generative model or a UNIVERSE model.
12. The dialogue estimation system of any one of claims 1-11, wherein the input dialogue estimate processor comprises: K single element dialogue estimate processors, wherein K ≥ 1, and wherein each single element dialogue estimate processor of the K single element dialogue estimate processors is configured to: receive, as input, a single channel input mix signal or a single object input mix signal; and determine a dialogue estimate of the single channel input mix signal or the single object input mix signal.
13. The dialogue estimation system of claim 12, wherein K = M + N .
14. The dialogue estimation system of any one of claims 1-11, wherein the input dialogue estimate processor comprises: a single element mask generator; and wherein the input dialogue estimate processor is configured to: downmix the input mix signal into a mono mix signal; determine a mask of the mono mix signal via the single element mask generator; and apply the mask to each channel of the input mix signal and/or each audio object of the input mix signal to determine a dialogue estimate of each channel and/or each audio object.
15. The dialogue estimation system of any one of claims 1-11, wherein M + N ≥ 3, wherein the input dialogue estimate processor comprises: a first single element mask generator; a second single element mask generator; and wherein the input dialogue estimate processor is configured to: downmix a first channel or object and a second channel or object of the input mix signal into a first mono mix signal;
downmix the second channel or object and a third channel or object of the input mix signal into a second mono mix signal; determine a first mask of the first mono mix signal via the first single element mask generator; determine a second mask of the second mono mix signal via the second single element mask generator; apply the first mask to the first channel or object to determine a dialogue estimate of the first channel or object of the input mix signal; apply the second mask to the third channel or object to determine a dialogue estimate of the third channel or object of the input mix signal.
16. The dialogue estimation system of claim 15, wherein the input dialogue estimate processor is further configured to: apply the first mask to the second channel or object to determine a first masked version of second channel or object; apply the second mask to the second channel or object to determine a second masked version of second channel or object; and combine the first masked version and the second masked version of the second channel or object to determine a dialogue estimate of the second channel or object of the input mix signal.
17. The dialogue estimation system of any one of claims 1-11, wherein M + N ≥ 3, wherein the input dialogue estimate processor comprises: a first single element mask generator; a second element mask generator; and wherein the input dialogue estimate processor is configured to: downmix a first channel or object and a second channel or object of the input mix signal into a mono mix signal or; determine a first mask of the mono mix signal via the first single element mask generator; determine a second mask of a third channel or object of the input mix signal or via the second single element mask generator; apply the first mask to the first channel or object to determine a dialogue estimate of the first channel or object of the input mix signal;
apply the first mask to the second channel or object of the input mix signal to determine a dialogue estimate of the second channel or object of the input mix signal; and apply the second mask to the third channel or object of the input mix signal to determine a dialogue estimate of the third channel or object of the input mix signal.
18. The dialogue estimation system of any one of claims 1 and 3-17, wherein the output dialogue estimate processor comprises a mask-based system configured to determine the second dialogue estimate by: determining a time and frequency varying mask based on the downmixed signal; and applying the mask to the downmixed signal to determine the second dialogue estimate of the downmixed signal.
19. The dialogue estimation system of claim 18, wherein the mask-based system comprises one or more neural networks.
20. The dialogue estimation system of claim 18 or 19, wherein the mask-based system comprises at least one of a deep neural network or a LensNet model.
21. The dialogue estimation system of any one of claims 1 and 3-17, wherein the output dialogue estimate processor comprises a direct-mapped system configured to determine the second dialogue estimate by directly mapping the downmixed signal to the second dialogue estimate.
22. The dialogue estimation system of claim 21, wherein the direct-mapped system comprises one or more neural networks.
23. The dialogue estimation system of claim 21 or 22, wherein the direct-mapped system comprises at least one of a generative model or a UNIVERSE model.
24. The dialogue estimation system of any one of claims 1-23, wherein the spatial parameters comprise at least one of adaptive downmixing parameters and adaptive upmixing parameters.
25. The dialogue estimation system of any one of claims 1-24, wherein the spatial analyzer is configured to determine the spatial parameters based on the first dialogue estimate of the input mix signal via correlation based spatial analysis and/or amplitude ratio based spatial analysis.
26. The dialogue estimation system of any one of claims 1-25, wherein the spatial parameters comprise one scalar value per channel of the input mix signal or one scalar value per audio object of the input mix signal.
27. The dialogue estimation system of any one of claims 1-26, further comprising a dialogue detector configured to determine one or more dialogue confidence values.
28. The dialogue estimation system according to any one of the preceding claims, wherein the spatial parameters are panning parameters.
29. The dialogue estimation system according to any one of the preceding claims, wherein the input mix signal comprises a subset of the channels and/or audio objects of an original mix signal, the original mix signal comprising at least one residual channel and/or audio object in addition to the at least one of N ≥ 1 channels and/or M ≥ 1 audio objects of the mix input signal.
30. The dialogue estimation system according to claim 29, wherein the dialogue estimation system is further configured to: combine the output dialogue signal with the at least one residual channel and/or audio object to form a hybrid dialogue estimate signal.
31. A dialogue enhancement system comprising the dialogue estimation system of any one of claims 1-30, wherein the dialogue enhancement system is configured to: apply a dialogue gain to the output dialogue signal of the dialogue estimation system to compute a dialogue estimation; and determine a dialogue enhanced mix signal based on a combination of the dialogue estimation and the input mix signal.
32. A dialogue enhancement system comprising the dialogue estimation system of any one of claims 1-30, wherein the dialogue enhancement system is configured to:
determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; apply a dialogue gain to the output dialogue signal to determine a first signal; apply a background gain to the background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
33. A dialogue enhancement system comprising the dialogue estimation system of any one of claims 1-30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; process the background estimate and the output dialogue signal; apply a dialogue gain to the processed dialogue signal to determine a first signal; apply a background gain to the processed background estimate to determine a second signal; and determine a dialogue enhanced mix signal based on a combination of the first and second signals.
34. A dialogue enhancement system comprising the dialogue estimation system of any one of claims 1-30, wherein the dialogue enhancement system is configured to: apply parametric coding to the output dialogue signal of the dialogue estimation system to determine a dialogue parameter bitstream; and encode the input mix signal to determine a mix waveform bitstream.
35. A dialogue enhancement system comprising the dialogue estimation system of any one of claims 1-30, wherein the dialogue enhancement system is configured to: determine a background estimate based on the input mix signal and the output dialogue signal of the dialogue estimation system; encode the output dialogue signal to determine a dialogue bitstream; and encode the background estimate to determine a background bitstream.
36. A dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the method comprising: determining a first dialogue estimate of the input mix signal; determining spatial parameters based on the first dialogue estimate of the input mix signal; downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; determining a second dialogue estimate of the downmixed signal; and upmixing the second dialogue estimate based on the spatial parameters to generate an output dialogue signal.
37. A dialogue estimation method for extracting dialogue from an input mix signal comprising at least one of N ≥ 1 channels and/or M ≥ 1 audio objects, wherein M + N ≥ 2, the method comprising: determining a first dialogue estimate of the input mix signal; determining spatial parameters based on the first dialogue estimate of the input mix signal; downmixing the input mix signal based on the input mix signal and the spatial parameters to generate a downmixed signal; upmixing the downmixed signal based on the spatial parameters to generate an upmixed signal; and determining a second dialogue estimate of the upmixed signal to generate an output dialogue signal.
38. The dialogue estimation method according to claim 36 or claim 37, wherein the spatial parameters are panning parameters.
39. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of claims 36-38.
40. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 36-38.
41. A computer-readable storage medium storing the program according to claim 40.
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463563651P | 2024-03-11 | 2024-03-11 | |
| US63/563,651 | 2024-03-11 | ||
| US202463678508P | 2024-08-01 | 2024-08-01 | |
| US63/678,508 | 2024-08-01 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025190810A1 true WO2025190810A1 (en) | 2025-09-18 |
Family
ID=94970248
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2025/056307 Pending WO2025190810A1 (en) | 2024-03-11 | 2025-03-07 | Systems and methods for spatial fidelity improving dialogue estimation |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025190810A1 (en) |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10453467B2 (en) | 2014-10-10 | 2019-10-22 | Dolby Laboratories Licensing Corporation | Transmission-agnostic presentation-based program loudness |
| WO2021252823A1 (en) | 2020-06-11 | 2021-12-16 | Dolby Laboratories Licensing Corporation | Methods, apparatus, and systems for detection and extraction of spatially-identifiable subband audio sources |
| WO2023052523A1 (en) | 2021-09-29 | 2023-04-06 | Dolby International Ab | Universal speech enhancement using generative neural networks |
| WO2023192036A1 (en) | 2022-03-29 | 2023-10-05 | Dolby Laboratories Licensing Corporation | Multichannel and multi-stream source separation via multi-pair processing |
| US20230368807A1 (en) | 2020-10-29 | 2023-11-16 | Dolby Laboratories Licensing Corporation | Deep-learning based speech enhancement |
| US20240007817A1 (en) * | 2022-06-30 | 2024-01-04 | Amazon Technologies, Inc. | Real-time low-complexity stereo speech enhancement with spatial cue preservation |
-
2025
- 2025-03-07 WO PCT/EP2025/056307 patent/WO2025190810A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10453467B2 (en) | 2014-10-10 | 2019-10-22 | Dolby Laboratories Licensing Corporation | Transmission-agnostic presentation-based program loudness |
| WO2021252823A1 (en) | 2020-06-11 | 2021-12-16 | Dolby Laboratories Licensing Corporation | Methods, apparatus, and systems for detection and extraction of spatially-identifiable subband audio sources |
| US20230368807A1 (en) | 2020-10-29 | 2023-11-16 | Dolby Laboratories Licensing Corporation | Deep-learning based speech enhancement |
| WO2023052523A1 (en) | 2021-09-29 | 2023-04-06 | Dolby International Ab | Universal speech enhancement using generative neural networks |
| WO2023192036A1 (en) | 2022-03-29 | 2023-10-05 | Dolby Laboratories Licensing Corporation | Multichannel and multi-stream source separation via multi-pair processing |
| US20240007817A1 (en) * | 2022-06-30 | 2024-01-04 | Amazon Technologies, Inc. | Real-time low-complexity stereo speech enhancement with spatial cue preservation |
Non-Patent Citations (1)
| Title |
|---|
| AARON MASTER ET AL: "DeepSpace: Dynamic Spatial and Source Cue Based Source Separation for Dialog Enhancement", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 16 February 2023 (2023-02-16), XP091439550 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7161564B2 (en) | Apparatus and method for estimating inter-channel time difference | |
| US12198705B2 (en) | Apparatus, method or computer program for estimating an inter-channel time difference | |
| CN102160113B (en) | Multichannel audio coder and decoder | |
| CN101809655B (en) | Apparatus and method for encoding a multi channel audio signal | |
| CN101138274B (en) | Device and method for processing decoherent or combined signals | |
| KR101984115B1 (en) | Apparatus and method for multichannel direct-ambient decomposition for audio signal processing | |
| US8015018B2 (en) | Multichannel decorrelation in spatial audio coding | |
| KR20170042709A (en) | A signal processing apparatus for enhancing a voice component within a multi-channal audio signal | |
| KR101710544B1 (en) | Method and apparatus for decomposing a stereo recording using frequency-domain processing employing a spectral weights generator | |
| KR20170063657A (en) | Audio encoder and decoder | |
| CN105284133A (en) | Apparatus and method for center signal scaling and stereophonic enhancement based on a signal-to-downmix ratio | |
| US20250166654A1 (en) | Apparatus, method or computer program for generating an output downmix representation | |
| JP5053849B2 (en) | Multi-channel acoustic signal processing apparatus and multi-channel acoustic signal processing method | |
| EP4576071A1 (en) | Generation of multichannel audio signal | |
| JP2007025290A (en) | Device for controlling reverberation in a multi-channel acoustic codec | |
| CN116529813A (en) | Apparatus, method or computer program for processing encoded audio scenes using parameter conversion |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25711445 Country of ref document: EP Kind code of ref document: A1 |