FIELD OF THE INVENTION
-
The invention relates to generation of multichannel audio signals and in particular, but not exclusively, to generation of stereo signals from upmixing of a mono downmix audio signal.
BACKGROUND OF THE INVENTION
-
Spatial audio applications have become numerous and widespread and increasingly form at least part of many audiovisual experiences. Indeed, new and improved spatial experiences and applications are continuously being developed which results in increased demands for audio processing and rendering.
-
For example, in recent years, Virtual Reality (VR) and Augmented Reality (AR) have received increasing interest, and a number of implementations and applications are reaching the consumer market. Indeed, equipment is being developed for both rendering the experience as well as for capturing or recording suitable data for such applications. For example, relatively low-cost equipment is being developed for allowing gaming consoles to provide a full VR experience. It is expected that this trend will continue and indeed will increase in speed with the market for VR and AR reaching a substantial size within a short time scale. In the audio domain, a prominent field explores the reproduction and synthesis of realistic and natural spatial audio. The ideal aim is to produce natural audio sources such that the user cannot recognize the difference between a synthetic or an original one.
-
A lot of research and development effort has focused on providing efficient and high-quality audio encoding and audio decoding for spatial audio. A frequently used spatial audio representation is multichannel audio representations, including stereo representation, and efficient encoding of such multichannel audio based on downmixing multichannel audio signals to downmix channels with fewer channels have been developed. One of the main advances in low bit-rate audio coding has been the use of parametric multichannel coding where a downmix audio signal is generated together with parametric data that can be used to upmix the downmix audio signal to recreate the multichannel audio signal.
-
In particular, instead of traditional mid-side or intensity coding, parametric multichannel audio coding uses a downmix of a multichannel input signal to a lower number of channels (e.g. two to one) and multichannel image (stereo) parameters are extracted. Then the downmix audio signal is encoded using a more traditional audio coder (e.g. a mono audio encoder). The data of the downmix is combined with the encoded multichannel parameter data to generate a suitable audio bitstream. This bitstream is then transmitted to the decoder, where the process is inverted. First the downmix audio signal is decoded, after which the multichannel audio signal is reconstructed, guided by the encoded multichannel image upmix parameters.
-
An example of stereo coding is described in
E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, "Advances in Parametric Coding for High-Quality Audio", 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852. In the described approach, the downmixed mono signal is parametrized by exploiting the natural separation of the signal into three components (objects): transients, sinusoids, and noise. In
E. Schuijers, J. Breebaart, H. Pumhagen, J. Engdegård, "Low Complexity Parametric Stereo Coding", 116th AES, Berlin, Germany, 2004, Preprint 6073 more details are provided describing how parametric stereo was realized with a low (decoder) complexity when combining it with Spectral Band Replication (SBR).
-
In the described approaches, the decoding side multi-channel regeneration is based on the use of a so-called de-correlation process. The de-correlation process generates a decorrelated helper signal from the downmix audio signal, i.e. from the monaural signal of parametric stereo. The decorrelated signal is in particular used to control and reconstruct the coherence (stereo width in case of stereo) between/among the different channels during the upmix process. In the stereo reconstruction process, both the monaural signal and the decorrelated helper signal are used to generate the upmixed stereo signal based on the upmix parameters. Specifically, the two signals may be multiplied by a time- and frequency-dependent 2x2 matrix having coefficients determined from the upmix parameters to provide the output stereo signal.
-
However, although Parametric Stereo (PS) and similar downmix encoding/ decoding approaches were a leap forward from traditional stereo and multichannel coding, the approach is not optimal in all scenarios. In particular, known approaches tend to introduce some distortion, changes, artefacts etc. that may introduce differences between the (original) multichannel audio signal input to the encoder and the multichannel audio signal recreated at the decoder. Typically, the audio quality may be degraded and imperfect recreation of the multichannel occurs. Further, the data rate may still be higher than desired and/or the complexity/ resource usage of the involved processing may be higher than preferred. In particular, it is challenging to generate a decorrelated helper signal that results in optimum audio quality of the upmixed multichannel signal. For example, it is known that the decorrelation process often times smears out audio variations and may introduce additional artifacts such as post-echoes due to the filter characteristics of the decorrelator.
-
Hence, an improved approach would be advantageous. In particular, an approach allowing increased flexibility, improved adaptability, an improved performance, increased audio quality, improved audio quality to data rate trade-off, reduced complexity and/or resource usage, reduced computational load, facilitated implementation, improved dynamic audio processing and reproduction, and/or an improved spatial audio experience would be advantageous.
SUMMARY OF THE INVENTION
-
Accordingly, the invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.
-
According to an aspect of the invention there is provided an apparatus for generating a multichannel audio signal, the apparatus comprising: a receiver arranged to receive a downmix audio signal being a downmix of a plurality of audio signals and upmix parameters indicative of relative properties of the plurality of audio signals; a transient circuit arranged to generate a transient signal indicative of a likelihood of samples of the downmix audio signal representing a transient; an auxiliary audio signal circuit arranged to generate an auxiliary audio signal from a synthesized signal and in dependence on the transient signal, the circuit comprising a trained artificial neural network generating the synthesized signal, the trained artificial neural network receiving input data determined from the downmix audio signal and from the transient signal; and an output generator arranged to generate the multichannel audio signal by upmixing the downmix audio signal and the auxiliary audio signal in dependence on the upmix parameters.
-
The approach may provide an improved audio experience in many embodiments. For many signals and scenarios, the approach may provide improved generation/ reconstruction of a multichannel audio signal with an improved perceived audio quality. The approach may provide a particularly advantageous arrangement which may in many embodiments and scenarios allow a facilitated and/or improved possibility of utilizing artificial neural networks in audio processing, including typically audio encoding and/or decoding. The approach may allow an advantageous employment of artificial neural network(s) in generating a multichannel audio signal from a downmix audio signal.
-
The approach may provide an efficient implementation and may in many embodiments allow a reduced complexity and/or resource usage.
-
The approach may allow a particularly efficient and high performance approach for improving and increasing the correspondence of properties between a received downmix audio signal and an auxiliary (decorrelated) signal used to generate the multichannel audio signal. In particular, the approach may effectively mitigate and reduce the impact of generated decorrelated signals not having ideal temporal properties, including in particular errors and/or deviations in the signal level or envelope of the decorrelated signal.
-
The approach may in many cases in particular mitigate or avoid temporal smearing and/or frequency distortion thereby allowing improved quality of the resulting multichannel audio signal.
-
In many embodiments, the input data to the trained artificial neural network may include data determined from the transient signal, and in many cases specifically samples of the transient signal.
-
The upmix parametric data may comprise parameter (values) relating properties of the downmix audio signal to properties of the multichannel audio signal. The upmix parametric data may comprise data being indicative of relative properties between channels of the multichannel audio signal. The upmix parametric data may comprise data being indicative of differences in properties between channels of the multichannel audio signal. The upmix parametric data may comprise data being perceptually relevant for the synthesis of the multichannel audio signal. The properties may for example be differences in phase and/or intensity and/or timing and/or correlation. The upmix parametric data may comprise data including at least one of interchannel intensity differences (IID), interchannel timing differences (ITD), interchannel correlations (IIC) and/or interchannel phase differences (IPD) for channels of the multichannel audio signal. In particular, the upmix parameters may include one or more of an ICC, IPD, IID parameter as known from Parametric Stereo encoding/decoding.
-
The artificial neural network(s) may be a trained artificial neural network(s) trained by training data including training downmix audio signals and training upmix parametric data generated from training multichannel audio signals; the training employing a cost function dependent on a (spectro-temporal) envelope/shape/ difference between the downmix audio signal, and the synthesized signal generated by the artificial neural network. The cost function may be dependent on a correlation between the synthesized signal and the downmix audio signal, and specifically may provide an decreasing cost function for a decreasing correlation. The cost function may be dependent on a difference between statistical properties of the synthesized signal and the downmix audio signal, and specifically may provide an decreasing cost function for a decreasing correlation. Depending on the preferences and the decorrelation calculation/measurements/metric, it may in some embodiments be appropriate for the cost function to be dependent on a correlation between the synthesized signal and the downmix audio signal, and specifically to provide an increasing cost function for a decreasing correlation. The cost function may in some cases be dependent on a difference between statistical properties of the synthesized signal and the downmix audio signal, and specifically may provide an increasing cost function for a decreasing correlation. Such approaches may be used to ensure some decorrelation between the synthesized signal and the downmix audio signal.
-
The generator may be arranged to generate the multichannel audio signal by applying a matrix multiplication to the compensated (often frequency domain) decorrelated signal and the (frequency domain) downmix audio signal with the coefficients of the matrix being determined as a function of parameters of the upmix parametric data. The matrix may be time- and frequency-dependent.
-
The audio apparatus may specifically be an audio decoder apparatus. The auxiliary audio signal circuit generates an auxiliary audio signal from a synthesized signal which is generated by a trained artificial neural network.
-
The receiver may receive the downmix audio signal from any suitable internal or external source. For example, the audio apparatus may include a bitstream receiver for receiving a bitstream comprising a downmix audio signal and specifically a mono signal. The bitstream receiver may receive further data in the bitstream, such as the upmix data, positional information, acoustic environment data etc.
-
The trained artificial neural network may comprise input nodes receiving samples of the downmix audio signal and/or data values/features derived therefrom. The trained artificial neural network may comprise output nodes generating samples of the auxiliary audio signal and/or may generate data from which the auxiliary audio signal can be generated (e.g. using an analytical/predetermined function and/or not employing an artificial neural network).
-
The auxiliary audio signal may be generated directly as the synthesized signal. The auxiliary audio signal may be the synthesized signal.
-
The trained artificial neural network may be trained to generate the first signal to have statistical and/or spectro-temporal (envelope) properties approaching/similar to the spectro-temporal (envelope) properties of the downmix audio signal.
-
The input data may be a latent representation of the downmix audio signal such as that produced by an autoencoder, or a coarse spectro-temporal representation such as specifically a Mel spectrogram.
-
The transient signal may in some embodiments be a binary signal. The transient signal may in some embodiments indicate which samples of the downmix audio signal are part of, or not part of, a transient component/part. The transient signal may for each sample (instant) of the downmix audio signal indicate a probability of the sample (of that instant) being part of a transient part of the downmix audio signal.
-
According to an optional feature of the invention, the auxiliary signal circuit is arranged to generate the auxiliary audio signal to include a synthesized signal component being a component of the synthesized signal and a downmix audio signal component being a component of the downmix audio signal, the weight of the synthesized signal component relative to the downmix audio signal component being dependent on the transient signal.
-
This may provide a particularly efficient and high-performance operation in many scenarios and may typically result in substantially improved audio quality of the generated multichannel audio signal. It may in many scenarios allow a particularly advantageous implementation of audio upmixing utilizing an artificial neural network.
-
According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to: divide the downmix audio signal into a first signal component and a second signal component in dependence on the transient signal, the first signal component comprising faster changing parts of the downmix audio signal and the second signal component comprising slower changing parts of the downmix audio signal; generating the input data from the second signal component; and generating the auxiliary audio signal by combining the first signal component and the first signal in dependence on the transient signal.
-
This may provide a particularly efficient and high performance operation in many scenarios and may typically result in substantially improved audio quality of the generated multichannel audio signal. It may in many scenarios allow a particularly advantageous implementation of audio upmixing utilizing an artificial neural network.
-
According to an optional feature of the invention, the input data comprises a spectrogram of the downmix audio signal.
-
This may provide a particularly efficient and high performance operation and/or may facilitate implementation in many scenarios. The spectrogram may be a no-phase spectrogram (e.g. the spectrogram may only include amplitude values). The spectrogram may be a Mel-spectrogram. According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to generate the transient signal from a comparison of a signal level of the downmix audio signal in a first time window to a signal level of the downmix audio signal in a second time window, the second time window exceeding the first time window.
-
This may provide a particularly efficient and high performance operation in many scenarios, and may typically result in substantially improved audio quality of the generated multichannel audio signal. It may in many scenarios allow a particularly advantageous implementation of audio upmixing utilizing an artificial neural network.
-
The two signal levels may result from an averaging, or more generally from low pass filtering, of the downmix audio signal (or a signal derived therefrom). The averaging may have different durations and/or the low pass filtering may have different low pass characteristics. The second time window may have a duration exceeding the first time window by a factor of at least 2, 5, 10 times.
-
According to an optional feature of the invention, the trained artificial neural network comprises a generative model.
-
A generative model is particularly advantageous in many scenarios. It has been found that a diffusion model allows particularly efficient implementation and in particular tends to provide for a highly efficient and high performance approach for adapting the generation of the auxiliary audio signal based on the transient signal.
-
According to an optional feature of the invention, the trained artificial neural network comprises a diffusion model.
-
The trained artificial neural network may implement a diffusion model. A diffusion model is particularly advantageous in many scenarios. It has been found that a diffusion model allows particularly efficient implementation and in particular tends to provide for a highly efficient and high performance approach for adapting the generation of the auxiliary audio signal based on the transient signal.
-
According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to generate the input data to include a combination of samples of the downmix audio signal and noise samples, the combination being dependent on the transient signal.
-
This may provide a particularly advantageous operation in many embodiments and scenarios. It may allow a particularly efficient implementation of conditioning an artificial neural network on a transient signal.
-
In many embodiments, the weights for the samples of the downmix audio signal relative to the noise samples may be dependent on the transient signal. The noise samples may be randomly generated in accordance with a probability distribution, such as for example a Gaussian distribution.
-
In some embodiments, the diffusion model for a given iteration/(time) step is arranged to determine a mean value for a signal estimate for the first signal in dependence on the transient signal and a noise estimate for the given iteration.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
The diffusion model for a given iteration/(time) step may be arranged to determine a noise estimate in dependence on the transient signal.
-
The diffusion model for a given iteration/(time) step may be arranged to determine the signal estimate for the first signal as a combination of a signal estimate from the previous iteration of the given iteration and a noise compensated signal estimate for the given iteration, the noise compensated signal estimate being determined from the noise estimate for the given iteration, the combination being dependent on the transient signal.
-
According to an optional feature of the invention, the diffusion model is arranged to generate a mean value for a subsequent iteration as mean value for a signal estimate for the first signal for the current iteration offset by noise component determined in dependence on the transient signal.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to initialize the diffusion model with a noise level dependent on the transient signal.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to vary a noise level between sample times in dependence an energy level of the downmix audio signal for the sample times and the transient signal for the sample times.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
According to an optional feature of the invention, the auxiliary audio signal circuit is arranged to add a random change to the transient signal for some sample instants.
-
In some embodiments, the sample instants may be sample instants for which the transient signal meets a criterion indicative of the transient signal indicating a presence of a transient. In some embodiments, the sample instants may be sample instants for which the transient signal meets a criterion indicative of the transient signal not indicating a presence of a transient.
-
This may provide a particularly advantageous and high performance operation and/or implementation. The change may be such that the transient signal reduces the indication of a transient being present.
-
According to an optional feature of the invention, the diffusion model for at least a first iteration/ (time) step is arranged to generate a plurality of possible signal estimates for a following step by applying different noise components, and to select one signal estimate for the following step by selecting between the plurality of possible signal estimates in dependence on a correlation between each of the plurality of possible signal estimates and the downmix audio signal.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
According to an optional feature of the invention, the transient circuit is arranged to determine a time interval for which a transient is detected and to generate the transient signal to indicate a presence of the transient in a larger time interval than the time interval for which the transient is detected.
-
This may provide a particularly advantageous and high performance operation and/or implementation.
-
According to an aspect of the invention, there is provided a method of generating a multichannel audio signal, the method comprising: receiving a downmix audio signal being a downmix of a plurality of audio signals and upmix parameters indicative of relative properties of the plurality of audio signals; generating a transient signal indicative of a likelihood of samples of the downmix audio signal representing a transient; generating an auxiliary audio signal from a synthesized signal in dependence on the transient signal, the generating comprising a trained artificial neural network generating the synthesized signal, and the trained artificial neural network receiving input data determined from the downmix audio signal and from the transient signal; and generating the multichannel audio signal by upmixing the downmix audio signal and the auxiliary audio signal in dependence on the upmix parameters.
-
In many embodiments, the input data may be determined from the downmix audio signal and from the transient signal. In many embodiments, the trained artificial neural network may be conditioned on the transient signal, or on date determined from this.
-
These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.
BRIEF DESCRIPTION OF THE DRAWINGS
-
Embodiments of the invention will be described, by way of example only, with reference to the drawings, in which
- FIG. 1 illustrates some elements of an example of an audio apparatus in accordance with some embodiments of the invention;
- FIG. 2 illustrates some elements of an example of a time to frequency converter for an audio apparatus in accordance with some embodiments of the invention;
- FIG. 3 illustrates an example of a structure of an auxiliary audio signal circuit for an audio apparatus in accordance with some embodiments of the invention artificial neural network;
- FIG. 4 illustrates an example of a structure of an auxiliary audio signal circuit for an audio apparatus in accordance with some embodiments of the invention;
- FIG. 5 illustrates an example of a neuron for an artificial neural network;
- FIG. 6 illustrates an example of a structure of an artificial neural network;
- FIG. 7 illustrates an example of a structure of a diffusion model approach;
- FIG. 8 illustrates an example of elements of part of an audio apparatus in accordance with some embodiments of the invention;
- FIG. 9 illustrates an example of elements of an arrangement for training a diffusion-based generative artificial neural network;
- FIG. 10 illustrates some examples of an applause audio signal; and
- FIG. 11 illustrates some elements of a possible arrangement of a processor for implementing elements of an audio apparatus in accordance with some embodiments of the invention.
DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION
-
FIG. 1 illustrates some elements of an audio apparatus arranged to generate a multichannel audio signal.
-
The audio apparatus comprises a receiver 101 which is arranged to receive a downmix audio signal which is a downmix of a plurality of audio signals, such as specifically a downmix of a plurality of spatial audio channels each of which represents sound/audio at different locations in an audio environment. The receiver 101 further receives upmix parameters indicative of the relative properties between the plurality of audio signals, such as upmix parameters indicative of a phase difference, level difference, or coherence/correlation between audio signals of the plurality of audio signals.
-
The receiver 101 may specifically receive a data signal/ bitstream comprising an audio signal which is a downmix audio signal being a downmix of a multichannel audio signal which is to be recreated/generated by the audio apparatus. The following description will focus on a case where the multichannel audio signal is a stereo signal and the downmix audio signal is a mono signal, but it will be appreciated that the described approach and principles are equally applicable to the multichannel audio signal having more than two channels and to the downmix audio signal having more than a single channel (albeit fewer channels than the multichannel audio signal).
-
The received downmix audio signal is typically processed as time domain audio signal as will be described in the following. However, in some embodiments some or all of the processing may be performed using a frequency domain representation of the downmix audio signal and indeed in many embodiments the processing may include both processing of a time domain and frequency domain representations of the downmix audio signal.
-
Accordingly, in some embodiments, a frequency representation may be generated for the downmix audio signal and used in the processing. In some embodiments, the frequency domain audio signal may be received from a remote source which may generate a bitstream comprising a frequency representation of the downmix audio signal. In other embodiments, the remote source may generate a time domain representation of the audio signal and a local frequency transformer may be arranged to generate a (time-)frequency representation of the time domain representation. Thus, in some embodiments, the receiver 101 may receive the frequency domain audio signal from an internal source, such as an internal time-to-frequency domain transformer.
-
In particular, in the example of FIG. 1, the audio apparatus comprises a filter bank 103 which is arranged to generate a frequency (subband) representation of a received time domain downmix audio signal. Typically, the audio apparatus may comprise a filter bank 103 that is applied to the audio signal such that it is divided into frequency subbands.
-
The filter bank may be a Quadrature Mirror Filter (QMF) bank or may e.g. be implemented by a Fast Fourier Transform (FFT), but it will be appreciated that many other filter banks and approaches for dividing an audio signal into a plurality of subband signals are known and may be used. The filter-bank may specifically be a complex-exponential modulated pseudo QMF bank, resulting in e.g. 32 or 64 complex-valued sub-band signals.
-
The processing is furthermore typically performed in time segments or time slots. In most embodiments, the audio signal is divided into time intervals/segments with a conversion to the frequency/subband domain by applying e.g. an FFT or QMF filtering to the samples of each signal. For example, each channel of the downmix audio signal may be divided into time segments of e.g. 2048, 1024, or 512 samples. These signals may then be processed to generate samples for e.g. 64, 32 or 16 subbands. Thus, a set of samples may be determined for each subband of the downmix audio signal.
-
It should be noted that the number of time domain samples is not directly coupled to the number of subbands. Typically, for a so-called critically sampled filterbank of N bands, every N input samples will lead to N sub-band samples (one for every sub-band). An oversampled filterbank will produce more output samples. E.g. for every N input samples, it would generate k*N output samples, i.e., k consecutive samples for every band.
-
In some embodiments, the subbands are generated to have the same bandwidth but in other embodiments subbands are generated to have different bandwidths, e.g. reflecting the sensitivity of human hearing to different frequencies.
-
FIG. 2 illustrates an example of an approach where different bandwidths are generated by a hybrid analysis/synthesis filter bank approach.
-
In the example, this is realized by a combination of a complex-exponential modulated pseudo QMF bank 201 and a small filter bank 203 for the lower frequency bands to realize a higher frequency resolution, as desired for binaural perception of the human auditory system. The result is a hybrid filterbank with logarithmic filter band center-frequency spacings that follow that of human perception similar to equivalent rectangular bandwidths (ERBs). In order to compensate for the delay of the filtering by the small filter bank 203, a delay 205 is introduced for higher frequency subbands.
-
In the specific example, a time-domain signal x[n] is fed through a downsampled complex-exponential modulated QMF bank with K bands. Each frame of 64 time domain samples x[n] results in one slot of QMF samples X[k, m] with k = (0, ..., K - 1) at slot m. The lower slots are then filtered by additional complex-modulated filterbanks splitting the lower bands further. The higher slots are delayed ensuring that the filtered signals of the lower bands are in sync with the higher bands as the filtering introduces a delay. This finally results in a structure where for every 64 time-domain samples x[n], one slot m of hybrid QMF samples Y[l, m] is produced with l = (0, ..., L - 1) at slot m, e.g. with a total number of hybrid bands M = 77.
-
The received data signal further includes upmix parametric data for upmixing the downmix audio signal. The upmix parametric data may specifically be a set of upmix parameters that indicate relative properties/relationships between the signals of different audio channels of the multichannel audio signal (specifically the stereo signal) and/or between the downmix audio signal and audio channels of the multichannel audio signal. Typically, the upmix parameters may be indicative of time differences, phase differences, level/intensity differences and/or a measure of similarity, such as correlation. Typically, the upmix parameters are provided on a per time and per frequency basis (time frequency tiles). For example, new parameters may periodically be provided for a set of subbands. Parameters may specifically include Inter-channel phase difference (IPD), Overall phase difference (OPD), Inter-channel correlation (ICC), Channel phase difference (CPD) parameters as known from Parametric Stereo encoding (as well as from higher channel encodings).
-
Typically, the downmix audio signal is encoded and the receiver 101 may include a decoder function that decodes the downmix audio signal, i.e. the mono signal in the specific example.
-
The receiver 101 is coupled to a generator 105 which generates the multichannel audio signal from the frequency domain audio signal. In the example, the generator 105 is arranged to generate the multichannel audio signal by (up)mixing of (at least) the frequency domain audio signal and an auxiliary audio signal generated from the downmix audio signal. The upmixing is performed in dependence on the parametric upmix data.
-
The generator 105 may specifically for the stereo case generate the output multichannel audio signal by applying a 2x2 matrix multiplication to the samples of the downmix audio signal and the auxiliary audio signal. The coefficients of the 2x2 matrix are determined from the upmix parameters of the upmix parametric data, typically on a time and frequency band basis. For other upmix operations, such as from a mono or stereo downmix audio signal to a five channel multichannel audio signal, the generator 105 may apply matrix multiplications with matrices of suitable dimensions.
-
It will be appreciated that many different approaches of generating such a multichannel audio signal from a downmix audio signal and an auxiliary audio signal, and for determining suitable matrix coefficients from upmix parametric data, will be known to skilled person and that any suitable approach may be used. Specifically, various approaches for Parametric Stereo upmixing that are based on downmix and auxiliary audio signals are well known to the skilled person.
-
It has been found that by generating an auxiliary signal, and in particular a decorrelated signal, and mixing this with a downmix audio signal, and specifically a mono audio signal for stereo upmixing, an improved quality of the upmix signal is perceived and many decoders have been developed to exploit this.
-
The auxiliary audio signal is generated from the downmix audio signal and is specifically generated to be a different signal but with substantially the same spectro-temporal (envelope/shape) properties (for at least the majority of time). The auxiliary signal may be generated as a signal which is decorrelated with the downmix audio signal but with substantially the same spectro-temporal properties.
-
The audio apparatus specifically comprises an auxiliary audio signal circuit 107 which generates an auxiliary signal which, in a spectro-temporal sense, is similar to, but uncorrelated with, the downmix audio signal .
-
In the approach the generation of the auxiliary audio signal is based on a trained artificial neural network which receives input data that is generated from (and dependent on/reflecting properties of) the downmix audio signal and which based on the training of the artificial neural network proceeds to generate output samples from which the auxiliary audio signal is generated, for example by directly generating the auxiliary audio signal as the output of the trained artificial neural network. In some embodiments, the auxiliary audio signal circuit 107 may generate a set of features or parameters reflecting properties of the downmix audio signal, such as a spectrogram, and feed it to the trained artificial neural network which based on this, and the training, may proceed to directly generate samples of the auxiliary audio signal which can then be used for upmixing.
-
The input data to the trained artificial neural network in some cases directly include samples of the downmix audio signal. In some embodiments, the input data may alternatively or additionally be a latent space feature set which may be provided as input data to the trained artificial neural network of the auxiliary audio signal circuit 107. The input data, and specifically the latent space feature set, may for example be provided as feature inputs or the trained artificial neural network may be conditioned on the latent space feature set.
-
The audio apparatus further comprises a transient circuit 109 which is arranged to generate a transient signal indicative of a likelihood that samples of the downmix audio signal represent a transient. The transient signal may indicate which parts/samples of the downmix audio signal belong to/represent transient portions/parts of the downmix audio signal and which parts/samples belong to/represent non-transient portions/parts of the downmix audio signal.
-
In acoustics and audio, a transient is well known to be a sudden/quick increase in the signal level/amplitude for a short duration. For example, in many scenarios, a transient may be a quick high amplitude, short-duration sound at often at the beginning of a waveform. Transients often occur in audio such as musical sounds, noises, or speech. Transients are typically short bursts of increased energy in an audio signal, often occurring at the beginning of a sound, like a drum hit or a plucked guitar string. In the time-domain, transients display sharp, short duration peaks relative to surrounding samples. Viewed in the time-frequency domain, they can be seen as broadband (over frequency) bursts of energy of short duration.
-
Transients may in many cases be quite short in duration because of short attack times and can often range from 3-20 ms. For example, a transient may be considered to occur when the energy level in a time interval not exceeding, say, 1, 3, 5,10, 20, 50, 100 msec exceeds the energy level in a surrounding/neighbor time interval having a duration exceeding this by a factor of 2, 5, 10, or 20 times. It may e.g. be required that the energy level of the shorter time interval should exceed the energy level of the longer time interval by a factor of 2, 5, 10, 20, 50, 100 or more.
-
A transient may in many cases be considered to be a time interval/signal component for which the energy level is substantially higher (e.g. by a factor of 2, 5, 10) than in a surrounding/neighboring time interval. It may e.g. be required that the time interval/signal component/transient has a duration that does not exceed say, 1, 3, 5,10, 20, 50, 100 msec.
-
A transient may be a sufficiently high temporary increase in the energy level for a sufficiently short duration. A transient may be characterized by a high amplitude, sudden change in volume, and a brief duration.
-
The generated transient signal may in some cases be a binary signal that for samples of the downmix audio signal indicates whether they represent a transient or a non-transient part of the downmix audio signal/captured audio. In other embodiments, the generated transient signal may be a non-binary signal indicating confidence values for whether the samples of the downmix audio signal belong to transient or non-transient portions.
-
In the audio apparatus the transient signal is further fed to the auxiliary audio signal circuit 107 where it is further used to generate the auxiliary audio signal. Thus, in the approach both the downmix audio signal (and specifically various properties thereof) and the transient signal are used to generate the auxiliary audio signal.
-
In many embodiments, the transient signal can be considered a mask or weighting signal which is used by the auxiliary audio signal circuit 107 to determine relative weights for samples of the downmix audio signal and samples of a synthesized signal generated by the trained artificial neural network when generating/for contributing to the auxiliary audio signal. In the binary case, the transient signal may indicate whether the auxiliary audio signal should be generated from samples of the downmix audio signal or from samples of the synthesized signal. For example, for samples that the transient signal indicate belongs to a transient, the auxiliary audio signal circuit 107 may generate the auxiliary audio signal to (e.g. directly) include/copy the samples of the downmix audio signal whereas for samples that the transient signal indicates do not belong to a transient, the auxiliary audio signal circuit 107 may generate the auxiliary audio signal to (e.g. directly) include/copy the samples of the synthesized signal generated by the trained artificial neural network. Thus, in such a case, the auxiliary audio signal may be generated by selecting samples from the downmix audio signal or from the synthesized signal dependent on the value of the transient signal for the samples.
-
In many embodiments, the transient signal may be a non-binary signal and/or the auxiliary audio signal may be generated by a more complex and possibly indirect weighting of samples of the downmix audio signal in the generated auxiliary audio signal. For example, in some embodiments, a weighted combination may be performed between samples of the downmix audio signal and samples of the synthesized signal with the weights being dependent on the transient signal.
-
In some embodiments, the input data to the trained artificial neural network may include samples of the transient signal (or data values generated therefrom). The trained artificial neural network may be trained to generate a synthesized signal (that directly may be the auxiliary audio signal) where the contribution/component/proportion/part of the downmix audio signal which is included in the synthesized signal is dependent on the transient signal.
-
In many embodiments, the auxiliary signal circuit is arranged to generate the auxiliary audio signal to include a synthesized signal component which is generated/estimated by the trained artificial neural network signal and a downmix audio signal component which is a component of the downmix audio signal. The weight of the synthesized signal component relative to the downmix audio signal component can in such a case be made dependent on the transient signal.
-
The auxiliary audio signal circuit 107 may specifically be arranged to generate the auxiliary audio signal such that the contribution/component/proportion/part of the downmix audio signal in the auxiliary audio signal increases for the transient signal providing an indication of an increasing likelihood that the samples of the downmix audio signal represent/describe/reflect a transient.
-
The transient signal may thus be generated and used to indicate which parts of the original downmix audio signal should be preserved (e.g., transients) during a decorrelation process and in the auxiliary audio signal and which parts of the downmix audio signal should be replaced by a newly generated signal with some decorrelation to the downmix audio signal but with similar spectro-temporal properties. The transient signal can as such be considered a masking signal/function which is determined based on the properties of the downmix audio signal. For example, a value of the transient signal (which may also be referred to as a masking value) close to 1 may indicate a portion of the downmix audio signal that is not part of a transient and which should be replaced by trained artificial neural network generated sample values. In contrast, a masking value close to 0 may indicate that the auxiliary audio signal is better generated to be substantially identical to the downmix audio signal and typically with the signal samples of the auxiliary audio signal directly being copied from the downmix audio signal.
-
It will be appreciated that many different approaches and algorithms are known for detecting transients in an audio signal and that any suitable approach may be used by the transient circuit 109.
-
In many embodiments, the transient detection may be based on comparing (typically an average) signal level (e.g. power/amplitude/envelope level) determined for a shorter time interval to a signal level determined for a longer time interval (with the longer time interval often having a larger duration of at least 2,5, or 10 times the duration of the shorter time interval. The longer time interval may be a neighbor time interval of the shorter time interval, and/or may often be adjacent or surrounding the shorter time interval.
-
Thus, in many embodiments, the transient circuit 109 may be arranged to generate the transient signal from a comparison of a signal level of the downmix audio signal in a first time window to a signal level of the downmix audio signal in a second time window where the second time window exceeds the first time window (e.g. by a factor of at least 2, 5, 10 times). The signal level for a time interval may typically be an average, median, or mean value for the time interval.
-
One specific approach for generating the transient signal may be based on tracking fast changes in the signal envelope (either wideband or per frequency spectrum) relative to a slowly changing background. The difference between these signals may be used to detect onset and offsets of transients.
-
For example, the magnitude of the downmix audio signal's time-frequency representation may be determined, and the resulting frequency envelopes may be summed across frequency bands. Two smoothed versions of this envelope may then be created, with one tracking the envelope more slowly than the other. One envelope signal may represent the slowly varying background envelope, and the other may apply a time-constant that provides much faster tracking and thus also tracking of the signal levels of short term transients.
-
As a specific example, a first order exponential smoothing or a smoothing moving average filter may be used to create the smoothed envelope. In the case of the fast-tracking envelope, the instantaneous envelope can also be used. The envelope measures may specifically be determined as: where x̃s (n) and x̃f (n) are the slow and fast envelopes respective with αs, αf ∈ (0,1] and αf » αs. It will be appreciated that more generally two different filters with different low pass characteristics can be applied (with the filter for the slow envelope tracking having a stronger low pass filtering characteristic).
-
The ratio between the fast and slow-tracking envelopes can then be used to signal sharp changes with respect to the residual signal which provides an indication of the presence of a transient, and thus may be used as the transient signal. For example, the transient/mask signal may be generated as:
-
For transients, xf (n) » xs (n), and accordingly m(n) → 1. ε is a small constant to prevent division by zero. Thus, this approach may provide a transient signal m(n) which is close to zero for strong transients and substantially 0 if there is no transient.
-
In some embodiments the mask can be further binarized by thresholding (Th) on values of m(n) > 0, where m(n) = 1 if m(n) > Th, and zero otherwise. The value of Th can be set to 0.6, for example.
-
In other embodiments, an activation function such as a sigmoid function can be used to convert the value of m(n) in to a probability score between 0 and 1, e.g:
-
Another approach to detect transients is to track the background signal using minimum tracking of the envelopes (like the minimum statistics approach for stationary noise tracking) per frequency band (ref. e.g. Martin, R., 2001. Noise power spectral density estimation based on optimal smoothing and minimum statistics. IEEE Transactions on speech and audio processing, 9(5), pp.504-512).
-
The transient signal can then be written in the frequency domain as, where Xm (k, l) is the time-frequency representation of the (mono) downmix audio signal, Xs (k, l) is the time-frequency estimate of the background signal using minimum tracking and γ is an over-subtraction factor (≥ 1.0 ) to account for under-estimating the background signal. gm is a small constant > 0 to prevent negative value and division by zero errors.
-
As another example, the transient circuit 109 may include a trained artificial neural network which is arranged to receive the downmix audio signal and directly generate a transient signal. For example, an artificial neural network may be trained to generate a per-sample probability value indicating the likelihood of the samples belonging to a transient component or not based on labelled data or scores produced by the aforementioned methods.
-
In some embodiments, the transient circuit 109 may be arranged to generate the transient signal to indicate an increased duration for transient, and specifically the transient signal may be arranged to indicate that a transient starts before and/or ends after the actual detected start time and/or end time.
-
Thus, in many embodiments, the transient circuit 109 may be arranged to determine a time interval for which a transient is detected, such as e.g. a start time and an end time for which a given criterion is met (such as the criterion that the short term signal level exceeds the longer term signal level by a given factor). The transient circuit 109 may then be arranged to generate the transient signal to indicate the presence of the transient in a larger time interval than the time interval for which the transient is detected.
-
The transient circuit 109 may for example be arranged to generate the transient signal to indicate a transient for a minimum time interval and thus may offset the detected time instants to increase the time interval duration if this is below the minimum time interval.
-
As another example, the transient circuit 109 may be arranged to offset the detected time instants for the transient by a value that is dependent on the duration between the detected time instants, thereby e.g. increasing the duration of the transient time interval by a relative amount. As another example, the transient circuit 109 may simply offset the detected time instants by a predetermined, fixed mount thereby providing a margin for the detected transients.
-
As another example, the transient signal may be generated by convolving a window shape with a transient location indicator signal to increase the coverage of a transient component.
-
Such increases of the detected durations of transients may provide improved performance in many scenarios and have in particular been found to result in a perceived improved sound quality of the resulting upmixed signal, especially for transient audio.
-
In the approach, the audio apparatus thus includes an auxiliary audio signal circuit 107 which generates an auxiliary signal that is at least partially decorrelated with the downmix audio signal and with the upmixing being based on both the auxiliary audio signal and the downmix audio signal. The auxiliary audio signal is generated from the downmix audio signal using a trained artificial neural network and with the auxiliary audio signal being generated to include some components/elements of the original downmix audio signal with this being dependent on a transient signal that indicates the presence of transients in the signal.
-
An example of an auxiliary audio signal circuit 107 is illustrated in FIG. 3. In the example, the downmix audio signal is fed to a feature extractor 301 which is arranged to generate a set of features for the downmix audio signal. As a specific example, the feature extractor 301 may generate a spectrogram (e.g. a Mel spectrogram) for the downmix audio signal, for example by performing a frequency transformation for each processing time segment.
-
In other embodiments, e.g. another latent space feature set representing properties of the downmix audio signal may alternatively or additionally be used. In particular, the feature extractor could consist of an autoencoder that has been previously trained to produce latent features representing the properties of the downmix audio signal. In other embodiments where the feature extractor has been trained on a diverse set of signals, the latent feature can represent regions of the latent space depending on the type of downmix audio signal, e.g., applause, speech, etc. The latent feature can then be used by the artificial neural network to generate samples most similar to signals from this region of the latent space.
-
The feature extractor 301 is coupled to a trained artificial neural network 303 which is trained to generate a synthesized signal that is different from the downmix audio signal but which has properties that are derived from the downmix audio signal. The trained artificial neural network is typically arranged to generate a synthesized signal that is decorrelated with the downmix audio signal but which has the same or similar spectro-temporal properties. In particular, the artificial neural network 303 can be a generative model trained to learn a conditional data distribution, conditioned on the set of features generated by the feature extractor 301.
-
Spectro-temporal properties mainly typically include the spectral and temporal envelopes of the downmix and generated signals. The synthesized signal and/or the auxiliary audio signal may in many embodiments be generated to be transparent perceptually but produce a signal that is decorrelated with the input. The synthesized signal may be generated such that the time and spectral envelopes of this are similar to, and often substantially the same, as the time and spectral envelopes of the downmix audio signal. In many embodiments, the artificial neural network may be trained to generate a synthesized signal that has a similar time and(/or) spectral envelope to the time and(/or) spectral envelope of the downmix audio signal. However, the signal/sample correlation between the signals may be low, i.e. the signals may be decorrelated signals (despite having the same/similar envelop characteristics).
-
The auxiliary audio signal circuit 107 further comprises a combiner 305 which receives the downmix audio signal and the synthesized signal. The combiner 305 is in the example arranged to generate the auxiliary audio signal by combining the synthesized signal and the downmix audio signal and for this purpose the combiner 305 receives the transient signal which is used to control the combination.
-
In some embodiments, the combination by the combiner 305 may be a selection combining where the auxiliary audio signal is generated either as the downmix audio signal or as the synthesized signal depending on the transient signal. Specifically, for time intervals and samples for which the transient signal indicates that the audio signal does not represent a transient, the auxiliary audio signal is generated as the synthesized signal, i.e. the samples of the output signal generated by the artificial neural network are directly copied to the auxiliary audio signal. However, for time instants and samples for which the transient signal indicates that the audio signal does represent a transient, the auxiliary audio signal is generated as the downmix audio signal, i.e. the samples of the downmix audio signal are directly copied to the auxiliary audio signal. Thus, in this case, the auxiliary audio signal is generated as the artificial neural network output signal, and thus e.g. as a signal decorrelated with the downmix audio signal but having similar spectro-temporal properties, except for when transients are present in which case the auxiliary audio signal is generated to be identical to the downmix audio signal. Thus, the auxiliary audio signal is generated as a decorrelated signal with similar statistical properties as the downmix audio signal, but with the transient components of the downmix audio signal further being maintained in the auxiliary audio signal.
-
The auxiliary audio signal circuit 107 may in this way be arranged to generate the auxiliary audio signal to include a synthesized signal component estimated by the trained artificial neural network signal and a downmix audio signal component which is a component of the downmix audio signal, where the weight of the synthesized signal component relative to the downmix audio signal component is dependent on the transient signal.
-
In some embodiments, the auxiliary audio signal circuit 107 may as illustrated in FIG. 4 include a splitter 401 which is arranged to divide the downmix audio signal into at least a first and second signal component where the first signal component includes parts/components of the downmix audio signal that have faster changes and the second signal component includes parts/components of the downmix audio signal that have slower changes. The first signal component is thus generated to include faster changing parts of the downmix audio signal whereas the second signal component is generated to include slower changing parts of the downmix audio signal. The division into the two signals is based on the transient signal and the exact criterion for the splitting into the different signal components depend on the individual embodiment and the specific requirements and preferences of the individual embodiment. In many embodiments, the auxiliary audio signal circuit 107 may be arranged to generate the first signal component as the parts of the downmix audio signal for which the transient signal indicates that a transient is present (e.g. it indicates that the likelihood of the downmix audio signal representing a transient is above a given threshold) and the second signal component as the parts of the downmix audio signal for which the transient signal indicates that a transient is not present (e.g. it indicates that the likelihood of the downmix audio signal representing a transient is not above a given threshold). Thus, in some embodiments, the first signal component may be a transient signal component comprising transient components of the downmix audio signal and the second signal component may be a non-transient signal component comprising non-transient components of the downmix audio signal. The second signal component may be a residual signal after the extraction of the first signal component from the downmix audio signal.
-
In the example, the auxiliary audio signal circuit 107 may further include the features described with respect to FIG. 3 and thus may specifically include a combiner 305 that combines the output of the trained artificial neural network 303 and the downmix audio signal.
-
However, rather than generate inputs to the trained artificial neural network 303 from the downmix audio signal as such, the feature extractor 301 may be arranged to generate input data for the trained artificial neural network 303 from only the second signal component. Specifically, the feature extractor 301 may generate a latent space feature set from the second signal component and input this to the trained artificial neural network.
-
The combiner 305 may generate the auxiliary audio signal by combining the first signal component and the output signal from the trained artificial neural network. For example, the auxiliary audio signal may be generated to include the first signal component for time intervals which the transient signal indicates include transients, and the trained artificial neural network for time intervals which the transient signal indicates do not include transients.
-
As a specific example, the auxiliary audio signal circuit 107 may use the transient signal as a masking function splitting the downmix audio signal into a foreground (transient) and background (non-transient) component, i.e. where m represents the transient signal and xm represents the downmix audio signal (and the sample index n is left out for brevity).
-
The trained artificial neural network operation is then performed on the non-transient component (e.g. by setting the noise free version x 0 = xbgd (n) for a diffusion based network as described later). The result of the trained artificial neural network operation, denoted by x̂bgd (n) can then be recombined with xfgd (n) to produce the generated output sample,
-
The audio apparatus of FIG. 1 may specifically perform parametric stereo upmixing from a received encoded mono downmix audio signal and a set of parametric stereo parameters, with the encoded mono downmix audio signal first being decoded (e.g. by the receiver 101) to produce a time-domain downmix audio signal xm (n). The decoded mono downmix audio signal is analyzed to produce a transient/mask signal m(n). The decoded downmix audio signal, together with the transient signal and (alternatively or additionally) a feature representation of the downmix audio signal (e.g., Mel spectrogram or other latent representation) can be processed by a deep learning trained artificial neural network which synthesizes a decorrelated signal from which an auxiliary audio signal is generated (e.g. the trained artificial neural network may directly generate the auxiliary audio signal or a combination between the trained artificial neural network generated signal and the original downmix audio signal may be combined based on the transient/mask signal). The decoded mono downmix audio signal and the auxiliary audio signal are then processed in an upmix stage that is dependent on the parametric stereo parameters, resulting in a stereo output signal.
-
In many embodiments, the feature extractor 301 may be arranged to generate a latent space feature set which may be provided as input data to the trained artificial neural network of the auxiliary audio signal circuit 107. The input data, and specifically the latent space feature set, may for example be provided as feature inputs or the trained artificial neural network may be conditioned on the latent space feature set.
-
In some embodiments, the trained artificial neural network may receive input data generated from the downmix audio signal and the transient signal. For example, in some embodiments, the trained artificial neural network may directly receive samples of the downmix audio signal (or e.g. a latent space feature set generated therefrom) and samples of the transient signal with the artificial neural network being trained to therefrom directly generate a synthesized signal that can be used as the auxiliary audio signal. Specifically, it may be trained to directly generate a synthesized/auxiliary audio signal that is decorrelated with the downmix audio signal but which retains the transient components of the downmix audio signal and which additionally have similar statistical properties (specifically the same statistical spectro-temporal properties).
-
The trained artificial neural network may be implemented using different artificial neural network architectures, models, and processes in different embodiments.
-
In some embodiments, the trained artificial neural network may be a convolutional artificial neural network which for example may receive features (such as a spectrogram) representing the downmix audio signal as well as samples of the transient signal as inputs and may e.g. directly generate the auxiliary signal. In other examples, the artificial neural network may be applied only to the non-transient parts of the downmix audio signal and be used to generate samples of the auxiliary signal for these time intervals.
-
An artificial neural network as used in the described functions may be a network of nodes arranged in layers and with each node holding a node value. FIG. 5 illustrates an example of a section of an artificial neural network.
-
The node value for a given node may be calculated to include contributions from some or often all nodes of a previous layer of the artificial neural network. Specifically, the node value for a node may be calculated as a weighted summation of the node values of all the nodes output of the previous layer. Typically, a bias may be added and the result may be subjected to an activation function. The activation function provides an essential part of each neuron by typically providing a non-linearity. Such non-linearities and activation functions provides a significant effect in the learning and adaptation process of the artificial neural network. Thus, the node value is generated as a function of the node values of the previous layer.
-
The artificial neural network may specifically comprise an input layer 601 comprising a plurality of nodes receiving the input data values for the artificial neural network. Thus, the node values for nodes of the input layer may typically directly be the input data values to the artificial neural network and thus may not be calculated from other node values.
-
The artificial neural network may further comprise none, one, or more hidden layers 603, 605 or processing layers. For each of such layers, the node values are typically generated as a function of the node values of the nodes of the previous layer, and specifically a weighted combination and added bias followed by an activation function (such as a sigmoid, ReLU, or Tanh function may be applied).
-
Specifically, as shown in FIG. 6, each node, which may also be referred to as a neuron, may receive input values (from nodes of a previous layer) and therefrom calculate a node value as a function of these values. Often, this includes first generating a value as a linear combination of the input values with each of these weighted by a weight:
where w refers to weights, x refers to the nodes of the previous layer and n is an index referring to the different nodes of the previous layer.
-
An activation function may then be applied to the resulting combination. For example, the node value l may be determined as: where the function may for example be a Rectified Linear Unit (as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011) function:
-
Other often used functions include a sigmoid function or a tanh function. In many embodiments, the node output or value may be calculated using a plurality of functions. For example, both a ReLU and Sigmoid function may be combined using an activation function such as:
-
Such operations may be performed by each node of the artificial neural network (except for typically the input nodes).
-
The artificial neural network further comprises an output layer 607 which provides the output from the artificial neural network, i.e. the output data of the artificial neural network is the node values of the output layer. As for the hidden/ processing layers, the output node values are generated by a function of the node values of the previous layer. However, in contrast to the hidden/ processing layers where the node values are typically not accessible or used further, the node values of the output layer are accessible and provide the result of the operation of the artificial neural network.
-
A number of different networks structures and toolboxes for artificial neural networks have been developed and in many embodiments the artificial neural network may be based on adapting and customizing such a network. An example of a network architecture that may be suitable for the applications mentioned above is WaveNet by van den Oord et al which is described in
Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. "Wavenet: A generative model for raw audio." arXiv preprint arXiv: 1609.03499 (2016).
-
WaveNet is an architecture used for the synthesis of time domain signals using dilated causal convolution, and has been successfully applied to audio signals. For WaveNet the following activation function is commonly used: where * denotes a convolution operator, ⊙ denotes an element-wise multiplication operator, σ(·) is a sigmoid function, k is the layer index, f and g denote filter and gate, respectively, and W represents the weights of the learned artificial neural network. The filter product of the equation may typically provide a filtering effect with the gating product providing a weighting of the result which may in many cases effectively allow the contribution of the node to be reduced to substantially zero (i.e. it may allow or "cutoff" the node providing a contribution to other nodes thereby providing a "gate" function). In different circumstances, the gate function may result in the output of that node being negligible, whereas in other cases it would contribute substantially to the output. Such a function may substantially assist in allowing the artificial neural network to effectively learn and be trained.
-
An artificial neural network may in some cases further be arranged to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for a specific desired property or characteristics of the generated output. For example, a set of values may be provided to adapt the artificial neural network. These values may be included by providing a contribution to some nodes of the artificial neural network. These nodes may be specifically input nodes but may typically be nodes of a hidden or processing layer. Such adaptation values may for example be weighted and added as a contribution to the weighted summation/ correlation value for a given node. For example, for WaveNet such adaptation values may be included in the activation function. For example, the output of the activation function may be given as: where y is a vector representing the adaptation values and V represents suitable weights for these values.
-
The above description relates to an artificial neural network approach that may be suitable for many embodiments and implementations. However, it will be appreciated that many other types and structures of artificial neural network may be used. Indeed, many different approaches to artificial neural networks have been, and are being, developed including artificial neural networks using complex structures and processes that differ from the ones described above. The approach is not limited to any specific artificial neural network approach and any suitable approach may be used without detracting from the invention.
-
In many embodiments, the artificial neural network may advantageously be a generative model such as a diffusion model artificial neural network. This may provide a particularly efficient and high quality generation of an auxiliary signal. In particular, generative models may include generative adversarial networks, variational autoencoders, and diffusion models among others. The Inventors have realized that a diffusion model approach provides a particular efficient operation and addresses many issues of many conventional approaches and indeed provides particularly advantageous operation over other (more common) artificial neural network approaches and structures.
-
In the following, the description will focus on an approach where the artificial neural network is implemented using a diffusion-based generative deep learning model. Diffusion models have developed which are based on the diffusion process in thermodynamics, with a forward process of T steps that incrementally adds more and more Gaussian noise until a probability distribution is obtained that is standard normal Gaussian at the terminal time-step T, i.e., pT = N(0, I) where I is the identity matrix. During the reverse or generation process, an artificial neural network is trained to learn the standard normal Gaussian component at each time step and remove it, leaving an estimate of a sample from the true data distribution p 0.
-
At a high level, a typical diffusion model may during training start with a desired signal and through a series of steps/iterations, noise may be added until a given noise level is reached. The given noise level is typically such that it closely resembles a noise only sequence, such as a fully random Gaussian distribution. The diffusion model is then trained in reverse by the diffusion model being provided with a noise signal (typically a Gaussian noise - or indeed the resulting signal after adding noise may be the input signal) and with a series of steps/iterations each removing some noise. After a number of iterations, the resulting signal is compared to the original signal to provide a loss function and the diffusion model is based on this trained to minimize the loss function for a large set of audio signals and noise.
-
A typical diffusion model can be implemented using a denoising neural network that estimates the noise component of an input signal. During training, a desired signal is corrupted with a known level of noise, where a noise scaling factor is randomly selected from a diffusion noise schedule which is used to scale a standard Gaussian noise signal with the same dimension as the desired signal. Together with the noise level which is input to the neural network together with the corrupted signal, the parameters of the neural network are then optimized such that the network is able to estimate the original standard normal noise signal. By training on batches of randomly selected desired signals, standard Gaussian noise instances, and noise scales, the network learns to denoise a corrupted desired signal at each stage of the diffusion process. To optimize the network parameters a loss function compares the estimated standard Gaussian noise instances to their ground-truth values which are known when the desired signal is corrupted.
-
During a generation phase, such a trained network may then be started with a noise signal which is drawn from a standard distribution, and from this through a series of iterations that each remove some noise, generate/synthesize a signal. Essentially, if the neural network has learned to estimate noise correctly, it has implicitly learned the underlying data distribution. Typically, in each step/iteration, the diffusion model may generate an estimate of the total amount of noise present for that step/iteration. However, only a part of the estimated noise is removed for each step/iteration (except for the last step/iteration). Thus, through a series of iterations/(time) steps, estimated noise components may be removed with the resulting effect being that an estimated/synthesized signal is generated from the original noise. Thus, in the diffusion model the first step/iteration may substantially only be noise and with the noise level then gradually being reduced during each iteration until the last iteration/step of the diffusion model provides the estimate of the noise free signal.
-
Thus, generation is performed after training a de-noising artificial neural network to estimate and remove noise at steps of the reverse diffusion process (note that steps can be discrete or sampled from a continuum):
-
In the approach the down-mix signal may be considered the original non-noise signal, i.e. x 0 = xm. It should be noted that while t is defined as a timestep, it actually represents a step or instant that may also be coupled to a noise level (as will be discussed later).
-
Diffusion-based generative models rely on a predefined schedule of weights βt that controls the diffusion process at timestep t also known as the noise schedule), e.g. where ε ~ N(0, I) is the stochastic part of the process and the forward transition has mean and standard deviation . The parameter βt thus controls the amount of noise added/removed at each step/iteration.
-
If we let αt = 1 - βt and the cumulative product of α up to (time)step t, , then xt can be re-parameterized in terms of x 0, making it easier to draw samples for various t. An example of the corresponding structure and diffusion process during forward and reverse steps is illustrated in FIG. 7.
-
Different approaches are known and suitable for training diffusion model artificial neural networks. For example, during training, a timestep t may be selected randomly (e.g. based on a uniform distribution), and xt may be constructed using the above formula. The denoising artificial neural network has its parameters θ optimized to produce an estimate ε̂θ from inputs xt and t (embedding) that resembles ε. In other words, the goal of the denoising artificial neural network is to estimate the noise component added at time t-1 to produce the noisy signal at time t.
-
During the sampling or generation process, the mean value of the reverse process distribution at time-step t - 1 as a function of x 0 and xt is estimated, where ε̂θ (xt, t) is an estimate of the noise component produced by the artificial neural network with parameters θ optimized during the training step.
-
It should be noted that the diffusion model artificial neural network facilitates estimating the backward transition's mean, µq (xt, x 0) but that it is further desired to determine the variance estimate, for the iteration/step. This estimate may be given by,
-
Therefore, the estimate for x t-1 for the next time step in the generation process may be given by, where ε again is a standard normal random variable (vector).
-
It should also be noted that there are numerous approaches for diffusion-based models. The current description focuses on denoising networks/models, but it will be appreciated that other networks/models/structures can be used and can be trained to estimate x 0. Some may be based on modelling both forward and reverse diffusion processes described by stochastic differential equations (SDE), while others may e.g. convert the SDE into a deterministic ordinary differential equation (ODE) for faster generation/sampling.
-
In many embodiments, a diffusion model trained artificial neural network may be used that is only considering features of the downmix audio signal with the transient signal being used to process how this process is used to generate the auxiliary audio signal. For example, as previously described with reference to FIGs. 3 and 4, the transient signal may be used to combine or split signals but with the trained artificial neural network considering only features/properties of the downmix audio signal and not of the transient signal. For example, a diffusion model conditioned on a latent space feature set generated for the downmix audio signal (or part thereof) may be used.
-
However, in many embodiments, the trained artificial neural network and diffusion model may also receive input data derived from the transient signal. For example, features determined from the transient signal, or in many cases the transient signal directly, may be received and processed by the trained artificial neural network.
-
As a specific example, the transient signal m (the discrete time index n will for brevity not always be explicitly listed in the following description) may together with e.g. a spectrogram representation of the downmix audio signal be provided to a diffusion model neural network.
-
The equation for the step/iteration t of the diffusion model may in this case be modified as follows: where ⊙ denotes the element-wise product, and again with x 0 = xm.
-
In such a case, the diffusion model may e.g. be trained using the loss function of (shown using an L2 distance measure but of course other distance measures, such as an L1-based distance, may be used in other embodiments) for the denoising artificial neural network is given by:
-
In some embodiments, the loss function may be dependent on the transient signal, and in particular the loss value may be scaled based on the transient signal.
-
In some embodiments, the auxiliary audio signal may be arranged to generate the input data, and specifically input data to a diffusion model, to include a combination of samples of the downmix audio signal and noise samples, where the combination is dependent on the transient signal. In particular, the weighting of the noise samples relative to samples of the downmix audio signal is increased for the transient signal indicating a reduced likelihood of the downmix audio signal representing a transient.
-
Specifically, during sampling, the starting point/step may be based on a mix of Gaussian noise and relevant samples from the original downmix audio signal that are left untouched due to the transient signal indicating that the downmix audio signal represents a transient. For example: where xT is the input signal samples for the trained artificial neural network (the first step of the diffusion model), x 0 represents samples of the downmix audio signal, and ε represent noise samples (typically from a standard white Gaussian distribution)
-
For such an approach, for m -> 0, xT = ε, while xT = x 0 for m = 1, i.e. if the transient signal m indicates that the downmix audio signal represents a transient (i.e. m=1) the initial signal samples of the diffusion model are identical to the downmix audio signal and for the transient signal indicating that the downmix audio signal is not representing a transient the input signal is set to a pure noise/random signal.
-
The trained artificial neural network may then proceed to evaluate the diffusion model by iterating from timestep t to t - 1, until a clean estimate of x 0 is generated. If the diffusion model is trained to provide a decorrelated signal with similar statistical/spatial-temporal properties, an advantageous auxiliary audio signal with improved upmixing properties can be achieved while mitigating or preventing much of the smearing and a quality degradation known from many existing approaches.
-
The noise samples for the diffusion model may be generated as random signal values in accordance with a suitable distribution and noise level. For example, a random number generator with a uniform distribution may generate samples and a suitable function may convert these values to a desired distribution. In many embodiments, the noise samples may be generated to be samples from a standard normal distribution.
-
The applied diffusion model may be arranged to process one or more of the individual (time) steps/iterations taking into account the transient signal. The transient signal may e.g. be used to modify different aspects such as the estimated signal for the given operation, the signal that is forwarded to the next step/iteration, the noise samples, a weighting between the signal estimates and the original downmix audio signal, etc.
-
In many embodiments, a mean value for the signal being estimated/synthesized/generated by the diffusion model at a given (time) step/iteration t may be generated in dependence on the transient signal as well as on a noise estimate generated at the given (time) step/iteration t.
-
For a given iteration/time step, the diffusion model may be arranged to receive a noisy signal estimate from the previous iteration/step with the noise level being gradually reduced in the iterations/steps. The diffusion model may then proceed to generate a noise estimate for the total noise at the given iteration. The noise estimate may in many embodiments be scaled to reflect the noise energy/level at the given iteration. An estimate of the mean signal at iteration t may be generated by subtracting a (possibly scaled) version of the estimated noise from the generated signal estimate xt of the iteration. The scaling may be dependent on the expected noise variance/level at the specific step, and on the design parameter βt reflecting how much noise is removed in each step.
-
For example, the auxiliary audio signal circuit 107 may as previously described proceed to determine the following signal estimate for the current iteration t:
-
However, in some embodiments, the mean signal estimate for the iteration/step is further modified based on the transient signal m. Specifically, it may be arranged to generate the mean signal estimate as a combination of the signal estimate as generated above and of the generated signal estimate xt of the iteration. The auxiliary audio signal circuit 107 may be arranged to weigh the contributions based on the transient signal, such as e.g. by determining the mean value µq (xt, x 0 , c) as:
-
The first term of the summation preserves the regions of the signal that are transients. It is equivalent to x 0 ⊙ m for m equal to 1, i.e. in each step only this operation may be performed and contribute to the output. The second term corresponds to the decorrelated signal at timestep t. Because these samples include noise, the generated samples will be uncorrelated with the original downmix audio signal xm assuming the network has been trained on sufficient data.
-
In some embodiments, the diffusion model for a given iteration/step may be arranged to determine the signal estimate for the first signal as a combination of a signal estimate from the previous iteration of the given iteration and a noise compensated signal estimate for the given iteration. The noise compensated signal estimate may be determined from the noise estimate for the given iteration. Further, the combination is dependent on the transient signal.
-
The diffusion model for a given step/iteration is further arranged to generate a mean signal estimate for the next step/iteration. This signal is in accordance with the fundamental approach of diffusion models generated to include a noise component. In the specific case, the signal estimate for the next iteration may be generated as the combination of the (no noise) signal estimate for the current iteration and a noise component. The level of the noise component is dependent on the iteration/step (as noise is reduced for each step/iteration in order to reach a denoised signal at the output of the diffusion model). In some embodiments, the trained artificial neural network may further be arranged to adapt the noise component dependent on the transient signal, and specifically the level/amplitude of the noise contribution at a given iteration/level may be scaled/controlled by the transient signal.
-
As a specific example, the signal estimate passed on to the next iteration/step may specifically be determined by re-adding noise to the noise-free signal estimate of the current iteration, such as e.g. by where x t-1 is the input for the next iteration, µq (xt, x 0 , c) is the noise free estimate for the current iteration, m is the transient signal which remains fixed throughout the generation process, and βt is a design parameter and α f and α t-1 are the noise signal levels for the current and next iterations respectively as specified by a noise schedule.
-
It should be noted that if the specific equations provided above are implemented, then in the case where m=1, i.e. with the transient signal indicating that a transient is present, the downmix audio signal represented by x 0 will be carried though each iteration/step without any changes and the synthesized signal will be equal to x 0, i.e. to the downmix audio signal. Thus, for the transient signal indicating that a transient is definitely present, the synthesized signal is directly generated as the downmix audio signal. Thus, transients are maintained.
-
In some embodiments, the diffusion model may be arranged to generate multiple possible signal estimates xt-1 for the next iteration. It may then select a specific estimate of these possible estimates based on how correlated the estimates are to the downmix audio signal.
-
Specifically, the auxiliary audio signal circuit 107 may generate a plurality of noise components and generate a signal estimate xt-1 for each of these. Specifically, the above equation may be evaluated for a plurality of different versions of the noise component ε.
-
The resulting plurality of possible signal estimates may then be compared to the downmix audio signal and a correlation value may be determined for each signal estimate. The auxiliary audio signal circuit 107 may then proceed to select one of the signal estimates dependent on the correlation value, and specifically it may select the signal estimate which results in the lowest correlation with the downmix audio signal.
-
In particular, in embodiments where parallel processing is available, and particularly related to the diffusion models, at each time-step t of the signal generation process, one or more instances of additive noise can be generated to produce the input of the model at time step t-1, the sampling process can then continue with the instance that produces an output that has the lowest coherence with respect to the downmix audio signal.
-
In some cases, the auxiliary audio signal circuit 107 may be arranged to initialize the diffusion model with a noise level that is dependent on the transient signal. The noise level is gradually reduced for each new iteration/step of the diffusion model starting from an initial noise level for the first step/iteration of the diffusion model. In many embodiments, the initial level may be dependent on the transient signal and thus may be different for transient parts of the downmix audio signal than for non-transient parts. In many cases, the noise level may be decreased for the transient signal being indicative of an increased likelihood that the downmix audio signal represents a transient.
-
In some embodiments, the trained artificial neural network may be a denoising artificial neural network that is conditioned on a noise schedule/level for different steps/iterations rather than on a timestep t. For example, it may be conditioned on the scaling factor
representing the noise level at the different steps/iterations. This may often be advantageous since it allows sampling from a continuous distribution which may improve the robustness and performance of the trained artificial neural network and diffusion model.
-
Such conditioning may also allow the transient signal to be represented as/by a noise variance which may be used to generate a decorrelated auxiliary audio signal while preserving the transient components. Since this noise level is estimated and therefore known, the reverse sampling process does not need to begin at timestep t = T, but rather based on the scaling factor (or the noise scale . Therefore, the reverse noise schedule can be dynamically started at a specific noise level based on the estimated transient signal , e.g., the reverse process can be started with xα, where δ is an indicator function that equals 1 when the corresponding sample of x 0 satisfies the condition in parentheses and γ is a scaling factor.
-
In some embodiments, the auxiliary audio signal circuit 107 may be arranged to vary a noise level between sample times in dependence on an energy level of the downmix audio signal for the sample times and the transient signal for the sample times.
-
Thus, the noise level that is applied in the diffusion model (such as specifically the initial noise level) may be adapted depending on both the signal level of the downmix audio signal and on the transient signal. For example, a given function for determining an initial noise level, or a noise level profile for multiple steps, may be used when the transient signal indicates that a transient is present, and a different function may be used when the transient signal indicates that a transient is not present.
-
In some embodiments, diffusion models may be used in which the standard normal Gaussian noise is dynamically varied. For example, a low pass envelope signal may be estimated by low pass filtering the downmix audio signal and the variance of the Gaussian distribution may be varied to follow variations in this envelope. Such approaches may help preserve more long-term envelope information in the generated auxiliary audio signal. Further, due to the consideration of the transient signal, this can be achieved while mitigating any impact on the dynamic performance and specifically typically without smearing or removing transient effects and properties.
-
The above description has focused on a generative model/artificial neural network based on a diffusion generative model, which specifically may be conditioned on a Mel spectrogram of the input signal. However, the diffusion model may e.g. alternatively or additionally be conditioned on features embeddings. In particular, the codes learned by a Vector-Quantized Variational AutoEncoder (VQ-VAE) or latent embeddings by a VAE can be used to condition the diffusion model (hybrid embodiment). The VQ-VAE may first be trained to regenerate its input using a vector-quantizer bottleneck. The latent codes that are learned by the VQ represent centroids of regions in the latent space and can be viewed as rich classification information about the different signal types, i.e. input signals are mapped to points in the latent space that belong to regions identified by their centroids. The diffusion model can be trained with these codes as conditional inputs to guide the generation process. During sampling, the downmix signal is first input to the VQ-VAE encoder to produce the latent code. This code is then supplied to the diffusion model as a conditioning signal during generation. Unlike the VQ-VAE which produces discrete codes at its bottleneck, a VAE produces parametrizations of the Gaussian distribution (mean and variance). In another embodiment, instead of using a diffusion model for the generator, another generative model such as a VAE or generative adversarial network (GAN) generator can be used, e.g. conditioned on the centroid code returned by the VQ-VAE for the same input. Using a VAE or GAN can be advantageous in that they are able to generate samples faster, albeit typically of lesser quality than diffusion models. Their training procedures may be somewhat different as well.
-
Inherently both types of generators first sample a latent prior from a standard distribution like the Gaussian normal and combining this latent vector with the conditioning signal can thus similarly produce a synthesized signal. An example of elements of such an embodiment is illustrated in FIG 8.
-
FIG. 8 specifically shows an example of an artificial neural network approach for the background signal of the downmix audio signal (i.e. excluding transients), where a generator 801 (which may include a trained neural network, such as specifically a diffusion model, VAE, or GAN) is conditioned on a Vector Quantizer code learned by a VQ-VAE 803. The decoder 805 of the VQ-VAE 803 is not used during operation but is present to allow training of the VQ-VAE 803. Thus, while it is used during training of the VQ-VAE, it is not used to generate the synthesized background signal.
-
In some embodiments, the auxiliary audio signal circuit 107 may be arranged to introduce a random component to the transient signal. In particular, in many embodiments, the auxiliary audio signal circuit 107 may introduce a stochastic element where sometimes a transient signal indicating that a transient is present may be modified to indicate that no transient is present (or more generally, the transient signal may be modified to indicate that the likelihood of the downmix audio signal representing a transient is less likely). E.g., in some embodiments, the transient signal conditioning the trained artificial neural network/diffusion model may (occasionally) be modified to indicate that the downmix audio signal is not a transient even if a transient is in fact present.
-
The auxiliary audio signal circuit 107 may thus in some embodiments be arranged to add a random change to the transient signal for sample instants for which the transient signal meets a criterion indicative of the transient signal indicating a presence of a transient.
-
For example, for a binary transient signal, an indication of a transient may be modified to indicate that no transient is present. The change may in some embodiments be performed for the entire transient time interval, or for one or more specific subintervals thereof, or indeed in some cases for individual samples. For example, for each sample of the transient signal indicating the presence of a transient signal, a random decision (with a given probability) may be made whether to retain the indication or to change it to indicate that no transient is present.
-
The probability of a change of the transient signal may be fixed and predetermined or may e.g. in some embodiments be dynamically changed. For example, it may be dependent on how frequent (and long) transients are. As a specific example, the probability of a change of the transient signal may increase for an increasing presence of transients in the downmix audio signal. In some cases, an average value of the transient signal may be determined, and the probability may be changed to increase or decrease the average (of the modified transient signal) to approach a desired value.
-
For the above specific example, the samples for which the transient signal m = 0, a toggle/switch/change may be applied with a specific probability, i.e. with a certain probability pswitch, the value of a 0 is changed into a 1.
-
Such an approach may for example reduce the risk that an auxiliary audio signal is generated which is too correlated with the downmix audio signal. For example, if the transient signal includes too many transients, the previously described approach may have a risk of maintaining too much of the auxiliary audio signal to be similar (or the same) as the downmix audio signal resulting in a lack of desired decorrelation. The approaches of modifying the transient signal may mitigate and/or prevent such effects. It may be particularly suitable for signals for which there are many and frequent transients.
-
In some embodiments, the auxiliary audio signal circuit 107 may alternatively or additionally introduce a stochastic element where sometimes a transient signal indicating that a transient is not present may be modified to indicate that a transient is present (or more generally, the transient signal may be modified to indicate that the likelihood of the downmix audio signal representing a transient is more likely). E.g., in some embodiments, the transient signal conditioning the trained artificial neural network/diffusion model may (occasionally) be modified to indicate that the downmix audio signal is a transient even if a transient is not in fact present.
-
The auxiliary audio signal circuit 107 may thus in some embodiments be arranged to add a random change to the transient signal for sample instants for which the transient signal meets a criterion indicative of the transient signal indicating a presence of a transient.
-
For example, for a binary transient signal, an indication of a transient may be modified to indicate that a transient is present. The change may in some embodiments be performed for the entire transient time interval, or for one or more specific subintervals thereof, or indeed in some cases for individual samples. For example, for each sample of the transient signal indicating the absence of a transient signal, a random decision (with a given probability) may be made whether to retain the indication or to change it to indicate that no transient is present.
-
The probability of a change of the transient signal may be fixed and predetermined or may e.g. in some embodiments be dynamically changed. For example, it may be dependent on how frequent (and long) transients are. As a specific example, the probability of a change of the transient signal may increase for an increasing presence of transients in the downmix audio signal. In some cases, an average value of the transient signal may be determined, and the probability may be changed to increase or decrease the average (of the modified transient signal) to approach a desired value.
-
For the above specific example, the samples for which the transient signal m = 0, a toggle/switch/change may be applied with a specific probability, i.e. with a certain probability pswitch, the value of a 0 is changed into a 1.
-
Such an approach may for example reduce the risk that an auxiliary audio signal is generated which is too correlated with the downmix audio signal. For example, if the transient signal includes too many transients, the previously described approach may have a risk of maintaining too much of the auxiliary audio signal to be similar (or the same) as the downmix audio signal resulting in a lack of desired decorrelation. The approaches of modifying the transient signal may mitigate and/or prevent such effects.
-
Artificial neural networks are adapted to specific purposes by a training process which is used to adapt/ tune/ modify the weights and other parameters (e.g. bias) of the artificial neural network. It will be appreciated that many different training processes and algorithms are known for training artificial neural networks. A training setup that may be used for the described artificial neural networks is illustrated in FIG. 9.
-
Typically, training is based on large training sets where a large number of examples of input data are provided to the network. Further, the output of the artificial neural network is typically (directly or indirectly) compared to an expected or ideal result. A cost function may be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost (or loss) function often represents the distance between the prediction and the ground truth for a particular input data. Based on the cost function, the weights may be changed and by reiterating the process for the modified weights, the artificial neural network may be adapted towards a state for which the cost function is minimized.
-
Different approaches may be used to train the artificial neural network of the auxiliary audio signal circuit 107 and in particular an overall training may seek to result in the output of the audio apparatus being a multichannel audio signal that most closely correspond to the original multichannel audio signal. Thus, the trained artificial neural network may be trained to provide an auxiliary audio signal that most effectively results in accurate reconstruction of the multichannel audio signal. In such a case, a cost function based on the difference between a generated multichannel audio signal and an original training multichannel audio signal may be used. In other cases, more specific training of an artificial neural network may be used. For example, the training of the trained artificial neural network may be based on a cost function reflecting the difference between properties of the generated auxiliary audio signal and the downmix audio signal. Specifically, a cost function may be generated increases for an increasing statistical/spectro-temporal difference between the downmix audio signal and the signal generated by the trained artificial neural network. Specifically, the cost function may determine a frequency distribution (an amplitude distribution) for the downmix audio signal and for the generated signal in each time segment, and then determine a contribution to the cost function dependent on the difference between these, and typically as an increasing value for increasing difference. The training process may e.g. perform a transformation of the generated signal into frequency bins (e.g. using an FFT) and compare this to corresponding frequency bin values determined for the training downmix audio signal. Further, in many embodiments, the cost function may include a contribution which is dependent on the correlation between the generated signal and the downmix audio signal, and specifically an increasing contribution may be determined for a decreasing correlation. Such approaches may train the trained artificial neural network to generate a synthesized signal that has statistical and/or spectro-temporal properties that are similar to the downmix audio signal but with the cross-correlation being low. Thus, a decorrelated version of the downmix audio signal may in many cases be generated.
-
In more detail, during a training step an artificial neural network may have two different flows of information from input to output (forward pass) and from output to input (backward pass). In the forward pass, the data is processed by the artificial neural network as described above while in the backward pass the weights are updated to minimize the cost function. Typically, such a backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth for a batch of data input, one can estimate the direction in which the cost function is minimized and propagate backward, by updating the weights accordingly. Other approaches known for training artificial neural networks include for example Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method etc.
-
In the present case, training may specifically include a training set comprising a potentially large number of multichannel audio signals or corresponding downmix audio signals. The training sets may include audio signals representing a number of different audio sources including e.g. recording of videos, movies, telecommunications, etc. In some embodiments, the training data may even include non-audio data such as a training being performed in combination with training data from other sources, such as text data etc.
-
In some embodiments, training data may be multichannel audio signals in time segments corresponding to the processing time intervals of the trained artificial neural network being trained, e.g. the number of samples in a training multichannel audio signal may correspond to a number of samples corresponding to the input nodes of the artificial neural network(s) being trained. Each training example may thus correspond to one operation of the artificial neural network(s) being trained. Usually, however, a batch of training samples is considered for each step to speed up and smoothen the training process. Furthermore, many upgrades to gradient descent are possible also to speed up convergence or avoid local minima in the cost function landscape.
-
For each training multichannel audio signal, a training processor may perform a downmix operation to generate a downmix audio signal and corresponding upmix parametric data. Thus, the encoding process that is applied to the multichannel audio signal during normal operation may also be applied to the training multichannel audio signal thereby generating a downmix and the upmix parametric data.
-
Based on the cost value, the training processor may adapt the weights of the artificial neural networks. For example, a back-propagation approach may be used. In particular, the training processor may adjust the weights of one or all of the artificial neural networks based on the cost value. For example, given the derivative (representing the slope) of the weights with respect to the cost function the weights values are modified to go in the direction of the slope. For a simple/minima account one can refer to the training of the perceptron (single neuron) in case of backward pass of a single data input.
-
The process may be iterated until the artificial neural networks are considered to be trained. For example, training may be performed for a predetermined number of iterations. As another example, training may be continued until the weights change be less than a predetermined amount. Also very common, a validation stop is implemented where the network is tested again a validation metric and stopped when reaching the expected outcome.
-
Specifically, for the trained artificial neural network implementing a diffusion model, the training process includes randomly selecting timestamps or corresponding noise scaling values for each of the training examples in a given batch. For each training example in a batch, an instance of a standard Gaussian noise signal (can also correspond to some other standard distribution) is generated with the same dimension as the training example. Then after scaling each of the noise signals with their corresponding noise scaling values, these are added to the examples in the training batch. The model then processes this batch of noisy data with additional inputs of the noise scaling factors and features to estimate the original Gaussian noise signals. Therefore the loss between the estimated and real noise signals are back propagated through the network to optimize their weights using an optimizer such as the Adam optimizer.
-
The described approach may provide advantageous operation in many embodiments and scenarios. Conventional decorrelators used for upmixing of downmix audio signals tend to introduce smearing and coloration artefacts in their output signals, particularly for fast changing signals (temporal smearing) and broadband noise-like signals (coloration).
-
To further illustrate the effects of a conventional decorrelator, a time-frequency decomposition of a mono (downmix) and decorrelated applause stereo signal sampled at 44.1 kHz (spectrogram) is shown in Fig. 10. The signal consists of background and foreground applause, where the background applause is dominant below 3 kHz. The lower signal (decorrelator output) sees its transient components smeared out and furthermore certain spectral components are delayed with respect to the original signal. This leads to improved decorrelation, but at the cost of poor reconstruction of IID components for the sharp changes which require that their temporal sharpness and temporal occurrence are preserved. This is clearly not the case for the decorrelated signal.
-
It is desirable for a decorrelator to preserve spectro-temporal properties of such signals (i.e., spectro-temporal envelope) as this may result in an improved reconstructed stereo signal at the decoder. The described approach provides a decorrelator that may process the (e.g. mono downmix audio signal) a way that preserves the sharpness of transients, while producing a decorrelated background stereo image.
-
It will be appreciated that the processing may be performed in time segments and that for example the trained artificial neural network and specifically the diffusion model may be evaluated once in each time segment. In some cases, such time segments may have a duration not less than lor 10 msec and not exceeding 1000 msec.
-
The audio apparatus(s) may specifically be implemented in one or more suitably programmed processors. In particular, the artificial neural networks may be implemented in one more such suitably programmed processors. The different functional blocks, and in particular the artificial neural network, may be implemented in separate processors and/or may e.g. be implemented in the same processor. An example of a suitable processor is provided in the following.
-
FIG. 11 is a block diagram illustrating an example processor 1100 according to embodiments of the disclosure. Processor 1100 may be used to implement one or more processors implementing an apparatus as previously described or elements thereof (including in particular one more artificial neural network). Processor 1100 may be any suitable processor type including, but not limited to, a microprocessor, a microcontroller, a Digital Signal Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA has been programmed to form a processor, a Graphical Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC) where the ASIC has been designed to form a processor, or a combination thereof.
-
The processor 1100 may include one or more cores 1102. The core 1102 may include one or more Arithmetic Logic Units (ALU) 1104. In some embodiments, the core 1102 may include a Floating Point Logic Unit (FPLU) 1106 and/or a Digital Signal Processing Unit (DSPU) 1108 in addition to or instead of the ALU 1104.
-
The processor 1100 may include one or more registers 1112 communicatively coupled to the core 1102. The registers 1112 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and/or any memory technology. In some embodiments the registers 1112 may be implemented using static memory. The register may provide data, instructions and addresses to the core 1102.
-
In some embodiments, processor 1100 may include one or more levels of cache memory 1110 communicatively coupled to the core 1102. The cache memory 1110 may provide computer-readable instructions to the core 1102 for execution. The cache memory 1110 may provide data for processing by the core 1102. In some embodiments, the computer-readable instructions may have been provided to the cache memory 1110 by a local memory, for example, local memory attached to the external bus 1116. The cache memory 1110 may be implemented with any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS) memory such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), and/or any other suitable memory technology.
-
The processor 1100 may include a controller 1114, which may control input to the processor 1100 from other processors and/or components included in a system and/or outputs from the processor 1100 to other processors and/or components included in the system. Controller 1114 may control the data paths in the ALU 1104, FPLU 1106 and/or DSPU 1108. Controller 1114 may be implemented as one or more state machines, data paths and/or dedicated control logic. The gates of controller 1114 may be implemented as standalone gates, FPGA, ASIC or any other suitable technology.
-
The registers 1112 and the cache 1110 may communicate with controller 1114 and core 1102 via internal connections 1120A, 1120B, 1120C and 1120D. Internal connections may be implemented as a bus, multiplexer, crossbar switch, and/or any other suitable connection technology.
-
Inputs and outputs for the processor 1100 may be provided via a bus 1116, which may include one or more conductive lines. The bus 1116 may be communicatively coupled to one or more components of processor 1100, for example the controller 1114, cache 1110, and/or register 1112. The bus 1116 may be coupled to one or more components of the system.
-
The bus 1116 may be coupled to one or more external memories. The external memories may include Read Only Memory (ROM) 1132. ROM 1132 may be a masked ROM, Electronically Programmable Read Only Memory (EPROM) or any other suitable technology. The external memory may include Random Access Memory (RAM) 1133. RAM 1133 may be a static RAM, battery backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology. The external memory may include Electrically Erasable Programmable Read Only Memory (EEPROM) 1135. The external memory may include Flash memory 1134. The External memory may include a magnetic storage device such as disc 1136. In some embodiments, the external memories may be included in a system.
-
The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and/or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
-
Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the accompanying claims. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in accordance with the invention. In the claims, the term comprising does not exclude the presence of other elements or steps.
-
Furthermore, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and/or advantageous. Also the inclusion of a feature in one category of claims does not imply a limitation to this category but rather indicates that the feature is equally applicable to other claim categories as appropriate. Furthermore, the order of features in the claims do not imply any specific order in which the features must be worked and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. In addition, singular references do not exclude a plurality. Thus references to "a", "an", "first", "second" etc do not preclude a plurality. Reference signs in the claims are provided merely as a clarifying example shall not be construed as limiting the scope of the claims in any way.