EP4639533A1 - Method and decoder for stereo decoding with a neural network model - Google Patents
Method and decoder for stereo decoding with a neural network modelInfo
- Publication number
- EP4639533A1 EP4639533A1 EP23828198.4A EP23828198A EP4639533A1 EP 4639533 A1 EP4639533 A1 EP 4639533A1 EP 23828198 A EP23828198 A EP 23828198A EP 4639533 A1 EP4639533 A1 EP 4639533A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio signal
- mono audio
- mono
- stereo
- neural network
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
Definitions
- the present application relates to a method and decoder for stereo decoding, and particularly stereo reconstruction using a neural network model.
- Stereo audio is used to present many different types of audio content (e.g. music) and is suitable for rendering to earphones, stereo loudspeaker pairs or even surround sound loudspeaker arrangements with more than two loudspeakers using various upmixing techniques.
- Stereo audio consists of two audio signals, e.g. a left audio signal and a right audio signal, which together form a stereo pair.
- a left and right audio signal can be recorded using two microphones which are spatially displaced and/or using two directional microphones which are directed in different directions (e.g. at 90 degree angles).
- stereo audio can be used to produce immersive, three dimensional, spatial effects giving a listener a sense of direction in a rendered audio scene.
- a user listening to stereo audio via earphones can perceive that the source of the audio content is somewhere between the ears of the listener (with the source moving with the panning of the audio signal) or that the source is outside of the user (with the source moving as the left and right audio signals are provided with a relative delay or processed with a head related transfer function (HRTF)).
- HRTF head related transfer function
- stereo audio may be represented with a mid audio signal and a side audio signal forming a mid-side stereo pair.
- a mid-side audio pair can be captured or created.
- a left and right stereo pair can be converted into a mid-side stereo pair or mid-side stereo pair can be recorded using i an omnidirectional or forward directed microphone (recording the mid audio signal) and sidewards directed microphone recording the side audio signal.
- a benefit with a mid-side stereo pair is that the mid audio signal usually captures the most essential audio content making the mid-side stereo pair backwards compatible with mono playback systems which simply disregard the side audio signal and renders only the mid-audio signal.
- a stereo audio signal comprising two audio signals, carries more information than a mono audio signal meaning that e.g. transmission of a stereo audio signal requires higher bitrate and that storage of a stereo audio signal requires a larger data volume.
- encoders have been proposed which obtain a left and right stereo pair, converts it into a mid and side stereo pair and encodes the mid audio signal as a downmix audio signal which is transmitted to the decoder along with some side parameters indicating the correlation between the left and right audio signal.
- the decoder decodes the downmix mid audio signal and converts it to a left and right stereo pair guided by the side parameters.
- the downmix mid audio signal is passed through an all-pass filter with filter parameters selected to introduce a fixed temporal delay to generate a synthetic side signal from the downmix mid audio signal.
- An all-pass filter with fixed delays has proven to be a suitable method for producing a synthetic, yet convincing, side audio signal from a mid audio signal wherein the side audio signal has approximately the same temporal and spectral energy distribution as the downmix mid audio signal.
- the downmix mid audio signal and synthetic side audio signal are then used alongside the side parameters to convert this mid and side stereo pair into a left and right stereo pair.
- a problem with the above mentioned previous solutions is that the decoding process fails to reproduce a convincing stereo pair when the original left and right audio signals are strongly decorrelated.
- Examples of strongly decorrelated audio includes audio signals representing rain sounds, the sound of applause or even some types of music.
- a method for reconstructing a stereo audio signal comprising the steps of receiving a bitstream including an encoded first mono audio signal and a set of reconstruction parameters and decoding the encoded first mono audio signal to provide a first mono audio signal.
- the method further comprises reconstructing a second mono audio signal using a neural network system trained to predict samples of the second mono audio signal given samples of the first mono audio signal and the reconstruction parameters, wherein the first mono audio signal and the reconstructed second mono audio signal forms a stereo audio signal pair.
- stereo audio signal pair it is meant two audio signals which together form a stereo format.
- the two audio signals of the stereo audio signal pair may have been recorded using two microphones in a stereo recoding configuration. It is also possible that the stereo audio signal pairs have been generated in a mixing process.
- the most common format of stereo audio signal pairs is left and right stereo audio signals, however many alternative formats of stereo audio signal pairs exist, such as mid and side stereo audio signals.
- the bitstream is an encoded representation of an original stereo audio signal.
- the invention is at least partially based on the understanding that a trained neural network model will be able to reconstruct a second mono audio signal with higher quality compared to a second mono audio signal which has been calculated analytically in a conventional decoder. Especially, when there is low correlation between the audio signals of the original stereo audio signal pair the encoded representation simply does carry enough information to reconstruct the second mono audio signal accurately which leads to poor performance for conventional decoders. With the trained neural network model, on the other hand, a second mono audio signal can be reconstructed with perceptually much higher quality, even when there is low or no correlation between the original stereo audio signal. Thus, the efficient and highly compressed bitstream can still be used even when the correlation between the original stereo audio signals is low.
- a method for reconstructing a stereo audio signal comprising the steps of receiving a bitstream including an encoded first mono audio signal and a set of reconstruction parameters and decoding the encoded first mono audio signal to provide a first mono audio signal.
- the method further comprises reconstructing a second mono audio and a third mono audio signal using a neural network system trained to predict samples of the second mono audio signal and samples of the third mono audio signal given samples of the first mono audio signal and the reconstruction parameters, wherein the reconstructed second mono audio signal and the reconstructed third mono audio signals forms a stereo audio signal pair.
- the neural network model configured as a single output network, trained to reconstruct a second mono audio signal which forms a stereo audio signal pair with the first mono audio signal
- the neural network model could be configured as a double output network, trained to reconstruct a second and third mono audio signal directly, wherein the second and third mono audio signal forms stereo audio signal pair.
- the second aspect of the invention features the same or equivalent benefits as the first aspect of the invention.
- the method of the second aspect of the invention enables the neural network model to introduce additional enhancements in the reconstruction of the stereo audio signal pair.
- the second mono audio signal may still form a stereo audio signal pair with the first mono audio signal but the third mono audio signal is an enhanced version of the first mono audio signal which has been predicted by the neural network model and which offers enhanced perceptual quality.
- the first mono audio signal may be of a stereo audio signal format (e.g. mid-side format) which is different from the desired output of a decoder (e.g. a left and right format).
- a decoder e.g. a left and right format
- the second and third mono audio signal may be of a desired stereo audio signal format (e.g. left and right format) different from the stereo audio format of the first mono audio signal.
- the neural network system is trained to operate on flattened audio signal samples and the method further comprises envelope flattening the first mono audio signal, to produce a flattened first mono audio signal, and providing the flattened first mono audio signal to the neural network system.
- the method further comprises inverse-flattening at least the reconstructed second mono audio signal.
- the neural network model is trained to operate on flattened samples of the first mono audio signal. By operating on flattened samples the performance of the neural network model may be enhanced while also allowing less complex neural network models to be used which are easier to train. In most audio content, the spectral energy content is higher for lower frequencies compared to higher frequencies, i.e.
- the audio content has a high dynamic range. If the samples of the audio signal are not flattened the neural network model will inherently prioritize accurate reconstruction of low frequencies over accurate reconstruction of high frequencies which could lead to noticeably distorted or lower quality reconstructed audio signals for some types of audio content. By flattening the samples, the spectral energy content will be more evenly disturbed across all frequencies meaning that the neural network model will put equal priority to accurate reconstruction of all frequencies which increases the quality of the reconstructed audio signals.
- a computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to the first or second aspect of the invention.
- a computer- readable storage medium storing the computer program according to the third aspect of the invention.
- a decoder comprising a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method steps of the first or second aspect of the invention.
- a method for training a neural network for stereo reconstruction comprising obtaining training data, the training data comprising a stereo audio signal pair and encoding the stereo audio signal pair into an encoded stereo audio signal, the encoded stereo audio signal comprising a first mono audio signal and reconstruction parameters.
- the method further comprises reconstructing a second mono audio signal using a neural network system trained to predict samples of the second mono audio signal given samples of the first mono audio signal and the reconstruction parameters and determining a difference measure between the reconstructed second mono audio signal and a ground truth second mono audio signal associated with the stereo audio signal pair of the training data.
- the method comprises modifying internal weights of the neural network model based on the determined difference.
- the method for training a neural network according to the sixth aspect is suitable for training a neural network model according to the first aspect of the invention.
- the neural network model is further configured to predict samples of a third mono audio signal given samples of the first mono audio signal and said reconstruction parameters, and the method further comprises reconstructing the third mono audio signal using the neural network model and determining the difference measure between the reconstructed third mono audio signal and a ground truth third mono audio signal associated with the stereo audio signal pair of the training data.
- This implementation of the training method is suitable for training the neural network model used in the second aspect of the invention.
- the third to sixth aspects of the invention features the same or equivalent benefits as the first and second aspects of the invention. Any functions described in relation to a method, may have corresponding features in a system and vice versa.
- Fig. la depicts a stereo encoder transmitting an encoded stereo bitstream to a stereo decoder according to some implementations.
- Fig. lb depicts a detailed view of a stereo encoder according to some implementations.
- Fig. 2a depicts a detailed view of a stereo decoder according to some implementations.
- Fig. 2b depicts a detailed view of another stereo decoder according to some implementations.
- Fig. 2c depicts a detailed view of a stereo decoder with a double output neural network model according to some implementations.
- Fig. 2d depicts a detailed view of another stereo decoder with a double output neural network model also performing stereo format conversion according to some implementations.
- Fig. 3a depicts a single output neural network model according to some implementations.
- Fig. 3b depicts a single output neural network model operating on flattened samples according to some implementations.
- Fig. 3c depicts a double output neural network model operating on flattened samples according to some implementations.
- Fig. 3d depicts a double output neural network model with stereo format conversion operating on flattened samples according to some implementations.
- Fig. 3e depicts a double output neural network model with stereo format conversion into alternative formats according to some implementations.
- Fig. 4a is flowchart describing a method of decoding a stereo audio signal with a single output neural network model according to some implementations.
- Fig. 4b is flowchart describing a method of decoding a stereo audio signal with a double output neural network model according to some implementations.
- Fig. 5a depicts a training setup for training a neural network model according to some implementations.
- Fig. 5b is a flowchart describing a method for training a neural network model according to some implementations.
- Fig. 6 illustrates a neural network model wherein a one of the neural network model and an LTI filter is used selectively, based on the content of the first mono audio signal according to some implementations.
- Systems and methods disclosed in the present application may be implemented as software, firmware, hardware or a combination thereof.
- the division of tasks does not necessarily correspond to the division into physical units; to the contrary, one physical component may have multiple functionalities, and one task may be carried out by several physical components in cooperation.
- the computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware.
- PC personal computer
- PDA personal digital assistant
- cellular telephone a smartphone
- smartphone a web appliance
- network router switch or bridge
- processors that accept computer-readable (also called machine-readable) code containing a set of instructions that when executed by one or more of the processors carry out at least one of the methods described herein.
- Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken are included.
- a typical processing system i.e. a computer hardware
- Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit.
- the processing system further may include a memory subsystem including a hard drive, SSD, RAM and/or ROM.
- a bus subsystem may be included for communicating between the components.
- the software may reside in the memory subsystem and/or within the processor during execution thereof by the computer system.
- the one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s).
- a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
- WAN Wide Area Network
- LAN Local Area Network
- the software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media).
- computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.
- Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.
- communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
- an encoder 10 and decoder 20 for encoding and decoding a stereo audio signal pair is presented.
- An original left audio signal L and original right audio signal R forming a stereo audio signal pair is provided to the encoder 10 which encodes the original stereo signal pair L, R to an encoded signal representation which is included in the bitstream B.
- Transforming the original stereo signal pair L, R to an encoded representation may be a lossy process wherein some information present in the original stereo audio signal pair has been omitted.
- the encoder 10 omits one of the original left and right original audio signals L, R and includes the other one of the left and right original audio signal L, R in the bitstream B.
- the encoder 10 further extracts reconstruction parameters indicating a relationship (e.g.
- the bitstream B is then provided to the decoder 20 which reconstructs one of the left audio signal L* and/or reconstructed right audio signal R* from the contents of the bitstream B.
- the decoder 20 reconstructs both a reconstructed left audio signal L* and a reconstructed right audio signal R*.
- the decoder 20 reconstructs the audio signal being the complement to the audio signal included in the bitstream B, i.e., only one of a reconstructed left audio signal L* and a reconstructed right audio signal R*.
- the encoder 10 transforms the original left and right original audio signal L, R into a mid-side stereo format and includes only the mid audio signal in the bitstream B alongside the reconstruction parameters indicating a property of the side audio signal.
- the mid audio signal is expected to capture the most essential information of a stereo signal pair (e.g., most stereo audio signals are center panned meaning that most of the spectral energy will be comprised in the mid audio signal), including the mid audio signal in the bitstream B instead of one of the left and right audio signal L, R enables more accurate reconstruction in the decoder 20.
- the decoder 20 then reconstructs a reconstructed side signal or a reconstructed left audio signal L* and a reconstructed right audio signal R* using the content in the bitstream B.
- An original left and right stereo signal pair L, R is received and provided to a stereo downmixing unit 12.
- the stereo downmixing unit 12 performs two tasks, it extracts a first mono audio signal a (e.g. in the form of a mid audio signal M), also referred to as a downmix audio signal, and it extracts reconstruction parameters P indicating a property of a relationship of the original left and right audio signal L, R. Extracting a mid audio signal M from a left and right audio signal L, R is for example achieved by the following equation:
- M giL + g r R (1) wherein gi and g r are channel weights, and setting gi and g r equal to ’A yields a conventional mono audio signal. Similarly, a corresponding side audio signal, S, can be created as
- the encoder 10 extracts a target mid audio signal wherein weights gi and g r change with time and/or frequency.
- the parameter 0 is referred to as a target panning parameter 0 and ranges from 0 to as it dictates the panning of a target audio source in the stereo pair L, R and the resulting dynamic mid and side audio signals are referred to as target mid and side audio signals.
- This target mid and side audio signals relate to the left and right stereo pair via the time varying panning dictated by the target panning parameter 0.
- the target panning parameter 0 is transmitted with the reconstruction parameters P and used by the neural network model and/or mixer of the decoder when reconstructing the stereo audio signals in the left and right format.
- the target panning parameter 0 varies over time and frequency to extract a target mid audio signal which captures a dominating audio source in each frequency band.
- the target panning parameter 0 could be set to an estimated panning in each frequency band.
- the target panning parameter 0 is calculated as arctan(
- is the spectral energy of the left and right audio signal L, R for a particular time segment and frequency band. For instance, if the left signal contains most spectral energy for a certain frequency band and time segment, 0 arctan(
- a target phase difference parameter ⁇ t> may be obtained for each time segment and frequency band of the left and right stereo pair.
- the target panning parameter 0 and the target phase difference parameter ⁇ t> are transmitted with the reconstruction parameters P and used by the neural network model and/or mixer of the decoder when reconstructing the stereo audio signals in the left and right format.
- the target mid audio signal obtained with equation 3 or equation 5 above will dynamically target a source which varies over time in frequency, panning and phase in the left and right stereo audio signal to enable the most prominent audio source to always be present in the target mid audio signal.
- the target mid audio signal the risk of not including the dominating audio source in stereo mix is reduced, even if the dominating source is varying in panning, phase or frequency over time.
- Extracting reconstruction parameters P may involve extracting at least one of the Inter-channel Intensity Difference (IID), the Inter-channel Cross-Correlation (ICC), the Interchannel Phase Difference (IPD) and the Inter-channel Time Difference (ITD) of the original left and right audio signals.
- IID Inter-channel Intensity Difference
- ICC Inter-channel Cross-Correlation
- IPD Interchannel Phase Difference
- ITD Inter-channel Time Difference
- Inter-channel Intensity Difference or IID indicates the intensity difference between the two signals in the original stereo signal pair L, R.
- Inter-channel Cross-Correlation or ICC indicates the cross-correlation or the coherence of the two signals in the original stereo signal pair L, R.
- the coherence is determined as the maximum of the cross-correlation as a function of time or phase.
- Inter-channel Phase Difference or IPD indicates the phase difference between the two signals in the original stereo signal pair L, R.
- An alternative to the IPD is the Interchannel Time Difference or ITD which indicates the time difference between the two signals of the original stereo audio signals L, R.
- the reconstruction parameters indicate the target panning parameter 0 and/or the target phase difference parameter ⁇ t> for each time segment and frequency band. This allows the neural network model and/or mixer of the decoder to reconstruct the original left and right audio signal from a target mid audio signal extracted using the target panning parameter 0 and/or the target phase difference parameter ⁇ t>.
- the first mono audio signal a is provided to a mono signal encoder 13 configured to encode the first mono audio signal a into an encoded first mono audio signal E(a).
- the encoding performed by the mono signal encoder may be lossless or lossy. Lossy encoding enables the first mono audio a to be compressed.
- the mono signal encoder 13 may perform downsampling or quantization of the first mono audio signal a.
- the bitstream encoder 11 is depicted as being separate from the encoder 10 it is also possible that the bitstream encoder 11 is a part of the encoder 10 which then accepts a pair of stereo audio signals L, R as an input and outputs an encoded bitstream B.
- the reconstruction parameters P are also compressed using e.g., quantization, performed by the quantizer 14.
- the (optionally encoded) first mono audio signal and the (optionally encoded) reconstruction parameters P are provided to a bitstream encoder 11 which encodes the information into a bitstream B.
- the bitstream B is then stored or transmitted (e.g. over a network) to a decoder 20.
- bitstream B is received by a bitstream decoder 21 which decodes the bitstream B to obtain the first mono audio signal a and the reconstruction parameters P contained in the bitstream B.
- the bitstream decoder 21 may be provided separately from the stere decoder 20 or integrated therewith.
- the bitstream decoder 21 decodes the bitstream encoding and any encoding encapsulating the first mono audio signal a and the reconstruction parameters P, and provides the first mono audio signal a and the reconstruction parameters P to the neural network model 24a of the stereo decoder 20.
- samples of the first mono audio signal a and reconstruction parameters P are provided as input parameters to the neural network model 24a trained to predict samples of a reconstructed second mono audio signal *.
- the first mono audio signal a and the reconstructed second mono audio signal 0* forms a stereo audio signal pair.
- the first mono audio signal a is a mid audio signal
- the reconstructed second mono audio signal 0* is a side audio signal wherein these two audio signals forms a mid and side stereo audio signal pair.
- the first mono audio signal a and the reconstructed second mono audio signal 0* are provided to a mixing unit 26 which mixes the first mono audio signal a and the reconstructed second mono audio signal 0 to form a reconstructed left and right stereo audio signal pair L*, R* if the first mono audio signal a and the reconstructed second mono audio signal 0* are not already in the left and right stereo audio signal format.
- the mixing unit 26 is provided with the target panning parameter 0 X and uses this parameter to reconstruct the left and right audio signals L*, R*.
- the neural network model 24a may comprise any type of neural network.
- the neural network is a Recurrent neural network (RNN) or a convolutional neural network (CNN).
- the neural network may comprise a plurality of neural network layers.
- the neural network model 24a may comprise a generative model.
- a generative model is a neural network that implements probability distribution (e.g., a conditional probability distribution), which models the probability distribution of the dataset on which the neural network has been trained.
- the reconstruction of the second, and optionally third, mono audio signal P* is achieved by random sampling according to the probability distribution implemented by the trained neural network.
- the architecture of the generative model may e.g., resemble that of the generative model described in detail in “HIGH FREQUENCY RECONSTRUCTION USING NEURAL NETWORK SYSTEM” filed as U.S. Provisional Application No. 63/331,056 on April 14, 2022, hereby incorporated by reference in its entirety.
- This generative model reconstructs a filter bank domain high-band signal using a neural network system trained to predict samples of a high-band audio signal in the filter bank domain given samples of the filter bank domain low-band signal and high frequency reconstruction (HFR) parameters, wherein the HFR parameters describe properties of the higher frequency bands.
- the neural network system comprises an upper neural network tier and a neural network bottom tier.
- the upper neural network tier In the upper neural network tier, previously generated filter-bank samples are received together with the decoded low-band samples and the high frequency reconstruction parameters.
- the bottom neural network tier is divided into a plurality of sequentially executed sub-layers, each sub-layer is configured to generate a set of channels of the reconstructed high frequency band.
- the generative model also reconstructs an enhanced low-band audio signal.
- the decoder 20 comprises some components which are identical with the corresponding component of the decoder of fig. 2a (e.g., the mixing unit 26).
- the decoder 20 of fig. 2b further comprises an envelope estimator 22 and a flattening unit 23. Additionally, the decoder 20 may comprise or be associated with a bitstream decoder as described in connection to the embodiment of fig. 2a.
- the envelope estimator 22 is configured to obtain the first mono audio signal a and estimate the spectral envelope of this audio signal.
- the spectral envelope is estimated for a number of frequency bands.
- the spectral envelope is a parametric representation of the spectral energy of each QMF-band in the first mono audio signal a.
- the spectral envelope may be represented with one, two, or at least three parameter values per frequency band.
- the audio signals are represented with a predetermined number (e.g. 32) of QMF bands which vary over time in segments wherein each band is associated with one reconstruction parameter (e.g. an IID-, ICC-, or IPD-value) that is updated for each time segment.
- one reconstruction parameter e.g. an IID-, ICC-, or IPD-value
- the spectral envelope is provided to a flattening unit 23.
- the flattening unit 23 is configured to flatten the first mono audio signal a so as to provide flattened samples OLF of the first mono audio signal a to the neural network model 24b.
- the neural network model 24b is trained to predict flattened samples of the second mono audio signal P*F provided flattened samples OLF of the first mono audio signal a.
- the neural network model 24a of fig. 2a is trained to operate on original (non-flattened) samples of the first mono audio signal a
- the neural network model 24b is trained to operate on flattened samples OLF.
- the neural network model 24b also receives the reconstruction parameters P as an input.
- the neural network model 24b also receives the spectral envelope as input, wherein the neural network model 24b is trained to predict the reconstructed second mono audio signal P* based on three types of input data: the flattened samples of the first mono audio signal a, the reconstruction parameters P, and the spectral envelope.
- the neural network model 24b By allowing the neural network model 24b to operate in the flattened domain, the neural network model 24b can be made less complex (e.g. fewer layers and/or fewer parameters) and/or the training of the neural network model 24b is more efficient.
- the reconstructed flattened samples P*F are provided to an inverse-flattening unit 25 which performs the inverse operation of the flattening unit 23 to obtain non-flattened samples P*F.
- the spectral envelope is provided to the inverse-flattening unit 25 alongside the reconstructed flattened samples P*F.
- the inverse flattening unit 25 accepts as an input the flattened reconstructed second mono audio signal P*F (e.g. a flattened reconstructed side audio signal), and outputs inverse-flattened audio signal samples P* (i.e. the reconstructed audio signal with no flattening).
- the inverse flattened reconstructed second mono audio signal P* output by the inverse-flattening unit 25 is provided to the mixing unit 26 which mixes the inverse flattened reconstructed second mono audio signal P* with the first mono audio signal a to obtain a reconstructed left and right stereo audio signal pair L*, R*.
- a decoder 20 is schematically illustrated with a neural network model 24c trained to predict a (flattened) second reconstructed mono audio signal P*(F) and a (flattened) third reconstructed mono audio signal y*(F) given the (flattened) first mono audio signal a(F) and reconstruction parameters P.
- the reconstructed third mono audio signal y* is an enhanced version of the first mono audio signal a.
- the first mono audio signal a is a mid audio signal
- the reconstructed second mono audio signal P* is a side audio signal
- the reconstructed third mono audio signal y* is an enhanced mid audio signal.
- the first mono audio signal a may be compressed, quantized or processed with any form of lossy audio encoding technique.
- the neural network model 24c can be trained to produce an enhanced version of the first mono audio signal a in addition to the reconstructed second mono audio signal.
- the mixing unit 26 mixes the reconstructed third mono audio signal with the reconstructed second audio signal to produce a reconstructed left and right stereo audio signal L*, R* with enhanced quality.
- Fig. 2d shows yet another embodiment of the decoder 20 wherein the neural network model 24d has been trained to output samples of a (flattened) left and right stereo audio signal pair L*F, R*F directly provided samples of the first mono audio signal a and the reconstruction parameters P. This allows e.g. the mixing unit 26 to be omitted completely from the decoder 10.
- the first mono audio signal a may be any one of a mid audio signal, a side audio signal, a left audio signal and a right audio signal.
- the neural network model 24a, 24b, 24c, 24d receives a first mono audio signal a being a first part of a first stereo format and reconstruction parameters P describing a property of the second part of the first stereo format.
- the neural network model 24a, 24b, 24c, 24d is trained to output either (a) a reconstructed first format audio signal being the second part of the first stereo format or (b) two reconstructed second format audio signals being a first and second part of a second stereo format wherein the second stereo format is different from the first stereo format.
- the neural network model 24a, 24b, 24c, 24d obtains a left audio signal and reconstruction parameters associated with the right audio signal and outputs the right audio signal or the neural network model 24a, 24b, 24c, 24d obtains a mid audio signal and reconstruction parameters associated with the side audio signal and outputs a left and right stereo audio signal pair.
- the first mono audio signal a is a mid audio signal and the parameters P describe a property of the corresponding side audio signal, wherein the neural network model 24 d directly predicts a reconstructed (flattened) left and right stereo audio signal pair L*F, R*F.
- the decoder 20 of fig. 2d operates on flattened samples it is understood that the envelope estimator 22, flattening unit 23 and inverse-flattening unit 25 may be omitted to allow the neural network module 24d to operate on un-flattened samples.
- the encoder 10 receives an original left and right audio signal L, R.
- the encoder 10 instead receives original audio signals of a mid-side format or any other type of stereo audio signal format. Irrespective of the type of stereo format is received by the encoder 10, the encoder 10 encodes a bitstream B carrying a representation of a mono audio signal and reconstruction parameters describing at least one property of the relationship between the original audio signals.
- the encoder 10 receives a left and right audio signal L, R and includes in the bitstream B one of the left and right audio signals L, R and reconstruction parameters P describing a property of the other one of the left and right audio signals L, R.
- the encoder 10 and decoder 20 described in the above may operate on audio signals in the time domain and/or in the frequency domain (e.g. in the QMF-domain).
- the encoder 10 converts the first mono audio signal and reconstruction parameters into a time-frequency domain format (such as QMF).
- the neural network model 24a, 24b, 24c, 24d may be trained to predict the second (and optionally the third) mono audio signal based on time domain samples of the first mono audio and reconstruction parameters describing time domain properties.
- the neural network model 24a, 24b, 24c, 24d may be trained to predict the second (and optionally the third) mono audio signal based on frequency domain samples of the first mono audio and reconstruction parameters describing frequency domain properties.
- legacy decoders without a neural network model 24a, 24b, 24c, 24d it is common for the other components in the decoder (e.g. the upmixing unit) to operate in a filter-bank domain with a predetermined number of frequency bands.
- the neural network model 24a, 24b, 24c, 24d may then preferably be trained to operate in the same filter-bank domain to facilitate easy implementation in legacy decoders.
- Fig. 3a, 3b, 3c, 3d and 3e depicts some implementations of the different neural network models described in the above and specific examples of the first, second and third mono audio signals.
- the neural network model 24a of fig. 3a is trained to obtain samples of a mid audio signal M as well as reconstruction parameters Ps indicating a property of the associated side audio signal S and output a reconstructed side audio signal S*. That is, the neural network model 24a is a single output network and the first and second mono audio signal forms a stereo audio signal pair of a mid-side format.
- the neural network model 24b of fig. 3b is equal to that of the neural network model 24a from fig. 3 a besides the fact that the neural network model 24b of fig. 3b is trained to operate on flattened samples, and optionally reconstruction parameters Ps associated with a flattened audio signal.
- the spectral envelope (determined in connection to flattening the samples) is also provided to the neural network model 24b as additional input data. Experiments have shown that when the samples of the first mono audio signal (e.g. the mid audio signal M) are flattened, providing the envelope to the neural network model 24b facilitates performance of the neural network model 24b.
- the neural network model 24c of fig. 3c also operates on flattened audio signal samples, however it is envisaged that the same neural network model 24c may also be trained to operate on non-flattened samples.
- the neural network model 24c is a double output network trained to obtain a flattened mid audio signal MF (the first mono audio signal) and reconstruction parameters Ps associated with the corresponding side audio signal and outputs two flattened audio signals: the reconstructed side audio signal S*F (second mono audio signal) and an enhanced reconstructed mid audio signal M*F (third mono audio signal).
- These audio signals S*F, M*F forms their own stereo audio signal pair and may be outputted (after inverse-flattening) as a stereo audio signal or mixed to form a different stereo audio signal pair (e.g. a left and right stereo audio signal pair).
- the neural network model 24c is provided with the spectral envelope as additional input data in some implementations.
- the neural network model 24d of fig. 3d also operates on flattened audio signal samples, however it is envisaged that the same neural network model 24d may be trained to operate on non-flattened samples. Similar to the neural network model 24c, the neural network model 24d of fig. 3d outputs two audio signals, however, the outputted reconstructed audio signals L*F, R*F of the neural network model 24d are of a different stereo format than that of the first mono audio signal MF. AS seen in fig 3d, the first mono audio signal is a mid audio signal of a mid-side stereo format whereas the outputted reconstructed stereo audio signals are of a left and right stereo format.
- the neural network model 24d is provided with the spectral envelope as additional input data in some implementations.
- the first mono audio signal provided to a neural network model 24e is a left or right audio signal L, R with the reconstruction parameters PL/R indicating a property of the other one of the left or right audio signal L, R.
- the output of the neural network model 24e may be a single signal or double signals.
- the single signal being the reconstruction of the one of the left or right audio signal which does not constitute the first mono audio signal.
- the double output audio signals may be a reconstructed/enhanced left and right audio signal L, R (i.e.
- the stereo format is maintained by the neural network model) or a reconstructed stereo audio signal pair of a different stereo format such as a mid-side format M, S.
- the neural network model 24e is provided with the spectral envelope of the flattened samples representing the left or right audio signal L, R which may facilitate performance.
- a neural network 24e and stereo decoder which operates in an analogous manner to the neural networks and stereo decoder embodiments described in the above with the input first mono audio signal being anyone of a side audio signal S, a left audio signal L, and a right audio signal R.
- a bitstream B is received by the decoder 20, the bitstream comprising an encoded first mono audio signal a and reconstruction parameters P.
- the decoder 20 decodes the encoded first mono audio signal a and the encoded reconstruction parameters P provides the decoded first mono audio signal a to the neural network model 24a alongside the reconstruction parameters P.
- a second mono audio signal P* is reconstructed by the neural network model 24a and the first and second mono audio signal forms a stereo audio signal pair.
- the first and second mono audio signal, a, P is a mid and side stereo audio signal or a left and right stereo audio signal.
- the first and second mono audio signal are provided to, and mixed, by a mixing unit 26 at step S4a so as to convert the first and second mono audio signal into a different alternative stereo format.
- Embodiments are also envisaged in which two audio signals are reconstructed as will now be described with reference to fig. 2c and fig. 4b wherein the method also comprises receiving a bitstream at step SI and decoding the first mono audio signal a at step S2 as described in the above.
- step S3b is also carried out in which the neural network model 24b reconstructs the third mono audio signal y* in addition to the second mono audio signal P* reconstructed at step S3a.
- Steps S3a and S3b may occur in sequence or simultaneously.
- the second and third reconstructed mono audio signals P*, y* forms a stereo audio signal pair of the same format as the format to which the first mono audio signal belongs or a different stereo format which may be outputted by the decoder 20.
- training data is obtained, for example from a database 30 comprising training data.
- the training data comprises at least one example of a stereo audio signal pair (e.g. a left right stereo audio signal pair L, R) and is provided to a stereo encoder 10.
- the stereo encoder 10 encodes the stereo audio signal pair into a bitstream containing a first mono audio signal (such as a mid audio signal M) and reconstruction parameters P indicating at least one property of a stereo audio signal associated with the first mono audio signal.
- the encoding process implemented by the encoder 10 may be a lossy process, meaning that the information contained in the bitstream may not comprise sufficient information to perfectly reconstruct the stereo audio signal pair L, R input to the encoder 10.
- the first mono audio signal and the reconstruction parameters are provided to the decoder 20 which outputs a reconstructed stereo audio signal pair L*, R*.
- the decoder 20 may be any one of the decoders described in connection to fig. 2a, 2b, 2c, 2d in the above and comprises a neural network model 24 which is the subject of the training.
- the reconstructed stereo audio signal pair L*, R* is provided to a loss function unit 40 which compares the original stereo audio signal pair L, R to the reconstructed and reconstructed stereo audio signal pair L*, R* and determines a difference measure (also called a loss) between the original and the reconstructed stereo audio signal pairs.
- a difference measure also called a loss
- NLL Negative Log Likelihood
- a discriminator a so called learnable loss function
- the decoder 20 acts a generator and the loss function unit 40 comprises an additional neural network acting as a discriminator.
- the discriminator and generator are then trained used traditional generator/discriminator training.
- the internal weights and parameters of the neural network model 24 is updated at step T5 so as to reduce the difference measure.
- the audio signals of the training data and the audio signals output by the decoder 20 are in the same left-right L, R stereo format.
- the training data is in a different stereo format (e.g. mid-side format) and/or there is a mismatch between the stereo format outputted by the decoder 20 compared to the format training data in the database 30.
- the loss function unit 40 is configured to convert the stereo format output by the decoder 20 and/or the stereo format of the ground truth training data 30 to enable calculation of the loss.
- Fig. 6 shows yet another embodiment of a decoder 20 implementing a neural network module 24.
- the neural network module 24 could be anyone of the neural network modules 24a, 24b, 24c, 24d, 24e described in the above and may be variants thereof operating with flattened or non-flattened audio signal samples.
- the decoder 20 in fig. 6 comprises a content analyzer 27 configured to determine a correlation level for the audio content of the bitstream (embodied by the first mono audio signal and the reconstruction parameters).
- the correlation level indicates a level of correlation between the first mono audio signal a and a stereo audio signal associated with the first mono audio signal a. For example, if the bitstream carrying the first mono audio signal a and the reconstruction parameters P is obtained by encoding a left and right audio signal pair the correlation level will indicate the level of correlation between the left and right audio signal.
- the correlation level is determined by the content analyzer based on the reconstruction parameters P.
- the reconstruction parameters may comprise a parameter P which indicates the level of correlation directly.
- the content analyzer 27 comprises a neural network trained to predict the correlation level based on samples of the first mono audio signal a and optionally also the reconstruction parameters P. For instance, the content analyzer 27 may be trained to determine the type of audio content (e.g. speech, music, the sound of applause or the sound of rain) present in the first mono audio signal a wherein some content types (e.g. the sound of applause or rain) is associated with a low level of correlation.
- the type of audio content e.g. speech, music, the sound of applause or the sound of rain
- some content types e.g. the sound of applause or rain
- the correlation level is provided to a selection module 28 which selects whether to provide the first mono audio signal M and the reconstruction parameters to the neural network module or to a predetermined Linear Time Invariant (LTI) filter 29.
- the LTI filter 29 may be realized with a delay line which imposes a (optionally frequency varying) delay to the audio signals in the time domain.
- the LTI filter 29 comprises infinite impulse response (IIR) and/or finite impulse response (FIR) filters simulating reverberation.
- the delay in the time domain may be varying with frequency, e.g. with higher frequency bands being subjected to smaller delays.
- the selection module selects 28 the neural network module 24 and if the correlation level is above the predetermined threshold the selection module 28 selects the LTI filter 29.
- the neural network model 24 is especially well suited for reconstructing a second (and optionally third) mono audio signal when there is low correlation between the channels of the stereo audio signal pair which is described by the first mono audio signal a and the reconstruction parameters P.
- the neural network model 24 is used when is most needed and the simpler LTI filter 29 is used for reconstruction of correlated stereo audio signal pairs.
- an LTI filter 29 it is meant a filter comprising a set of filter coefficients determined analytically to perform a specific task, in this case the task of imposing a time domain delay and/or reverberation to the audio signal.
- the delay filter comprises a delay unit configured to introduce a predetermined (possibly frequency varying) delay.
- a predetermined (possibly frequency varying) delay e.g. 10 milliseconds or 2 milliseconds
- the filter coefficients of the LTI filter 29 can be calculated analytically. For example, if an LTI filter 29 introduces a delay which decreases linearly with frequency, the LTI filter would have the impulse response of a chirp signal.
- An example of an LTI filter 29 is e.g.
- the neural network model 24 comprises at least one of a plurality of (learnable) neural network layers, non-linear activation layers (with e.g. Rectified Linear Units, ReLU), at least one Long Short-Term Memory (LSTM) layer, at least one recurrent layer (such as a layer comprising Gated Recurrent Units, GRUs) meaning that the neural network model 24 is clearly distinguished from the LTI filter 29 which is both linear and time-invariant.
- non-linear activation layers with e.g. Rectified Linear Units, ReLU
- LSTM Long Short-Term Memory
- GRUs Gated Recurrent Units
- Each of these exemplary decoders may also be configured to output stereo audio signal pairs of different formats than the left and right format, such as a mid-side or target mid-side format.
- decoders with and without spectral flattening and the associated components are also envisaged.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Stereophonic System (AREA)
Abstract
Description
Claims
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263433737P | 2022-12-19 | 2022-12-19 | |
| EP23157900 | 2023-02-22 | ||
| PCT/EP2023/086156 WO2024132968A1 (en) | 2022-12-19 | 2023-12-15 | Method and decoder for stereo decoding with a neural network model |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4639533A1 true EP4639533A1 (en) | 2025-10-29 |
Family
ID=89308574
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23828198.4A Pending EP4639533A1 (en) | 2022-12-19 | 2023-12-15 | Method and decoder for stereo decoding with a neural network model |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4639533A1 (en) |
| JP (1) | JP2025541140A (en) |
| CN (1) | CN120418863A (en) |
| WO (1) | WO2024132968A1 (en) |
-
2023
- 2023-12-15 WO PCT/EP2023/086156 patent/WO2024132968A1/en not_active Ceased
- 2023-12-15 EP EP23828198.4A patent/EP4639533A1/en active Pending
- 2023-12-15 CN CN202380087021.9A patent/CN120418863A/en active Pending
- 2023-12-15 JP JP2025532921A patent/JP2025541140A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN120418863A (en) | 2025-08-01 |
| WO2024132968A1 (en) | 2024-06-27 |
| JP2025541140A (en) | 2025-12-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12205600B2 (en) | Methods, apparatus and systems for encoding and decoding of multi-channel Ambisonics audio data | |
| US8817991B2 (en) | Advanced encoding of multi-channel digital audio signals | |
| EP2002424B1 (en) | Device and method for scalable encoding of a multichannel audio signal based on a principal component analysis | |
| US7573912B2 (en) | Near-transparent or transparent multi-channel encoder/decoder scheme | |
| EP1989920B1 (en) | Audio encoding and decoding | |
| US9516446B2 (en) | Scalable downmix design for object-based surround codec with cluster analysis by synthesis | |
| CA3071208C (en) | Apparatus for encoding or decoding an encoded multichannel signal using a filling signal generated by a broad band filter | |
| US11501785B2 (en) | Method and apparatus for adaptive control of decorrelation filters | |
| CN117136406A (en) | Combine spatial audio streams | |
| EP2489036B1 (en) | Method, apparatus and computer program for processing multi-channel audio signals | |
| EP4639533A1 (en) | Method and decoder for stereo decoding with a neural network model | |
| CN120266204A (en) | Parameter Spatial Audio Coding | |
| JP2026508703A (en) | Joint Stereo Coding in the Complex-Valued Filterbank Domain | |
| WO2017148526A1 (en) | Audio signal encoder, audio signal decoder, method for encoding and method for decoding |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250709 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| P01 | Opt-out of the competence of the unified patent court (upc) registered |
Free format text: CASE NUMBER: UPC_APP_0012839_4639533/2025 Effective date: 20251111 |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |