EP4042418A1 - Détermination de corrections à appliquer a un signal audio multicanal, codage et décodage associés - Google Patents
Détermination de corrections à appliquer a un signal audio multicanal, codage et décodage associésInfo
- Publication number
- EP4042418A1 EP4042418A1 EP20792467.1A EP20792467A EP4042418A1 EP 4042418 A1 EP4042418 A1 EP 4042418A1 EP 20792467 A EP20792467 A EP 20792467A EP 4042418 A1 EP4042418 A1 EP 4042418A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signal
- decoded
- multichannel signal
- decoding
- corrections
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/26—Pre-filtering or post-filtering
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S3/00—Systems employing more than two channels, e.g. quadraphonic
- H04S3/008—Systems employing more than two channels, e.g. quadraphonic in which the audio signals are in digital form, i.e. employing more than two discrete digital channels
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S3/00—Systems employing more than two channels, e.g. quadraphonic
- H04S3/02—Systems employing more than two channels, e.g. quadraphonic of the matrix type, i.e. in which input signals are combined algebraically, e.g. after having been phase shifted with respect to each other
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/01—Multi-channel, i.e. more than two input channels, sound reproduction with two speakers wherein the multi-channel information is substantially preserved
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/03—Aspects of down-mixing multi-channel audio to configurations with lower numbers of playback channels, e.g. 7.1 -> 5.1
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/11—Positioning of individual sound objects, e.g. moving airplane, within a sound field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/13—Aspects of volume control, not necessarily automatic, in stereophonic sound systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/07—Synergistic effects of band splitting and sub-band processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2420/00—Techniques used stereophonic systems covered by H04S but not provided for in its groups
- H04S2420/11—Application of ambisonics in stereophonic audio systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S5/00—Pseudo-stereo systems, e.g. in which additional channel signals are derived from monophonic signals by means of phase shifting, time delay or reverberation
Definitions
- the present invention relates to the encoding / decoding of spatialized sound data, in particular in a surround sound context (hereinafter also referred to as “ambisonic”).
- the encoders / decoders which are currently used in mobile telephony are mono (a single signal channel for reproduction on a single loudspeaker).
- coded which are currently used in mobile telephony are mono (a single signal channel for reproduction on a single loudspeaker).
- the 3GPP EVS (for “Enhanced Voice Services”) code makes it possible to offer “Super-HD” quality (also called “High Definition Rus” or HD + voice) with a super-widened audio band (SWB for “super- wideband ”in English) for signals sampled at 32 or 48 kHz or full band (FB for“ Fullband ”) for signals sampled at 48 kHz; the audio bandwidth is 14.4 to 16 kHz in SWB mode (9.6 to 128 kbit / s) and 20 kHz in FB mode (16.4 to 128 kbit / s).
- the next quality evolution in conversational services offered by operators should be immersive services, using terminals such as smartphones equipped with several microphones or spatialized audio conferencing or videoconferencing equipment such as tele-presence or video. 360 °, or even “live” audio content sharing equipment, with spatialized 3D sound rendering that is far more immersive than a simple 2D stereo reproduction.
- terminals such as smartphones equipped with several microphones or spatialized audio conferencing or videoconferencing equipment such as tele-presence or video. 360 °, or even “live” audio content sharing equipment, with spatialized 3D sound rendering that is far more immersive than a simple 2D stereo reproduction.
- advanced audio equipment accessories such as a 3D microphone, voice assistants with acoustic antennas, virtual reality headsets, etc.
- the capture and rendering of spatialized sound scenes are now common enough to offer an immersive communication experience.
- IVAS Intelligent Voice And Audio Services
- Ambisonics is a recording method (“encoding” in the acoustic sense) of spatialized sound and a reproduction system (“decoding” in the acoustic sense).
- An ambisonic microphone (at order 1) comprises at least four capsules (typically of the cardioid or sub-cardioid type) arranged on a spherical grid, for example the vertices of a regular tetrahedron.
- the audio channels associated with these capsules are called “A-format”. This format is converted into a “B-format”, in which the sound field is broken down into four components (spherical harmonics) denoted W, X, Y, Z, which correspond to four coincident virtual microphones.
- the W component corresponds to an omnidirectional capture of the sound field while the X, Y and Z components, which are more directive, can be compared to microphones with pressure gradients oriented along the three orthogonal axes of space.
- An ambisonic system is a flexible system in the sense that recording and playback are separate and decoupled. It allows decoding (in the acoustic sense) on any speaker configuration (for example, binaural, 5.1-type “surround” sound or 7.1.4-type periphery (with elevation)).
- the ambisonics approach can be generalized to more than four channels in B-format and this generalized representation is commonly referred to as “HOA” (for “Higher-Order Ambisonics”).
- FOA First-Order Ambisonics
- There is also a so-called "planar" variant of ambisonics (W, X, Y) which decomposes the sound defined in a plane which is generally the horizontal plane. In this case, the number of components is K 2M + 1 channels.
- ambisonics 1st order ambisonics (4 channels: W, X, Y, Z), 1st order planar ambisonics (3 channels: W, X, Y), higher order ambisonics are all referred to here -after “ambisonics” indiscriminately to facilitate reading, the treatments presented being applicable independently of the planar type or not and of the number of ambisonic components.
- an “ambisonic signal” will be called a signal in B-format with a predetermined order with a certain number of ambisonic components.
- This also includes hybrid cases, where for example in order 2 there are only 8 channels (instead of 9) - more predsly, in order 2, we find the 4 channels of order 1 (W , X, Y, Z) to which we normally add 5 channels (usually denoted R, S, T, U, V), and we can for example ignore one of the higher order channels (for example
- the signals to be processed by the encoder / decoder are in the form of successions of blocks of sound samples called "frames" or “sub-frames” below.
- the notations A T and A H indicate respectively the transposition and the Hermitian transposition (transposed and conjugated) of A.
- a multidimensional discrete-time signal, b (i), defined over a time interval i 0, .., L-1 of length L and with K dimensions is represented by a matrix of size
- a 3D point of Cartesian coordinates (x, y, z) can be converted to spherical coordinates (r, ⁇ , ⁇ ), where r is the distance to the origin, ⁇ is the azimuth and ⁇ the elevation.
- r is the distance to the origin
- ⁇ is the azimuth
- ⁇ the elevation.
- the first component of an ambisonic signal generally corresponds to the omnidirectional component W .
- the simplest approach to encoding an ambisonic signal is to use a mono encoder and apply it in parallel to all channels, possibly with different bit allocation depending on the channel. This approach is referred to herein as “multi-mono”.
- the multi-mono approach can be extended to multi-stereo coding (where pairs of channels are coded separately by a stereo coded) or more generally to the use of several parallel instances of the same core coded.
- the input signal is divided into channels (a mono channel or several channels) by the block 100. These channels are coded separately by the blocks 120 to 122 according to a distribution and of a predetermined binary allocation. Their bit stream is multiplexed (block 130) and after transmission and / or storage, it is demultiplexed (block 140) to apply a decoding to reconstruct the decoded channels (blocks 150 to 152) which are recombined (block 160).
- the associated quality varies depending on the core encoding and decoding used (blocks 120 to 122 and 150 to 152), and it is generally only satisfactory at very high speed.
- the multi-mono coding approach does not take into account the correlation between channels, it produces spatial distortions with the addition of different artefacts such as the appearance of phantom sound sources, diffuse noise or movements of the trajectories of sound sources. .
- the encoding of an ambisonic signal according to this approach generates degradation of spatialization.
- An alternative approach to coding all channels separately is given, for a stereo or multichannel signal, by parametric coding.
- the input multichannel signal is reduced to a smaller number of channels, after a processing called “downmix”, these channels are encoded and transmitted and additional spatialization information is also encoded.
- Parametric decoding consists in increasing the number of channels after decoding of the transmitted channels, by using a processing called “upmix” (typically implemented by decorrelation) and a spatial synthesis as a function of the additional decoded spatialization information.
- upmix typically implemented by decorrelation
- 3GPP e-AAC + codec An example of stereo parametric coding is given by the 3GPP e-AAC + codec. It should be noted that the downmix operation also generates degradation of spatialization; in this case, the spatial image is changed.
- the invention improves the state of the art.
- the determined set of corrections, to be applied to the decoded multichannel signal makes it possible to limit the spatial degradations due to the coding and possibly to channel reduction / increase operations.
- the implementation of the correction thus makes it possible to find a spatial image of the decoded multichannel signal closest to the spatial image of the original multichannel signal.
- the determination of the set of corrections is performed in the full band time domain (a frequency band). In variants, it is performed in the time domain by frequency sub-band. This makes it possible to adapt the corrections according to the frequency bands. In other variants, it is performed in a real or complex transformed domain (typically frequency) of the short-term discrete Fourier transform (STFT), modified discrete cosine transform (MDCT), or other type.
- STFT short-term discrete Fourier transform
- MDCT modified discrete cosine transform
- the invention also relates to a method for decoding a multichannel sound signal, comprising the following steps:
- the decoder is able to determine the corrections to be made to the decoded multichannel signal, from information representative of the spatial image of the original multichannel signal, received from the encoder.
- the information received from the encoder is thus limited. It is the decoder that takes care of both determining and applying corrections.
- the invention also relates to a method for encoding a multichannel sound signal, comprising the following steps:
- the encoder which determines the set of corrections to be made to the decoded multichannel signal and which transmits it to the decoder. It is therefore the coder who initiates this determination of corrections.
- the information representative of a spatial image is a covariance matrix and the determination of the set of corrections comprises in in addition to the following steps:
- the determination of the set of corrections of the decoding method comprises furthermore the following steps: - Obtaining a weighting matrix comprising weighting vectors associated with a set of virtual loudspeakers;
- the decoding method or the encoding method comprises a step of limiting the values of gains obtained according to at least one threshold.
- Get set of gains constitutes the set of corrections and may for example be in the form of a correction matrix comprising all of the gains thus determined.
- the information representative of a spatial image is a covariance matrix and the determination of the set of corrections comprises a step of determining a matrix of transformation by matrix decomposition of the two covariance matrices, the transformation matrix constituting the set of corrections.
- This embodiment has the advantage of making the corrections directly in the ambisonic domain in the case of an ambisonic multichannel signal. The steps of transforming the signals reproduced on loudspeakers into the ambisonic domain are thus avoided.
- the correction of the multi-channel signal decoded by the determined set of corrections is performed by the application of the set of corrections to the decoded multichannel signal, that is to say directly in the ambisonic domain in the case of an ambisonic signal.
- the correction of the multichannel signal decoded by the determined set of corrections is performed according to the following steps:
- the steps of decoding, applying gains and encoding / summing above are grouped together in a direct correction operation by a correction matrix.
- This correction matrix can be applied directly to the decoded multichannel signal, which has the advantage as described above of making the corrections directly in the ambisonic domain.
- the decoding method comprises the following steps:
- the encoder which determines the corrections to be made to the decoded multichannel signal, directly in the ambisonic domain and it is the decoder which implements the application of these corrections to the decoded multichannel signal, directly in the ambisonic domain.
- the set of corrections can in this case be a transformation matrix or else a correction matrix comprising a set of gains.
- the decoding method comprises the following steps:
- the encoder which determines the corrections to be made to the signals resulting from the acoustic decoding on a set of virtual loudspeakers and it is the decoder which implements the application of these corrections to the signals.
- signals resulting from acoustic decoding then which transforms these signals to return to the ambisonic domain in the case of an ambisonic multichannel signal.
- the steps of decoding, applying gains and encoding / summing above are grouped together in a direct correction operation by a correction matrix.
- the correction is then carried out directly by applying a correction matrix to the decoded multichannel signal, for example the ambisonic signal. As described previously, this has the advantage of making corrections directly in the Ambisonic domain.
- the invention also relates to a decoding device comprising a processing circuit for implementing the decoding methods as described above.
- the invention also relates to a decoding device comprising a processing circuit for implementing the coding methods as described above.
- the invention relates to a computer program comprising instructions for implementing decoding methods or encoding methods as described above, when they are executed by a processor.
- the invention relates to a storage medium, readable by a processor, storing a computer program comprising instructions for carrying out the decoding methods or the encoding methods described above.
- Figure 1 illustrates multi-mono coding according to the state of the art and as described above;
- Figure 2 illustrates in flowchart form the steps of a method for determining a set of corrections according to one embodiment of the invention
- FIG. 3 illustrates a first embodiment of an encoder and a decoder, an encoding method and a decoding method according to the invention
- FIG. 4 illustrates a first detailed embodiment of the block for determining the set of corrections
- FIG. 5 illustrates a second detailed embodiment of the block for determining the set of corrections
- FIG 6 Figure 6 illustrates a second embodiment of an encoder and a decoder, a coding method and a decoding method according to the invention
- Figure 7 illustrates examples of structural embodiments of a coder and of a decoder according to one embodiment of the invention.
- the method described below is based on the correction of spatial degradations, in particular to ensure that the spatial image of the decoded signal is as close as possible to the original signal.
- the invention is not based on a perceptual interpretation of spatial image information because the ambisonic domain is not directly “listenable”.
- FIG. 2 represents the main steps implemented to determine a set of corrections to be applied to the encoded and then decoded multichannel signal.
- the original multichannel signal B of dimension KxL (ie K components of L time or frequency samples) is input to the determination method.
- step S1 information representative of a spatial image of the original multichannel signal is extracted.
- the invention can also be applied for other types of multichannel signal such as a B-format signal with modifications, such as, for example, the removal of certain components (e.g. removal of the R component at order 2 in order to keep only 8 channels) or the matrixing of the B-format to pass into an equivalent domain (called “Equivalent Spatial Domain”) as described in the specification 3GPP TS 26.260 - another example of matrixing is given by the “channel mapping 3” of the IETF Opus coded and in the 3GPP TS 26.918 specification (dause 6.1.6.3).
- spatial image The distribution of sound energy from the ambisonic soundstage to different directions in space is referred to here as a "spatial image"; in variants, this spatial image describing the sound scene generally corresponds to positive quantities evaluated at different predetermined directions in space, for example in the form of a pseudo-spectrum of the MUSIC (Multiple Signal Classification) type sampled at these directions or a histogram of directions of arrival (where the directions of arrival are counted according to the discretization given by the predetermined directions); these positive quantities can be interpreted as energies and are seen as such hereafter to simplify the description of the invention.
- MUSIC Multiple Signal Classification
- a spatial image associated with an ambisonic sound scene therefore represents the sound energy (or more generally a positive quantity) relative as a function of different directions in space.
- a piece of information representative of a spatial image can be for example a covariance matrix calculated between the channels of the multichannel signal or else an information of energy associated with directions of origin of the sound (associated with directions of height. - virtual speakers distributed over a unity sphere).
- the set of corrections to be applied to a multichannel signal is a piece of information which can be defined by a set of gains associated with directions of origin of the sound which can be in the form of a matrix of corrections comprising this set of gains or a transformation matrix.
- a covariance matrix of a multichannel signal B is for example obtained in step S1. As described later with reference to FIGS. 3 and 6, this matrix is for example calculated as follows:
- the covariance can be estimated recursively (sample by sample) in the form:
- energy information is obtained in different directions (associated with directions of virtual loudspeakers distributed over a unit sphere).
- SRP for “Steered-Response Power”
- MUSIC pseudo-spectrum, arrival direction histogram can be used.
- multi-stereo coding where the channels b k are coded in separate pairs is also possible.
- a typical example for a 5.1 input signal is to use two separate stereo encodings of L / R and Ls / Rs with mono encodings of C and LFE (low frequencies only); for the ambisonic case, the multi-stereo coding can be applied to the ambisonic components (B-format) or to an equivalent multichannel signal obtained after matrixing of the B-format channels - for example at order 1 the channels W, X, Y, Z can be converted to four transformed channels and two pairs of channels are encoded separately and converted back to B-format on decoding.
- An example is given in recent versions of the Opus code (“channel mapping 3”) and in specification 3GPP TR 26.918 (dause 6.1.6.3).
- step S2 it is also possible to use in step S2 a joint multichannel coding, such as for example the MPEG-H 3D Audio coded for the ambisonic format (scene-based); in this case, the codec performs coding of the input channels jointly.
- this joint coding is broken down for an ambisonic signal into several steps such as the extraction and coding of predominant mono sources, the extraction of an ambience (typically reduced to an ambisonic signal of order 1 ), the coding of all the extracted channels (called “transport channels”) and of metadata describing the acoustic beamforming vectors for the extraction of predominant channels.
- Joint multichannel encoding makes it possible to exploit the relationships between all channels to, for example, extract predominant audio sources and ambience or perform global bit allocation taking into account all audio content.
- step S2 is taken as a multi-mono coding which is carried out using the 3GPP EVS code as described above.
- the method according to the invention can thus be used independently of the core coded (multi-mono, multi-stereo, joint coding) used to represent the channels to be coded.
- the signal thus encoded in the form of a bitstream can be decoded in step S3 either by a local decoder of the encoder, or by a decoder after transmission.
- the signal is decoded to find the channels of the multichannel signal S (for example by several instances of decoder EVS according to a multi-mono decoding).
- Steps S2a, S2b, S3a, S3b represent an alternative embodiment of the encoding and decoding of the multichannel signal B.
- the difference with the encoding of step S2 described above lies in the use of additional processing operations for reducing the number. of channels (“downmix” in English) in step S2a and increase in the number of channels (“upmix” in English) in step S3b.
- Ges encoding and decoding steps are similar to steps S2 and S3 except that the number of respective input and output channels is lower in steps S2b and S3a
- An example of a downmix for a first-order ambisonic input signal is to keep only the W channel; for an ambisonic input signal of order> 1, we can take as a downmix the first 4 components W, X, Y, Z (therefore truncate the signal to order 1).
- An example of upmixing a mono signal consists of applying different room spatial impulse responses (SRIR for "Spatial Room Impulse Response") or different decorrelator filters (of the all-pass type) in the time or frequency domain.
- SRIR Room spatial impulse responses
- decorrelator filters of the all-pass type
- An exemplary embodiment of decorrelation in a frequency domain is given for example in document 3GPP S4-180975, pCR to 26.118 on Dolby VRStream audio profile candidate (dause X6.2.3.5).
- the signal B * resulting from this “downmix” processing is coded in step S2b by a core coded (multi-mono, multi-stereo, joint coding), for example by a mono or multi-mono approach with the coded 3GPP EVS .
- the audio signal input from encoding step S2b and output from decoding step S3 has fewer channels than the original multi-channel audio signal.
- the spatial image represented by the core coded is already significantly degraded even before the coding.
- the number of channels is reduced to a single mono channel, by encoding only the W channel; the input signal is then limited to a single audio channel and the spatial image is therefore lost.
- the method according to the invention makes it possible to describe and reconstruct this spatial image as close as possible to that of the original multichannel signal.
- step S4 information representative of the spatial image of the decoded multichannel signal.
- this information can be a covariance matrix calculated on the decoded multichannel signal or else an information of energy associated with directions of origin of the sound (or in an equivalent way, with virtual points on a unit sphere ).
- the information representative of the original multichannel signal and of the decoded multichannel signal is used in step S5 to determine a set of corrections to be made to the decoded multichannel signal in order to limit the spatial degradations.
- the method described in FIG. 2 can be implemented in the time domain, in full frequency band (with a single band) or else by frequency sub-bands (with several bands), this does not change the operation of the process, each sub-band then being treated separately. If the method is carried out by sub-band, the set of corrections is then determined by sub-band, which causes an additional cost of calculation and of data to be transmitted to the decoder compared to the case of a single band.
- the division into sub-bands can be uniform or non-uniform. For example, we can divide the spectrum of a signal sampled at 32 kHz according to different variants:
- Bark bands (100 Hz wide at low frequencies to 3.5-4 kHz for the last sub-band)
- the 24 Bark bands can optionally be grouped into blocks of 4 or
- ERB bands - for "equivalent rectangular bandwidth" in English - or in 1/3 octave
- sampling frequency for example 16 or 48 kHz
- the invention may also be implemented in a transform domain, for example in the domain of the short-term discrete Fourier transform (STFT) or the domain of the modified discrete cosine transform
- STFT short-term discrete Fourier transform
- modified discrete cosine transform for example in the domain of the short-term discrete Fourier transform (STFT) or the domain of the modified discrete cosine transform
- a mono sound source can be artificially spatialized by multiplying its signal by the values of the spherical harmonics associated with its direction of origin (assuming the signal carried by a plane wave) to obtain as many ambisonic components. For this, we calculate the coefficients for each spherical harmonic for a position determined in azimuth ⁇ and in elevation ⁇ to the desired order:
- ⁇ Y ( ⁇ , ⁇ ) .s
- s the mono signal to spatialize
- Y ( ⁇ , ⁇ ) the encoding vector defining the coeffidents of the spherical harmonics associated with the direction ( ⁇ , ⁇ ) for the order M.
- An example of an encoding vector is given below for order 1 with the SN3D convention and the order of the SI D or FuMa channels:
- the Y ( ⁇ , ⁇ ) coefficients of the spherical harmonics can be found in the book by B. Rafaely, Fundamentals of Spherical Array Processing, Springer, 2015.
- such matrices will serve as a matrix for forming directional beams ("beamforming" in English) describing how to obtain signals characteristic of directions of space in order to carry out an analysis and / or transformations. space.
- beamforming in English
- We therefore define the reciprocal conversion as involving the pseudo-inverse of D: pinv (D) .S D T (DD T ) -1 .S
- FIG. 3 represents a first embodiment of an encoding device and of a decoding device for the implementation of an encoding and decoding method including a method for determining a set of corrections as described. with reference to figure 2.
- the encoder calculates information representative of the spatial image of the original multichannel signal and transmits it to the decoder to enable it to correct the spatial degradation caused by the encoding. This allows during decoding to attenuate spatial artefacts in the decoded ambisonic signal.
- the encoder receives a multichannel input signal of, for example, an FOA ambisonic representation, or HOA, or a hybrid representation with a subset of ambisonic components up to a given partial ambisonic order - the latter case is in fact undue. equivalent way in the case of FOA or HOA where the missing ambisonic components are zero and the ambisonic order is given by the order minimum required to indure all defined components.
- FOA or HQA cases are considered in the remainder of the description.
- the input signal is sampled at 32 kHz.
- the coding is performed in the time domain (on one or more bands), however in variants, the invention can be implemented in a transformed domain, for example after a short discrete Fourier transform. term (STFT) or modified discrete cosine transform (MDCT).
- STFT short discrete Fourier transform
- MDCT modified discrete cosine transform
- a block 310 for reducing the number of channels can be implemented; the input of block 311 is signal B * at the output of block 310 when the downmix is implemented or signal B otherwise.
- the downmix if the downmix is applied, it consists, for example, for an ambisonic input signal of order 1 to keep only the channel W and for an ambisonic input signal of order> 1, to not keep only the first 4 ambisonic components W, X, Y, Z (therefore to truncate the signal at order 1).
- Other types of downmix (such as those described above with a selection of a subset of channels and / or matrixing) can be implemented without modifying the process according to the invention.
- Block 311 encodes the audio signal b'k of B * at the output of block 310 in the case where the downmix step is performed or the audio signal bk of the original multichannel signal B. This signal corresponds to the ambisonic components of the signal. original multichannel if no channel count reduction processing has been applied.
- block 311 uses multi-mono coding (COD) with fixed or variable allocation, where the core codec is the 3GPP EVS standardized codec.
- CDD multi-mono coding
- each bk or b'k channel is coded separately by an instance of the coded; however, in variations other coding methods are possible, for example multi-stereo coding or joint multichannel coding. Therefore, at the output of this coding block 311, an encoded audio signal originating from the original multichannel signal is obtained, in the form of a binary train which is sent to the multiplexer 340.
- block 320 performs a sub-band division.
- this division into sub-bands could reuse equivalent processing operations carried out in blocks 310 or 311; the separation of block 320 is here functional.
- the channels of the original multichannel audio signal are divided into 4 frequency sub-bands of respective width 1 kHz, 3 kHz, 4 kHz, 8 kHz (which amounts to a division of the frequencies according to the 0 -1000, 1000- 4000, 4000-8000 and 8000-16000 Hz.
- Oe slicing can be implemented by means of a short-term discrete Fourier transform (STFT), band-pass filtering in the Fourier domain (by application of a frequency mask), and inverse transform with overlap addition
- STFT discrete Fourier transform
- the sub-bands remain sampled at the same original frequency and the processing according to the invention is applied in the time domain; variants, it is possible to use a filter bank with a critical sampling.
- the sub-band cutting operation generally involves a processing delay which is a function of the type of filter bank used; invention a time alignment can be applied ique before or after encoding-decoding and / or before the extraction of spatial image information, so that the spatial image information is well synchronized in time with the corrected signal.
- full-band processing may be carried out, or the sub-band cutting may be different as explained previously.
- the signal from a transform of the original multichannel audio signal is directly used and the invention is applied in the transformed domain with subband slicing in the transformed domain.
- a high-pass filtering (with a cut-off frequency typically at 20 or 50 Hz), for example in the form of an elliptical IIR filter of order 2 whose frequency of cut-off is preferably set at 20 or 50 Hz (50 Hz in some variants).
- Ge preprocessing avoids a potential bias for the subsequent estimation of covariance during coding; without this preprocessing, the correction implemented in block 390 described later will tend to amplify the low frequencies during full band processing.
- Block 321 determines (Inf. B) information representative of a spatial image of the original multichannel signal.
- this information is energy information associated with directions of origin of sound (associated with directions of virtual speakers distributed over a unit sphere).
- this 3D sphere is discretized by N points (“point” virtual speakers) whose position is defined in spherical coordinates by the directions ( ⁇ n , ⁇ n ) for the nth speaker.
- the loudspeakers are typically placed (almost) uniformly on the sphere.
- a “Lebedev” type quadrature method can for example be used to perform this discretization, according to the references Vl Lebedev, and DN Laikov, “A quadrature formula for the sphere of the 131st algebraic order of accuracy”, Doklady Mathematics, vol. 59, no. 3, 1999, pp. 477- 481 or Pierre Lecomte, Philippe-Aubert Gauthier, Christophe Langrenne, Alexandre Garcia and Alain Berry, On the use of a Lebedev grid for Ambisonics, AES Convention 139, New York, 2015.
- the spatial image of the multichannel signal is for example the SRP method (for "Steered- Response Power ”in English). Indeed, this method consists in calculating the short-term energy coming from different directions defined in terms of azimuth and elevation. For this, as explained previously, similarly to rendering on N speakers, a weighting matrix of the ambisonic components is calculated, then this matrix is applied to the multichannel signal to sum the contribution of the components and produce a set of N acoustic beams (or “beamformers” in English).
- SRP method for "Steered- Response Power ”in English.
- this method consists in calculating the short-term energy coming from different directions defined in terms of azimuth and elevation. For this, as explained previously, similarly to rendering on N speakers, a weighting matrix of the ambisonic components is calculated, then this matrix is applied to the multichannel signal to sum the contribution of the components and produce a set of N acoustic beams (or “beamformers” in English).
- the d n values may vary depending on the type of acoustic beam forming used (delay-sum, MVDR, LCMV, etc.).
- the invention also applies to these variant calculations of the matrix D and of the spatial image.
- the MUSIC method also provides another way of calculating a spatial image, with a subspace approach.
- the invention also applies in this variant of calculation of the spatial image.
- the spatial image can be calculated from a histogram of the intensity vector (at order 1) as for example in the article by S. Tervo, Direction estimation based on sound intensity vectors, Proc. EUSI PCOO, 2009, or its generalization into a pseudo-intensity vector.
- 'histogram (whose values are the number of occurrences of values of arrival directions according to the predetermined directions ( ⁇ n , ⁇ n )) is interpreted as a set of energies according to the predetermined directions.
- Block 330 then quantizes the spatial image thus determined, for example with 16-bit scalar quantization by coefficients (directly using the 16-bit truncated floating point representation). In variations, other scalar or vector quantization methods are possible.
- the information representative of the spatial image of the original multichannel signal is a covariance matrix (of the subbands) of the input channels B. This matrix is calculated as:
- the covariance matrix C (of size Kx (K) being, by definition, symmetric, only one of the lower or upper triangles is transmitted to the quantization block 330 which codes (Q) K (K + 1) / 2 coefficients, K being the number of ambisonic components.
- This block 330 performs a quantization of these coefficients, for example with a scalar quantization on 16 bits by coefficient (by using directly the floating point representation truncated on 16 bits).
- scalar or vector quantization of the covariance matrix can be implemented.For example, we can calculate the maximum value (maximum variance) of the covariance matrix then code by scalar quantization with a logarithmic step, on a number of bits more low (for example 8 bits), the values of the upper (or lower) triangle of the covariance matrix normalized by its maximum value.
- the covariance matrix C could be regularized before quantification in the form C + ⁇ l.
- the quantized values are sent to multiplexer 340.
- the decoder receives in the demultiplexer block 350, a bit stream comprising an encoded audio signal from the original multichannel signal and information representative of a spatial image of the original multichannel signal.
- Block 360 decodes (Q 1 ) the covariance matrix or other information representative of the spatial image of the original signal.
- Block 370 decodes (DEC) the audio signal as represented by the bit stream.
- the decoded multichannel signal is obtained at the output of decoding block 370.
- the decoding implemented in block 370 provides a decoded audio signal which is input to upmix block 371.
- block 371 implements an optional step (UPMIX) of increasing the number of channels.
- this step for the channel of a mono signal , it consists in changing the signal by different responses room spatial impulses (SRIR for “Spatial Room Impulse Response”); these SRIRs are defined in the original ambisonic order of B.
- SRIR room spatial impulses
- Other decorrelation methods are possible, for example the application of all-pass decorrelator filters to the different channels of the signal.
- the block 372 implements an optional step (SB) of division into sub-bands to obtain either sub-bands in the time domain or in a transformed domain.
- SB optional step
- Block 375 determines (Inf ) information representative of a spatial image of the decoded multichannel signal in a manner similar to that described for block 321 (for the original multichannel signal), this time applied to the decoded multichannel signal obtained at the output of the block 371 or block 370 depending on the embodiments decoding.
- this information is energy information associated with directions of origin of the sound (associated with the directions of virtual loudspeakers distributed over a unit sphere).
- an SRP (or other) type method can be used to determine the spatial image of the decoded multichannel signal.
- this information is a covariance matrix of the channels of the decoded multichannel signal. This covariance matrix is then obtained as follows:
- the covariance matrices C and block 380 implements the method of determination (Det.Corr) of a set of corrections as described with reference to FIG. 2.
- a method using rendering (explicit or not) on a virtual loudspeaker is used and in the embodiment of FIG. 5, a method implemented based on a factorization of the Cholesky type is used.
- Block 390 of Figure 3 implements a correction (CORR) of the multichannel signal decoded by the set of corrections determined by block 380 to obtain a corrected decoded multichannel signal.
- CORR correction
- FIG. 4 therefore represents an embodiment of the step of determining a set of corrections. This embodiment is accomplished through the use of virtual speaker rendering.
- the information representative of the spatial image of the original multichannel signal and of the decoded multichannel signal are the respective covariance matrices C and
- blocks 420 and 421 respectively determine the spatial images of the original multichannel signal and the decoded multichannel signal.
- the spatial image of the multichannel signal we can determine the spatial image of the multichannel signal.
- one possible method is the SRP (or other) method which consists in calculating the short-term energy coming from different directions defined in terms of azimuth and elevation.
- the information representative of the spatial image of the original signal (Inf B) received and decoded in 360 by the decoder is the spatial image itself, that is to say information of energy (or a positive quantity) associated with directions of origin of the sound (associated with directions of virtual loudspeakers distributed over a unit sphere), it is then no longer necessary to calculate it at 420.
- This spatial image is then used directly by block 430 described below.
- the determination at 375 of the information representative of the spatial image of the decoded multichannel signal (I nf ) is the spatial image itself of the decoded multichannel signal, then it is no longer necessary to calculate it at 421. This spatial image is then used directly by block 430 described below.
- Block 440 optionally makes it possible to limit (Limit g n ) the maximum value that a gain g n can take. It is recalled here that the positive quantities noted ⁇ ⁇ 2 and can correspond more generally to quantities resulting from of a MUSIC pseudo-spectrum or of the values resulting from a histogram of directions of arrival according to the discretized directions ( ⁇ n , ⁇ n ).
- a threshold is applied to the value of g n . Any value greater than this threshold is forced to be equal to this threshold value.
- the threshold can be for example fixed at 6 dB, so that a gain value outside the range ⁇
- 6 dB is saturated to ⁇ 6 dB.
- This set of gains g n therefore constitutes the set of corrections to be made to the decoded multichannel signal.
- This set of gains is received at the input of the correction block 390 of FIG. 3.
- Block 390 applies, for each virtual loudspeaker, the corresponding gain g n , determined previously. The application of this gain makes it possible to obtain, on this loudspeaker, the same energy as the original signal.
- An acoustic encoding step for example ambisonic encoding by the matrix E, is then implemented to obtain components of the multichannel signal, for example ambisonic components. These ambisonic components are finally summed to obtain the multichannel output signal, corrected (Corr). It is therefore possible to calculate explicitly the channels associated with the virtual loudspeakers, to apply a gain to them, then to recombine the processed channels, or in an equivalent manner to apply the matrix G to the signal to be corrected.
- the normalization factor g norm can be determined without calculating the entire matrix R, because it suffices to calculate only a subset of matrix elements to determine R 00 and therefore g norm ).
- the matrix G or G norm rm thus obtained corresponds to the set of corrections to be made to the decoded multichannel signal.
- Figure 5 now shows another embodiment of the method for determining the set of corrections implemented in block 380 of Figure
- the information representative of the spatial image of the original multichannel signal and of the decoded multichannel signal are the respective covariance matrices C and
- a transformation matrix T to be applied to the decoded signal is determined, so that the spatial image modified after application of the transformation matrix T to the decoded signal is the same as that of the original signal B.
- C BB T is the covariance matrix of B and is the covariance matrix of , in the current frame.
- the matrix A must be a positive definite symmetric matrix (real case) or a definite Hermitian matrix. positive (complex case); in the real case, the diagonal coefficients of L are strictly positive.
- Ax b
- the Cholesky factorization cannot be used as is.
- the matrices L and are lower triangular (respectively upper)
- the transformation matrix T is also lower triangular (respectively upper).
- block 510 forces the covariance matrix C to be positive definite.
- block 520 forces the covariance matrix to be positive definite, by modifying this matrix in the form, where ⁇ is a weak value set for example at 10 -9 and I is the identity matrix.
- block 530 calculates the associated Cholesky factorizations and finds (Det.T) the optimal transformation matrix T in the form
- an alternative resolution can be made with an eigenvalue decomposition.
- the decomposition into eigenvalues consists in factoring a real or complex matrix A of size n x n in the form:
- A Q ⁇ Q -1
- A is a diagonal matrix containing the eigenvalues ⁇ i and Q is the matrix of eigenvectors.
- the stability of the solution from one frame to another is typically poorer than with a Cholesky factorization approach. To this instability are added larger approximations of calculation potentially larger during the decomposition into eigenvalues.
- Block 640 optionally takes care of normalizing (Norm. T) this correction.
- a normalization factor is therefore calculated so as not to amplify frequency zones.
- the normalization factor g norm can be determined without calculating the entire matrix R, because it suffices to calculate only a subset of matrix elements to determine R 00 (and therefore g norm ).
- the matrix T or T norm thus obtained corresponds to the set of corrections to be made to the decoded multichannel signal.
- the block 390 of FIG. 3 performs the step of correcting the decoded multichannel signal by applying the transformation matrix T or T norm directly to the decoded multichannel signal, in the ambisonic domain, to obtain the ambisonic signal of output corrected (corr).
- FIG. 6 A second embodiment of an encoder / decoder according to the invention will now be described in which the method for determining the set of corrections is implemented at the encoder.
- Figure 6 describes this embodiment.
- This figure therefore represents a second embodiment of an encoding device and of a decoding device for the implementation of a coding and decoding method. including a method for determining a set of corrections as described with reference to FIG. 2.
- the method of determining the set of corrections is carried out to the encoder which then transmits this set of corrections to the decoder.
- the decoder decodes this set of corrections to apply it to the decoded multichannel signal.
- Oe embodiment therefore involves implementing a local decoding at the encoder, this local decoding is represented by blocks 612 to 613.
- the blocks 610, 611, 620 and 621 are identical respectively to the blocks 310, 311, 320 and 321 described with reference to FIG. 3.
- Block 612 implements local decoding (DEc_loc) in connection with the coding performed by block 611.
- the local decoding can consist of a complete decoding from the binary train coming from the block 611 or, preferably, it can be integrated into the block 611.
- the decoded multichannel signal is obtained at the output of local decoding block 612.
- the local decoding implemented in block 612 makes it possible to obtain a decoded audio signal which is sent as input to block 613 of upmix.
- block 613 implements an optional step (UPMIX) of increasing the number of channels.
- this step for the channel of a mono signal , it consists in convolving the signal by different room spatial impulse responses (SRIR for “Spatial Room Impulse Response”); these SRIRs are defined in the original ambisonic order of B.
- SRIR room spatial impulse responses
- Other decorrelation methods are possible, for example the application of all-pass decorrelator filters to the different channels of the signal.
- the block 614 implements an optional step (SB) of division into sub-bands to obtain either sub-bands in the time domain or in a transformed domain.
- Block 615 determines (Inf) information representative of a spatial image of the decoded multichannel signal similarly to what has been described for blocks 621 and 321 (for the original multichannel signal), applied this time. to the decoded multichannel signal obtained at the output of block 612 or of block 613 according to the modes for performing local decoding. This block 615 is equivalent to block 375 of figure
- this information is energy information associated with directions of origin of sound (associated with directions of virtual speakers distributed over a unit sphere) .
- an SRP or other type method can be used to determine the spatial image of the decoded multichannel signal.
- this information is a covariance matrix of the channels of the decoded multichannel signal. This covariance matrix is then obtained as follows: up to a normalization factor (in the real case) or up to a normalization factor (in the complex case)
- the covariance matrices C and , block 680 implements the method for determining (Det.Gorr) a set of corrections as described with reference to FIG. 2.
- a method using speaker rendering is used and in the embodiment of FIG. 5, a method implemented directly in the ambisonic domain based on a factorization of the Cholesky type or by eigenvalue decomposition is used.
- the determined set of corrections is a set of gains g n for a set of directions ( ⁇ n , ⁇ n ) defined by a set of virtual loudspeakers.
- This set of gains can be determined in the form of a correction matrix G as described with reference to FIG. 4.
- This set of gains (Gorr.) Is then coded at 640.
- the coding of this set of gains can consist in coding the correction matrix G or G norm .
- the matrix G of size KxK is symmetrical, so according to the invention it is possible to code only the lower or upper triangle of G or G norm , i.e.
- Kx (K + 1) / 2 values In general, the values on the diagonal are positive.
- the coding of the matrix G or G norm is carried out by scalar quantization (with or without a sign bit) depending on whether the values are outside the diagonal or not.
- G norm the coding of the matrix G or G norm is carried out by scalar quantization (with or without a sign bit) depending on whether the values are outside the diagonal or not.
- G norm we can omit to code and transmit the first value of the diagonal (corresponding to the omnidirectional component) of G norm because it is always at 1; for example in the ambisonic case of order 1 to
- other scalar or vector quantization methods (with or without prediction) could be used.
- the determined set of corrections is a transformation matrix T or T norm which is then coded at 640.
- the matrix T of size KxK is triangular in the variant using Cholesky factorization and symmetric in the variant using the eigenvalue decomposition; thus according to the invention it is possible to code only the lower or upper triangle of T or T norm , ie Kx (K + 1) / 2 values.
- the values on the diagonal are positive.
- the coding of the T or T norm matrix is performed by scalar quantization (with or without a sign bit) depending on whether the values are outside the diagonal or not.
- other scalar or vector quantization methods could be used.
- Block 640 thus encodes the determined set of corrections and sends the encoded set of corrections to multiplexer 650.
- the decoder receives in the demultiplexer block 660, a bit stream comprising an encoded audio signal from the original multichannel signal and the encoded set of corrections to be applied to the decoded multichannel signal.
- Block 670 decodes (Q -1 ) the encoded set of corrections.
- Block 680 decodes (DEC) the encoded audio signal received in the stream.
- the decoded multichannel signal is obtained at the output of decoding block 680.
- the decoding implemented in block 680 provides a decoded audio signal which is input to upmix block 681.
- block 681 implements an optional step (UPMIX) of increasing the number of channels.
- this step for the channel of a mono signal, it consists in convolving the signal by different responses room spatial impulses (SRIR for “Spatial Room Impulse Response”); these SRIRs are defined in the original ambisonic order of B.
- SRIR room spatial impulses
- Other decorrelation methods are possible, for example the application of all-pass decorrelator filters to the different channels of the signal.
- the block 682 implements an optional step (SB) of division into sub-bands to obtain either sub-bands in the time domain or in a transformed domain and the block 691 groups the sub-bands to find the output multichannel signal .
- SB optional step
- Block 690 implements a correction (CORR) of the multi-channel signal decoded by the set of corrections decoded at block 670 to obtain a corrected decoded multi-channel signal (Corr).
- CORR correction
- the set of corrections is a set of gains as described with reference to FIG. 4, this set of gains is received at the input of the correction block 690.
- the set of gains is in the form of a correction matrix directly applicable to the decoded multichannel signal, defined, for example in the form
- G E.diag ([g 0 ... g N-1 ]).
- D or G norm g norm .G, this matrix G or G norm is then applied to the decoded multichannel signal S to obtain the ambisonic output signal corrected (Corr).
- the block 690 receives a set of gains g n , the block 690 applies for each virtual loudspeaker, the corresponding gain g n.
- the application of this gain makes it possible to obtain, on this loudspeaker, the same energy as the original signal.
- An acoustic encoding step for example ambisonic encoding, is then implemented to obtain components of the multichannel signal, for example ambisonic components. These ambisonic components are then summed to obtain the multichannel output signal, corrected (Corr).
- the transformation matrix T decoded at 670 is received at the input of the correction block 690.
- block 690 performs the step of correcting the decoded multichannel signal by applying the T or T norm transformation matrix directly to the decoded multichannel signal, in the ambisonic domain, to obtain the corrected ambisonic output signal ( Corr).
- FIG. 7 shows a DCOD encoding device and a DDEC decoding device; within the meaning of the invention, these devices being dual from each other (in the sense of “reversible”) and connected to each other by a communication network RES.
- the DCOD coding device comprises a processing circuit typically including:
- a memory ⁇ EM1 for storing instruction data of a computer program within the meaning of the invention (these instructions can be distributed between the DOOD encoder and the DDEC decoder);
- an interface INT1 for receiving an original multichannel signal B for example an ambisonic signal distributed over different channels (for example four channels W, Y, Z, X at order 1) with a view to its coding in compression within the meaning of the invention;
- processor PROC1 for receiving this signal and processing it by executing the computer program instructions stored in the memory ⁇ BM1, with a view to its coding
- COM 1 communication interface for transmitting the coded signals via the network.
- the DDEC decoding device comprises its own processing circuit, typically including:
- a memory ⁇ EM2 for storing instruction data of a computer program within the meaning of the invention (these instructions can be distributed between the DOOD encoder and the DDEC decoder as indicated above);
- a PAOC2 processor for processing these signals by executing the computer program instructions stored in the memory ⁇ EM2, with a view to their decoding;
- an output interface INT2 to deliver the corrected decoded signals (Corr) for example in the form of ambisonic channels W..X, with a view to their reproduction.
- FIG. 7 illustrates an example of a structural embodiment of a codec (encoder or decoder) within the meaning of the invention.
- Figures 3 to 6 commented above describe in detail rather functional embodiments of these coded.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Multimedia (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Mathematical Physics (AREA)
- Theoretical Computer Science (AREA)
- Pure & Applied Mathematics (AREA)
- Mathematical Optimization (AREA)
- Mathematical Analysis (AREA)
- General Physics & Mathematics (AREA)
- Algebra (AREA)
- Stereophonic System (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR1910907A FR3101741A1 (fr) | 2019-10-02 | 2019-10-02 | Détermination de corrections à appliquer à un signal audio multicanal, codage et décodage associés |
| PCT/FR2020/051668 WO2021064311A1 (fr) | 2019-10-02 | 2020-09-24 | Détermination de corrections à appliquer a un signal audio multicanal, codage et décodage associés |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| EP4042418A1 true EP4042418A1 (fr) | 2022-08-17 |
| EP4042418B1 EP4042418B1 (fr) | 2023-09-06 |
Family
ID=69699960
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP20792467.1A Active EP4042418B1 (fr) | 2019-10-02 | 2020-09-24 | Détermination de corrections à appliquer a un signal audio multicanal, codage et décodage associés |
Country Status (10)
| Country | Link |
|---|---|
| US (1) | US12051427B2 (fr) |
| EP (1) | EP4042418B1 (fr) |
| JP (1) | JP7664232B2 (fr) |
| KR (1) | KR20220076480A (fr) |
| CN (1) | CN114503195B (fr) |
| BR (1) | BR112022005783A2 (fr) |
| ES (1) | ES2965084T3 (fr) |
| FR (1) | FR3101741A1 (fr) |
| WO (1) | WO2021064311A1 (fr) |
| ZA (1) | ZA202203157B (fr) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117041856A (zh) * | 2021-03-05 | 2023-11-10 | 华为技术有限公司 | Hoa系数的获取方法和装置 |
Family Cites Families (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| SE0400998D0 (sv) * | 2004-04-16 | 2004-04-16 | Cooding Technologies Sweden Ab | Method for representing multi-channel audio signals |
| KR20070005468A (ko) * | 2005-07-05 | 2007-01-10 | 엘지전자 주식회사 | 부호화된 오디오 신호의 생성방법, 그 부호화된 오디오신호를 생성하는 인코딩 장치 그리고 그 부호화된 오디오신호를 복호화하는 디코딩 장치 |
| KR100644715B1 (ko) * | 2005-12-19 | 2006-11-10 | 삼성전자주식회사 | 능동적 오디오 매트릭스 디코딩 방법 및 장치 |
| TW200742275A (en) * | 2006-03-21 | 2007-11-01 | Dolby Lab Licensing Corp | Low bit rate audio encoding and decoding in which multiple channels are represented by fewer channels and auxiliary information |
| US9025775B2 (en) * | 2008-07-01 | 2015-05-05 | Nokia Corporation | Apparatus and method for adjusting spatial cue information of a multichannel audio signal |
| EP2175670A1 (fr) * | 2008-10-07 | 2010-04-14 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Rendu binaural de signal audio multicanaux |
| EP2345027B1 (fr) * | 2008-10-10 | 2018-04-18 | Telefonaktiebolaget LM Ericsson (publ) | Codage et décodage audio multicanal conservant l'énergie |
| WO2010097748A1 (fr) * | 2009-02-27 | 2010-09-02 | Koninklijke Philips Electronics N.V. | Codage et décodage stéréo paramétriques |
| EP2600612A4 (fr) * | 2010-07-30 | 2015-06-03 | Panasonic Ip Man Co Ltd | Dispositif de décodage d'image, procédé de décodage d'image, dispositif de codage d'image, et procédé de codage d'image |
| JP5949270B2 (ja) | 2012-07-24 | 2016-07-06 | 富士通株式会社 | オーディオ復号装置、オーディオ復号方法、オーディオ復号用コンピュータプログラム |
| EP2717261A1 (fr) * | 2012-10-05 | 2014-04-09 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Codeur, décodeur et procédés pour le codage d'objet audio spatial à multirésolution rétrocompatible |
| CN104282309A (zh) * | 2013-07-05 | 2015-01-14 | 杜比实验室特许公司 | 丢包掩蔽装置和方法以及音频处理系统 |
| SG11201600466PA (en) * | 2013-07-22 | 2016-02-26 | Fraunhofer Ges Forschung | Multi-channel audio decoder, multi-channel audio encoder, methods, computer program and encoded audio representation using a decorrelation of rendered audio signals |
| KR102244379B1 (ko) | 2013-10-21 | 2021-04-26 | 돌비 인터네셔널 에이비 | 오디오 신호들의 파라메트릭 재구성 |
| EP3007167A1 (fr) | 2014-10-10 | 2016-04-13 | Thomson Licensing | Procédé et appareil de compression à faible débit binaire d'une représentation d'un signal HOA ambisonique d'ordre supérieur d'un champ acoustique |
| EP3067886A1 (fr) * | 2015-03-09 | 2016-09-14 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Codeur audio de signal multicanal et décodeur audio de signal audio codé |
| FR3048808A1 (fr) * | 2016-03-10 | 2017-09-15 | Orange | Codage et decodage optimise d'informations de spatialisation pour le codage et le decodage parametrique d'un signal audio multicanal |
-
2019
- 2019-10-02 FR FR1910907A patent/FR3101741A1/fr active Pending
-
2020
- 2020-09-24 CN CN202080069491.9A patent/CN114503195B/zh active Active
- 2020-09-24 WO PCT/FR2020/051668 patent/WO2021064311A1/fr not_active Ceased
- 2020-09-24 ES ES20792467T patent/ES2965084T3/es active Active
- 2020-09-24 EP EP20792467.1A patent/EP4042418B1/fr active Active
- 2020-09-24 US US17/764,064 patent/US12051427B2/en active Active
- 2020-09-24 BR BR112022005783A patent/BR112022005783A2/pt unknown
- 2020-09-24 JP JP2022520097A patent/JP7664232B2/ja active Active
- 2020-09-24 KR KR1020227013459A patent/KR20220076480A/ko active Pending
-
2022
- 2022-03-16 ZA ZA2022/03157A patent/ZA202203157B/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| KR20220076480A (ko) | 2022-06-08 |
| US20220358937A1 (en) | 2022-11-10 |
| CN114503195B (zh) | 2024-12-31 |
| US12051427B2 (en) | 2024-07-30 |
| JP7664232B2 (ja) | 2025-04-17 |
| FR3101741A1 (fr) | 2021-04-09 |
| ES2965084T3 (es) | 2024-04-10 |
| BR112022005783A2 (pt) | 2022-06-21 |
| ZA202203157B (en) | 2022-11-30 |
| WO2021064311A1 (fr) | 2021-04-08 |
| JP2022550803A (ja) | 2022-12-05 |
| CN114503195A (zh) | 2022-05-13 |
| EP4042418B1 (fr) | 2023-09-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11832080B2 (en) | Spatial audio parameters and associated spatial audio playback | |
| EP2374123B1 (fr) | Codage perfectionne de signaux audionumeriques multicanaux | |
| US8817991B2 (en) | Advanced encoding of multi-channel digital audio signals | |
| EP2143102B1 (fr) | Procede de codage et decodage audio, codeur audio, decodeur audio et programmes d'ordinateur associes | |
| CN117136406A (zh) | 组合空间音频流 | |
| EP2145167A2 (fr) | Procede de codage et decodage audio, codeur audio, decodeur audio et programmes d'ordinateur associes | |
| EP4042418B1 (fr) | Détermination de corrections à appliquer a un signal audio multicanal, codage et décodage associés | |
| EP4172986B1 (fr) | Codage optimisé d'une information représentative d'une image spatiale d'un signal audio multicanal | |
| EP4226368B1 (fr) | Quantification de paramètres audio | |
| EP4627572B1 (fr) | Codage audio spatial paramétrique | |
| EP4533449A1 (fr) | Titre: codage audio spatialisé avec adaptation d'un traitement de décorrélation | |
| CN120418863A (zh) | 神经网络模型进行立体声解码的方法及解码器 | |
| EP4371108A1 (fr) | Quantification vectorielle spherique optimisee | |
| FR3118266A1 (fr) | Codage optimisé de matrices de rotations pour le codage d’un signal audio multicanal |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20220415 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| GRAP | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOSNIGR1 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: GRANT OF PATENT IS INTENDED |
|
| INTG | Intention to grant announced |
Effective date: 20230512 |
|
| GRAS | Grant fee paid |
Free format text: ORIGINAL CODE: EPIDOSNIGR3 |
|
| GRAA | (expected) grant |
Free format text: ORIGINAL CODE: 0009210 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE PATENT HAS BEEN GRANTED |
|
| AK | Designated contracting states |
Kind code of ref document: B1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| REG | Reference to a national code |
Ref country code: GB Ref legal event code: FG4D Free format text: NOT ENGLISH |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: EP |
|
| REG | Reference to a national code |
Ref country code: IE Ref legal event code: FG4D Free format text: LANGUAGE OF EP DOCUMENT: FRENCH |
|
| REG | Reference to a national code |
Ref country code: DE Ref legal event code: R096 Ref document number: 602020017369 Country of ref document: DE |
|
| REG | Reference to a national code |
Ref country code: LT Ref legal event code: MG9D |
|
| RAP4 | Party data changed (patent owner data changed or rights of a patent transferred) |
Owner name: ORANGE |
|
| REG | Reference to a national code |
Ref country code: NL Ref legal event code: MP Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: GR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20231207 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: SE Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: RS Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: NO Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20231206 Ref country code: LV Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: LT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: HR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: GR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20231207 Ref country code: FI Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| REG | Reference to a national code |
Ref country code: AT Ref legal event code: MK05 Ref document number: 1609638 Country of ref document: AT Kind code of ref document: T Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: NL Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: IS Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20240106 |
|
| REG | Reference to a national code |
Ref country code: ES Ref legal event code: FG2A Ref document number: 2965084 Country of ref document: ES Kind code of ref document: T3 Effective date: 20240410 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: AT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: SM Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: RO Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: IS Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20240106 Ref country code: EE Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: CZ Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: AT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: SK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: PT Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20240108 |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: PL |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: LU Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230924 |
|
| REG | Reference to a national code |
Ref country code: BE Ref legal event code: MM Effective date: 20230930 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: PL Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: LU Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230924 |
|
| REG | Reference to a national code |
Ref country code: DE Ref legal event code: R097 Ref document number: 602020017369 Country of ref document: DE |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: MC Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| REG | Reference to a national code |
Ref country code: IE Ref legal event code: MM4A |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: IE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230924 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: DK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PLBE | No opposition filed within time limit |
Free format text: ORIGINAL CODE: 0009261 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: CH Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230930 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: MC Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: IE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230924 Ref country code: DK Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 Ref country code: CH Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230930 Ref country code: SI Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| 26N | No opposition filed |
Effective date: 20240607 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: BE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20230930 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: BG Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: BG Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: CY Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT; INVALID AB INITIO Effective date: 20200924 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: HU Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT; INVALID AB INITIO Effective date: 20200924 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: DE Payment date: 20250820 Year of fee payment: 6 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: IT Payment date: 20250820 Year of fee payment: 6 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: GB Payment date: 20250820 Year of fee payment: 6 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: FR Payment date: 20250821 Year of fee payment: 6 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: TR Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT Effective date: 20230906 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: ES Payment date: 20251001 Year of fee payment: 6 |