Technical Field
-
Embodiments according to the invention comprise apparatuses, methods and computer programs for audio signal processing based on inter-channel-level-difference and side signal component manipulation.
-
Embodiments according to the invention comprise audio processing systems, methods and computer programs for audio signal processing based on transient enhancement, stage width enhancement and ambience enhancement.
-
Embodiments according to the invention comprise apparatuses, systems, methods and computer programs for performing an acoustic boost.
Background of the Invention
-
There are many tools and techniques for processing audio signals. Yet, in view of the plethora of different kinds of audio signals to be processed, it is still focus of ongoing research how to process such signals having different and in particular contrasting characteristics, so as to enable rendering of a respective audio scene with good quality and no, or at least only few, artifacts. Beyond that, a quality of the audio reproduction is additionally dependent on a surrounding of a respective listener, e.g. whether the listener is located in an opera hall or within a more confined space, such as a car.
-
Hence, there is a need for improved concepts for audio signal processing that allow achieving a better compromise between a quality of the reproduced audio scene and a robustness and a flexibility of the concept, e.g. regarding audio signals with different characteristics and/or different listening locations.
-
This is achieved by the subject matters of the independent claims of the present application.
-
Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application.
Summary of the Invention
-
Embodiments of the invention will be presented according to a first and second aspect. This structuring of embodiments is provided in order to facilitate understanding the different details, functionalities and features of embodiments.
-
In this regard, it is to be noted that this structuring according to a first and second aspect is not to be understood in a limiting manner. In other words, details, functionalities and features disclosed in the context of an embodiment according to the first aspect may be incorporated individually or in combination in an embodiment according to the second aspect and vice versa.
-
The same applies for embodiments that are not discussed with regard to a specific aspect.
-
In particular, an audio processing system according to the second aspect of the invention may comprise any of the apparatuses as disclosed in the context of the first aspect, for example for applying the stage width enhancement.
First aspect of the invention:
-
Embodiments according to the first aspect of the invention comprise an apparatus for processing an audio signal, e.g. a stage width enhancement or a stage width enhancement block. The apparatus is configured to provide, on the basis of an input audio signal (e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement) a first modified audio signal (e.g. a first modified multi-channel audio signal; e.g. a stage width enhanced stereo output (audio) signal in a time domain) in which inter-channel-level-differences, e.g. between corresponding audio signal components in different channels, are increased (wherein, for example, an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent) when compared to inter-channel-level-differences in the input audio signal, e.g. using an inter-channel-level-difference increaser; e.g. using an inter-channel-level difference booster.
-
The apparatus is further configured to provide, on the basis of the input audio signal, e.g. on the basis of the stereo input in the time domain, a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components, e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference, are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
-
The apparatus is further configured to combine, e.g. using a linear combination; e.g. using a weighted linear combination, the first modified audio signal and the second modified audio signal, e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or, for example, in order to obtain a stage width enhanced stereo output (audio) signal.
-
This embodiment is based on the finding that boosting or increasing inter-channel-level-differences allows manipulating an audio scene, so as to off center a perceived position of an audio source in the audio scene, e.g. stereo image, further to sides of the audio scene and hence away from the center of the audio scene.
-
The inventors recognized that based on a manipulation of the inter-channel-level-differences, this may be achieved, for example, without modifying the loudness of the sources and, for example, without modifying the timbre, loudness and/or position of sources panned to the center. In other words, a manipulation of positional sound cues for a listener may be performed without altering characteristics, or at least some characteristics of the sounds themselves.
-
Furthermore, manipulation of side signal components relative to centered signal components allows a manipulation of components so as to increase the perceived width of the audio scene for the listener.
-
Hence, in simple words, increasing inter-channel-level-difference and side signal extraction and manipulation may both achieve an increase in perceived width of the audio scene.
-
However, the inventors recognized that these two concepts are susceptible to producing artifacts for different kinds of audio signals or different portions of an audio signal having different characteristics. The inventors hence recognized that a combination of inter-channel-level-difference increase and side signal extraction, although seemingly achieving a same or similar functionality, allows providing said functionality with increased robustness.
-
The inventors recognized that, surprisingly, the two concepts may even cancel each other out with regard to artifacts whilst still allowing efficiently increasing a stage width of the audio scene to be rendered.
-
Hence, embodiments may not only allow addressing a wide plurality of different audio signals (e.g. to enhance a perceived width thereof), but as well allow synergistically reducing respective artifacts in the perceived audio scene. Therefore, a processing of inter-channel-level-difference and side signal extraction to provide first and second signal components may be performed in parallel or consecutively, and it has been found to provide high audio quality for many different types of signals.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to determine inter-channel-level differences (e.g. inter-channel level difference values Dx(n,k) (also referred to as D(n,k or as Dx(f,k) or D(f,k))); wherein is should be noted that, at least in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices (e.g. bin (f) and band (n) indices), and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of the input audio signal, e.g. between corresponding time-frequency bins of different channels of the input audio signal.
-
It is to be noted that in any embodiment band-wise and bin-wise processing may be performed. Hence, embodiments discussed with regard to a bin-wise processing may as well be processed band-wise and vice versa. Hence, in some embodiments, indices n and f may be used interchangeably.
-
The apparatus is further configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g. Y1(n,k),Y2(n,k), is increased when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X1(n,k),X2(n,k), associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
-
The inventors recognized that adapting the inter-channel-level differences frequency-wise and optionally portion-wise, hence, for example per set of corresponding frequency bins and for example per set of corresponding time indices, allows manipulating spatial cues of the audio scene efficiently, for example, for providing a particularly immersive experience of the audio scene. Accordingly, the inventive approach may be applied on different frequency bands independently.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that a total energy in corresponding spectral bins of the first modified audio signal, e.g. Y1(n,k), Y2(n,k), is, for example at least substantially, e.g. with a deviation of no more than 10 percent, equal to a total energy in corresponding spectral bins of the input audio signals, e.g. X1(n,k), X2(n,k), associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
-
The inventors recognized that a quality of the perceived audio scene may be improved by maintaining the total energy in corresponding spectral bins, so as to change a relationship of the channels in the form of level differences between the channels, without changing the overall loudness, and hence also without changing the perceived loudness.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in such a manner that an inter-channel level difference, when considered on a logarithmic scale, is scaled, e.g. linearly scaled; e.g. scaled by multiplication, by a predetermined value.
-
This may, for example be performed, such that an inter-channel level difference between corresponding time frequency bins of the first modified audio signal (e.g. between corresponding time frequency bins of different channels of the first modified audio signal) on a logarithmic scale is linearly scaled (e.g. multiplied by a predetermined factor, e.g. alpha) when compared to an inter-channel level difference between corresponding time frequency bins of input audio signal.
-
The inventors recognized that a scaling performed on a logarithmic scale may allow for a precisely selectable (because of the scale) impact of a audio scene width enhancement.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to determine spectral weights, e.g. G1(n,k), G2(n,k), in dependence on inter-channel-level differences, e.g. inter-channel level difference values Dx(n,k), for a plurality of corresponding time-frequency bins of the input audio signal, e.g. between corresponding time-frequency bins of different channels of the input audio signal.
-
Furthermore, the apparatus may be further configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined spectral weights, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal, e.g. Y1(n,k), Y2(n,k)).
-
This approach may allow for a good compromise between computational complexity and audio scene enhancement.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain a logarithmic representation, e.g. Dx(n,k), of an inter-channel level difference between corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal (e.g. using a computation of powers or energies P1(n,k) and P2(n,k) the corresponding time frequency bins of the input audio signal and using a formation of a logarithm of a ratio between the computed powers or energies). Furthermore, the apparatus is configured to multiply the logarithmic representation, e.g. Dx(n,k), of the inter-channel level difference between the corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal with a predetermined scaling value, e.g. alpha, or with an negative version of the predetermined scaling value, e.g. - alpha, in order to obtain an intermediate scaling value, e.g. H1(n,k) or H2(n,k), in a logarithmic domain. In addition, the apparatus is configured to derive one or more intermediate scaling values in a linear domain, e.g. K1(n,k) or K2(n,k), from the one or more intermediate scaling values in the logarithmic domain (e.g. from H1(n,k) or H2(n,k); e.g. using an evaluation of an exponential function, wherein a respective exponent is defined by an respective intermediate scaling value in the logarithmic domain, or is defined in dependence on a respective intermediate scaling value in the logarithmic domain. Moreover, the apparatus is configured to apply an energy normalization, in order to obtain the spectral weights, e.g. G1(n,k), G2(n,k), in dependence on the intermediate scaling value in the linear domain, e.g. K1(n,k) and/or K2(n,k).
-
This approach may allow for a good compromise between computational complexity and audio scene enhancement.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain an inter-channel level difference value Dx(n,k) according to wherein P1(n,k) is a power or an energy in a time frequency bin of a first channel signal of the multi-channel input audio signal, wherein P2(n,k) is a power or an energy in a time frequency bin of a second channel signal of the multi-channel input audio signal, wherein n is a frequency index, e.g. a frequency bin index, and wherein k is a time index.
-
Furthermore, the apparatus is configured to obtain a first intermediate scaling value in a logarithmic domain, H1(n,k), and a second intermediate scaling value in a logarithmic domain, H2(n,k), according to otherwise 0 , otherwise 0, wherein α is a predetermined value, wherein the apparatus is configured to obtain a first intermediate scaling value in a linear domain, K1, and a second intermediate scaling value in a linear domain, K2, according to wherein the apparatus is configured to obtain a first spectral weight G1(n,k) and a second spectral weight G2(n,k) according to wherein the apparatus is configured to obtain a first channel signal Y1(n,k) and a second channel signal Y2(n,k) of the first modified audio signal according to wherein X1(n,k) is a spectral value of a time-frequency bin of the first channel signal of the input audio signal, wherein X2(n,k) is a spectral value of a time-frequency bin of the second channel signal of the input audio signal, wherein Y1(n,k) is a spectral value of a time-frequency bin of the first channel signal of the first modified audio signal, and wherein Y2(n,k) is a spectral value of a time-frequency bin of the second channel signal of the first modified audio signal.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to derive a factor, by which the inter-channel level difference is to be changed, e.g. K1(n,k) or K2(n,k), from an inter-channel level difference value (e.g. from a logarithmic representation of a ratio between energies in corresponding time frequency bins of two channels of the input audio signal; e.g. from Dx(n,k), e.g. such that a logarithmic representation of factor, by which the inter-channel level difference is to be changed, e.g. H1(f,m) or H2(f,m), is proportional to the logarithmic representation of a ratio between energies in corresponding time frequency bins of two channels of the input audio signal, e.g. to Dx(f,m)).
-
Furthermore, the apparatus is configured to provide spectral weights, e.g. G1(n,k), G2(n,k), for scaling corresponding time frequency bins of two channels of the input audio signal, such that ratio of the spectral weights is only determined by the factor. This may allow determining the spectral weights with low computational effort.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to derive spectral weights for scaling corresponding time frequency bins of two channels of the input audio signal, such that ratio of the spectral weights is equal to a potency, e.g. having a real-valued exponent, of a ratio of powers in the corresponding time frequency bins. The inventors recognized that such an approach may allow for an efficient weighting of the channel signals.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to map, e.g. directly, e.g. without a computation of a panning index or the like, a value representing the inter-channel level difference, e.g. Dx(n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain, using a linear or piecewise-linear mapping.
-
Optionally, this may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign. Optionally, such an approach may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
-
The inventors recognized that the linear or piecewise-linear mapping may allow for an efficient determination of signal weights for inter channel level adjustment.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to, for example directly, derive spectral weights for scaling corresponding time frequency bin without computing a panning index. This may allow increasing an efficiency of the audio processing.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain the first modified audio signal using an, optionally selective (e.g. total energy maintaining or total power maintaining) opposite modification of intensities of signal components in corresponding time frequency bins of different channels of the input audio signal, which increases inter-channel-level differences. This may allow providing a balanced adaptation of inter channel levels, e.g. so as to maintain an overall loudness of the audio signal.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain the second modified audio signal using an increase of energies of off-center signal components (e.g. of corresponding signal components having a comparatively larger inter-channel level difference or of corresponding signal components having a comparatively large inter-channel time difference or of corresponding signal components having a comparatively large inter-channel-phase difference) in corresponding time frequency bins of different channels of the input audio signal.
-
Alternatively or in addition, the apparatus is configured to obtain the second modified audio signal using a reduction of energies of centered signal components (e.g. of corresponding signal components having a comparatively small inter-channel level difference or of corresponding signal components having a comparatively small inter-channel time difference or of corresponding signal components having a comparatively small inter-channel-phase difference) in corresponding time frequency bins of different channels of the input audio signal.
-
Hence, the inventors recognized that an increase in signal energy differences between center and off center components may allow a manipulating in accordance with a signal manipulation based on inter channel level differences (for the first modified audio signal), but with a different or even complementary behavior regarding artifacts, so as to supplement the first modified audio signal.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain the second modified audio signal using a processing in which a ratio between energies of off-center signal components and centered signal components is increased, e.g. by more than 20 percent, or by more than 50 percent.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to obtain the first modified audio signal using a processing in which a ratio between energies of off-center signal components and centered signal components is maintained, within a tolerance of +/-10 percent.
-
Hence, in view of the above-discussed features, it is to be noted that the inventors recognized that a combination of energy conserving and energy altering approaches, for a same or similar type of manipulation, e.g. stage width enhancement, may be combined so as to address diverse input audio signals (e.g. to achieve a good width enhancement for many types of signals) and/or to cancel out approach-specific artifact behavior and/or to shift an emphasis on the one or the other approach depending on the type or the characteristics of the input signal.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to scale corresponding signal components in corresponding time frequency bins of different channels of the input audio signal in dependence on an inter-channel-level difference between the respective corresponding signal components, and/or in dependence on an inter-channel time difference between the respective corresponding signal components, and/or in dependence on an inter-channel phase difference between the respective corresponding signal components, in order obtain the second modified audio signal.
-
This may allow for a particularly improved perceptual quality of a respective rendered audio scene.
-
According to embodiments of the first aspect of the invention, the apparatus is configured to compute a gain mask (e.g. a gain mask in a time-frequency domain, e.g. a gain mask which extracts side signal components) for a scaling of time-frequency bins of the input audio signal on the basis of the input audio signal (e.g. in order to obtain the second modified audio signal, e.g. on the basis of a time-frequency domain representation of the input audio signal).
-
Furthermore, the apparatus is configured to scale time-frequency bins of the input audio signal using respective entries of the gain mask, in order to obtain the second modified audio signal, e.g. in order to obtain a time-frequency domain representation of the second modified audio signal.
-
According to embodiments of the first aspect of the invention, the apparatus comprises an inter-channel level difference modifier, wherein the inter-channel level difference modifier is configured to determine inter-channel-level differences (e.g. inter-channel level difference values Dx(n,k); wherein it should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal).
-
The inter-channel level difference modifier is configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of a modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g. Y1(n,k), Y2(n,k), is modified when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X1(n,k), X2(n,k), associated with a same frequency (e.g. having frequency index n) and a same time portion (e.g. having time index k)
-
Furthermore, the inter-channel level difference modifier is configured to map (e.g. directly, e.g. without a computation of a panning index or the like) a value representing the inter-channel level difference, e.g. Dx(n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain, using a linear or piecewise-linear mapping.
-
Optionally the same may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign. Optionally the approach may comprise using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
-
Further embodiments according to the first aspect comprise a method for processing an audio signal, e.g. a method for stage width enhancement, wherein the method comprises providing, on the basis of an input audio signal (e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement) a first modified audio signal (e.g. a first modified multi-channel audio signal; e.g. a stage width enhanced stereo output (audio) signal in a time domain; e.g. a stage width) in which inter-channel-level-differences, e.g. between corresponding audio signal components in different channels, are increased (wherein, for example, an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent) when compared to inter-channel-level-differences in the input audio signal (e.g. using an inter-channel-level-difference increaser; e.g. using an inter-channel-level difference booster).
-
The method further comprises providing, on the basis of the input audio signal, e.g. on the basis of the stereo input in the time domain, a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components (e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference) are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
-
The method further comprises combining (e.g. using a linear combination; e.g. using a weighted linear combination) the first modified audio signal and the second modified audio signal (e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or in order to obtain a stage width enhanced stereo output (audio) signal).
-
The method as described above is based on the same considerations as the above-described apparatus. The method can, by the way, be completed with all features and functionalities, which are also described with regard to the apparatus.
-
Further embodiments according to the first aspect comprise a computer program for performing a method according to an embodiment according to the first aspect, when the computer program runs on a computer.
Second aspect of the invention:
-
Embodiments according to the second aspect comprise an audio processing system for obtaining a processed audio signal on the basis of an input audio signal, wherein the audio processing system is configured, in order to obtain the processed audio signal, to apply a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
-
Furthermore, in order to obtain the processed audio signal, the audio processing system is configured to apply a stage width enhancement, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference) relative to audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene (e.g. having a comparatively small inter-channel level difference, e.g. using a side signal extraction), and/or which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference, e.g. while suppressing audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference).
-
Furthermore, in order to obtain the processed audio signal, the audio processing system is configured to apply an ambience enhancement, which provides decorrelated audio signal components on the basis of the input audio signal or on the basis of a processed version of the input audio signal.
-
The inventors recognized that a combination of transient enhancement, stage width enhancement and ambience enhancement allows improving a wide variety of audio signals, e.g. having different characteristics, and providing an improved quality of a respective rendered audio scene, for example with reduced artifacts.
-
In particular, the inventors recognized that in such a combination of audio processing techniques, in contrast to conventional approaches, emphasizing transients may allow maintaining the sonic characteristic of the input signal (e.g. punch) while still allowing for further modifications such as width enhancement.
-
In other words, in the particular combination, the transient enhancement, stage width enhancement and ambience enhancement synergistically introduce degrees of freedom for improvement of a respective rendered audio scene, while cancelling, at least partially, shortcomings and/or artifacts of each other.
-
The ambience enhancement, for example comprising a transient attenuation, may cancel artifacts introduced by the transient enhancement, for example comprising a transient boost, and vice versa. It was further recognized that such a supplementary effect may be further improved for audio signals which are to be perceptually widened, hence using a stage width enhancement.
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured, for applying the ambience enhancement, to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
-
Furthermore, the audio processing system is configured for applying the ambience enhancement to obtain, on the basis of the input audio signal, a second modified audio signal using an ambience extraction, which may, for example prefer audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation and/or which may, for example, provide the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal.
-
Furthermore, the audio processing system is configured, for applying the ambience enhancement, to combine, e.g. using a weighted combination, the first modified audio signal and the second modified audio signal, in order to obtain a combined audio signal and to apply a decorrelation to the combined audio signal, in order to obtain a processed audio signal, e.g. an ambience-enhanced signal.
-
The inventors recognized that in the context of the ambience enhancement a decomposition and separate processing of a background portion of the input signal (e.g. first modified audio signal) and an ambience portion of the input signal (e.g. second modified audio signal) may be performed. Both processings may comprise a transient suppression, which, in combination with the additional transient enhancement, e.g. in a separate processing step, allows balancing out respective artifacts. In particular a "punch" of the scene may be upheld (e.g. via transient enhancement), whilst still being able to improve a background of the audio scene (e.g. via ambience and background signal processing).
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured to apply the transient enhancement, in order to emphasize transients in the input audio signal relative to non-transient audio signal components of the input audio signal, in order to obtain a transient enhanced audio signal.
-
The audio processing system is further configured to apply the stage width enhancement, to increase inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, and/or to emphasize audio signal components of the input audio signal panned to one of the sides of an audio scene relative to audio signal components of the input audio signal panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, in order to obtain a stage-width-enhanced audio signal.
-
The audio processing system is further configured to apply the ambience enhancement to provide decorrelated audio signal components on the basis of the input audio signal, in order to obtain a decorrelated audio signal. Furthermore, the audio processing system is configured to combine the transient enhanced audio signal, the stage width enhanced audio signal and the decorrelated audio signal, in order to obtain a combined audio signal, wherein the combined audio signal constitutes the processed audio signal, or wherein the audio processing system is configured to derive the processed audio signal from the combined audio signal using a post-processing.
-
Hence, the three distinct enhancements may be performed in parallel, so as to combine respective results thereof. Such an approach may comprise a good computational efficiency because of parallel processing. Furthermore, such an approach may allow for a relatively easy adaptation of different aspects of the audio scene because of the relative independence of the improvements from one another. For making the scene appear "wider", parameters of the stage width enhancement may be tweaked, at least approximately independently, and vice versa for the other enhancements.
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured to apply the transient enhancement, in order to emphasize transients in the input audio signal relative to non-transient audio signal components of the input audio signal, in order to obtain a transient enhanced audio signal, and to apply the ambience enhancement to provide decorrelated audio signal components on the basis of the input audio signal, in order to obtain a decorrelated audio signal.
-
The audio processing system is further configured to combine the transient enhanced audio signal and the decorrelated audio signal, in order to obtain a combined audio signal, and the audio processing system is configured to apply the stage width enhancement, to increase inter-channel-level differences between corresponding audio signal components of different channels of the combined audio signal, and/or to emphasize audio signal components of the combined audio signal panned to one of the sides of an audio scene relative to audio signal components of the combined audio signal panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, in order to obtain a stage-width-enhanced audio signal.
-
Furthermore, the stage width enhanced audio signal may constitute the processed audio signal, or the audio processing system may be configured to derive the processed audio signal from the stage width enhanced audio signal using a post-processing.
-
Hence, the stage width enhancement may be performed after the transient enhancement and the ambience enhancement. The inventors recognized that for signals with significant differences between a foreground and background of the audio scene, a transient enhancement (e.g. for foreground improvement) and ambience enhancement (e.g. for background improvement) before a width enhancement, may allow improved artifact cancellation between transient enhancement and ambience enhancement, because of the targeting of the transients.
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured to apply the stage width enhancement, to increase inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, and/or to emphasize audio signal components of the input audio signal panned to one of the sides of an audio scene relative to audio signal components of the input audio signal panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, in order to obtain a stage-width-enhanced audio signal.
-
The audio processing system is further configured to apply the transient enhancement, in order to emphasize transients in the stage-width-enhanced audio signal relative to non-transient audio signal components of the stage-width-enhanced audio signal, in order to obtain a transient enhanced audio signal, and to apply the ambience enhancement to provide decorrelated audio signal components on the basis of the stage-width-enhanced audio signal, in order to obtain a decorrelated audio signal.
-
The audio processing system is further configured to combine the transient enhanced audio signal and the decorrelated audio signal, in order to obtain a combined audio signal, wherein the combined audio signal constitutes the processed audio signal, or wherein the audio processing system is configured to derive the processed audio signal from the combined audio signal using a post-processing.
-
Hence, the stage width enhancement may be performed before the transient enhancement and the ambience enhancement. This may allow the same advantages as performing the stage width enhancement after transient enhancement and ambience enhancement.
-
According to an embodiment of the second aspect of the invention, the transient enhancement is configured to increase inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, in order to obtain an inter-channel-level-difference enhanced audio signal. Furthermore, the transient enhancement is configured to emphasize transients in the inter-channel-level-difference-enhanced audio signal, relative to non-transient audio signal components of the inter-channel-level-difference-enhanced audio signal, in order to obtain a transient enhanced audio signal.
-
The inventors recognized that a joint modification of transient and inter-channel-level-differences may provide good results regarding a quality of a respective, rendered audio scene. In other words, emphasizing transients in the inter-channel-level-difference-enhanced audio signal, relative to non-transient audio signal components of the inter-channel-level-difference-enhanced audio signal may allow widening a perception of the acoustic foreground of a scene.
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured to apply a post-processing, e.g. to the combined audio signal, or to the stage-width-enhanced audio signal. The post processing comprises one or more of the following functionalities: a dynamic range compression; an equalization; a loudness compensation.
-
It was recognized that each of these functionalities, applied individually or in combination, may allow refining a result of the three parallel or consecutive audio scene enhancements.
-
According to an embodiment of the second aspect of the invention, the post-processing is configured to split-up a signal to be post-processed into a plurality of frequency ranges (e.g. into a low frequency range, a mid frequency range and a high frequency range, e.g. using a crossover filter bank). The post-processing may further be configured to separately post-process the different frequency ranges, e.g. using a respective dynamic range compression and a respective equalization.
-
It was recognized that frequency range-wise post-processing may improve a quality of a respective rendered audio scene.
-
According to an embodiment of the second aspect of the invention, the audio processing system is configured to determine inter-channel-level differences (e.g. inter-channel level difference values Dx(n,k); wherein is should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal).
-
The audio processing system is further configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of a modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g. Y1(n,k), Y2(n,k), is modified when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X1(n,k), X2(n,k), associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
-
The audio processing system is further configured to map (e.g. directly, e.g. without a computation of a panning index or the like) a value representing the inter-channel level difference, e.g. Dx(n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain) using a linear or piecewise-linear mapping.
-
Optionally, this may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign. Optionally this approach may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
-
According to an embodiment of the second aspect of the invention, for applying the stage width enhancement, the audio processing system is configured to provide, on the basis of an input audio signal (e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement) a first modified audio signal (e.g. a first modified multi-channel audio signal; e.g. a stage width enhanced stereo output (audio) signal in a time domain; e.g. a stage width) in which inter-channel-level-differences, e.g. between corresponding audio signal components in different channels, are increased (wherein, for example, an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent) when compared to inter-channel-level-differences in the input audio signal (e.g. using an inter-channel-level-difference increaser; e.g. using an inter-channel-level difference booster).
-
Furthermore, for applying the stage width enhancement, such an audio processing system is configured to provide, on the basis of the input audio signal (e.g. on the basis of the stereo input in the time domain), a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components (e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference) are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
-
Furthermore, for applying the stage width enhancement, such an audio processing system is configured to combine, e.g. using a linear combination; e.g. using a weighted linear combination, the first modified audio signal and the second modified audio signal (e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or in order to obtain a stage width enhanced stereo output (audio) signal).
-
It is to be noted that, embodiments according to the second aspect of the invention may comprise the functionalities, details and/or features (and/or even an apparatus, e.g. as a portion of a respective audio processing system) of any of the embodiments of the first aspect, both individually or taken in combination, for example, for applying the stage width enhancement.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
-
For applying the ambience enhancement, such an audio processing system is further configured to obtain, on the basis of the input audio signal, a second modified audio signal using an ambience extraction (which prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation, for example which provides the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal).
-
Furthermore, for applying the ambience enhancement, such an audio processing system is further configured to combine, e.g. using a weighted combination, the first modified audio signal, or a post-processed version of the first modified audio signal, and the second modified audio signal, or a post-processed version of the second modified audio signal, in order to obtain a combined audio signal.
-
Moreover, for applying the ambience enhancement, such an audio processing system is further configured to apply a decorrelation to the combined audio signal, in order to obtain a processed audio signal, e.g. an ambience-enhanced signal.
-
Hence, the previously discussed decomposition of the input signal in a transient portion and an ambience portion may be supplemented by a decorrelation for further improvement of the perceived quality, e.g. with regard to immersion, of the respective rendered audio signal.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement (or for example for performing the ambience extraction), the audio processing system is configured to extract from the input audio signal ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation, e.g. left-right correlation, (e.g. an inter-channel correlation which is below a predetermined threshold value, or an inter-channel correlation which is smaller than an inter-channel correlation of direct signal components, e.g. such that direct signal components (and also non-transient direct signal components) are at least partially suppressed in the second modified audio signal, while ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation are non-suppressed or enhanced in the second modified audio signal).
-
Alternatively or in addition, for applying the ambience enhancement (or for example for performing the ambience extraction), such an audio processing system is further configured to emphasize ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation over, e.g. relative to, direct signal components and/or signal components having a comparatively high, or for example higher, inter-channel correlation
-
Alternatively or in addition, for applying the ambience enhancement (or for example for performing the ambience extraction), such an audio processing system is further configured to attenuate direct signal components and/or signal components having a comparatively high inter-channel correlation relative to ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to compute a covariance matrix on the basis of a frequency domain representation of the input audio signal (e.g. on the basis of a stereo input signal in a frequency domain, e.g. on the basis of X(f,k) or X(f,m) or X(n,k)), to compute a multichannel parametric Wiener filter on the basis of the covariance matrix, to estimate a direct signal using the multichannel parametric Wiener filter, and to subtract the estimated direct signal from the input audio signal, in order to obtain the second modified audio signal, e.g. Y(f,k) or Y(f,m) or Y(n,k), e.g. to thereby obtain the second modified audio signal such that signal components of the input audio signal having a comparatively higher inter-channel correlation are reduced when compared to signal components having a comparatively lower inter-channel correlation.
-
The inventors recognized that using a Wiener filter on the basis of the covariance for direct signal estimation is computationally efficient. Furthermore, for some applications, or to be more specific for some input signals, the inventors recognized that, in the specific inventive structure with transient enhancement and stage width enhancement, artifacts caused by the Wiener filtering may be compensated by the transient enhancement and stage width enhancement.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to apply a transient suppression to the second modified audio signal, in order to obtain a post-processed version of the second modified audio signal for combination with the first modified audio signal.
-
It was recognized that the transient suppression in the ambience enhancement may counteract artifacts from the transient enhancement and vice versa, whilst still allowing to improve an acoustic foreground (e.g. via the transient enhancement) and an acoustic background (e.g. via ambience enhancement).
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain the first modified audio signal using a spectral weighting process with real weights (e.g. real valued weights).
-
This may allow reducing a computational effort and facilitate implementation of such a computation, e.g. in contrast to using complex weights, whilst still providing sufficiently good results.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain an estimate (e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate) of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals, e.g. X(f,m) for a given frequency index f and for a plurality time indices m, of the input audio signal (wherein subband signals may be signals associated with a single frequency bin, or signals associated with a frequency band comprising a plurality of frequency bins). Furthermore, for applying the ambience enhancement, such an audio processing system is configured to obtain the first modified audio signal in dependence on the estimate of the sustained signal component.
-
The inventors recognized that a subband-wise sustained signal estimation provides good audio scene enhancement, for example with regard to its acoustic background.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T̂(f,m) or T̂(n,k), (e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g. X(f,m) or X(n,k) for a given frequency index f or n and for a plurality time indices m or k) of the input audio signal, and to obtain the first modified audio signal in dependence on the estimate of the transient signal component.
-
The inventors recognized that a subband-wise sustained signal estimation provides good audio scene enhancement, for example with regard to its acoustic foreground.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to combine a plurality of frequency bins of a time-frequency-domain representation of the input audio signal, in order to obtain frequency band values associated with a plurality of frequency bands (wherein the frequency band values are examples of the subband values), and to obtain the estimate of the sustained signal component and/or the estimate of the transient signal component in dependence on the frequency band values.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain spectral weights, e.g. Gg(f,m) or Gg(n,k), associated with frequency bins, e.g. generally speaking with frequency subbands, on the basis of the frequency band values (e.g. using frequency-band-wise estimates of sustained signal components and/or of transient signal components derived from the frequency band values).
-
It was recognized that the weight determination, selective for different frequency bands is particularly efficient.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain temporally smoothened, e.g. low-pass filtered subband signals and/or subband signals purged from peaks, subband signals or temporally smoothened frequency-band signals, e.g. using a combination of frequency bins (frequency subbands) to frequency bands, or temporally smoothened subband signal envelopes or temporally smoothened frequency band signal envelopes (e.g. low-pass filtered subband signal envelopes and/or subband signal envelopes purged from peaks; e.g. low-pass filtered frequency-band signal envelopes or frequency band signal envelopes purged from peaks) on the basis of subband signals representing the input audio signal.
-
As an example, frequency-band signals may, for example, be a special example of subband signals, wherein the term "subband signals" comprises both signals associated with a single frequency bin, signals associated with multiple adjacent frequency bins, and signals associated with a frequency band comprising a plurality of frequency bins.
-
Furthermore, for applying the ambience enhancement, such an audio processing system is configured to obtain the first modified audio signal in dependence on the temporally smoothened (e.g. purged from peaks) subband signals or the temporally smoothened frequency band signals or the temporally smoothened subband signal envelopes or the temporally smoothened frequency band signal envelopes.
-
For example, instead of low-pass filtering the signal (e.g. instead of attenuating high frequencies), the envelope of the signal may be low pass filtered. In other words and as an example, when we low-pass filter the transformed, e.g. STFT, magnitudes coefficients of frequency bins corresponding to one frequency band (e.g. the bin indices 11-15), then we low-pass filter a signal representation and this may result in low-pass filtered / smoothed time trajectories (envelopes).
-
It was recognized that the temporal smoothing allows improving the audio signal enhancement.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain sustained signal envelopes, e.g. Ŝ(f,m) or Ŝ(n,k), using a temporal smoothing, e.g. low pass filtering, of subband time trajectories or frequency band time trajectories (wherein frequency band time trajectories are a special case of subband time trajectories, wherein the term "subband time trajectories" comprises time trajectories comprising only a single frequency bin and time trajectories comprising a plurality of frequency bins) of the input audio signal, e.g. of a subband representation X(f,m) or X(n,k) of the input audio signal, and using an application of a constraint, e.g. Ŝ(f,m)<X(f,m) or Ŝ(n,k)<=magnitude(X(n,k)), requiring that respective sustained signal envelopes, e.g. Ŝ(f,m) or Ŝ(n,k), do not exceed magnitudes of respective, for example associated, subband signals, e.g. |X(f,m)|, or magnitudes of respective frequency band signals of the input audio signal or subband time trajectories of magnitudes of the input audio signal or frequency band time trajectories of magnitudes of the input audio signal.
-
As an example, the subband time trajectories may, for example, represent a temporal evolution of energies in respective subbands of the input audio signal, wherein subbands may, for example, be frequency bands of a filterbank or may, for example, be frequency bins of a time-domain-to-spectral-domain transform or of a time-domain-to-frequency-domain transform.
-
It was recognized that the temporal smoothing together with the introduction of the constraints allows improving the audio signal enhancement.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain transient signal envelopes, e.g. T̂(f,m) or T̂(n,k), using a subtraction of sustained signal envelopes, e.g. Ŝ(f,m) or Ŝ(n,k), from respective subband signals, e.g. X(f,m), or from magnitudes or respective subband signals or from frequency band signals of the input audio signal or from magnitudes of frequency band signals of the input audio signal or from or the subband time trajectories, e.g. of magnitudes, of the input audio signal or from frequency band time trajectories, e.g. of magnitudes, of the input audio signal.
-
The inventors recognized that efficiency may be improved by estimating sustained signal envelopes and determining transient signal envelopes by subtraction, e.g. instead of directly estimating transient signal envelopes. Furthermore, sustained signal envelopes may be estimated more robustly.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain spectral weights, e.g. Gg(f,m) or Gg(n,k), in dependence on sustained signal envelope values (e.g. in dependence on the sustained signal envelope values mentioned before, e.g. Ŝ(n,k)).
-
This may allow determining the weights in efficient and robust manner.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain spectral weights, e.g. Gg(f,m) or Gg(n,k), in dependence on transient signal envelope values (e.g. in dependence on the transient signal envelope values mentioned before, e.g. T̂(n,k)).
-
Hence, an alternative weight determination may be provided, for example, for signal having strong transients, wherein it may me more efficient and/or robust to directly estimate the transients.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain respective spectral weights, e.g. Gg(f,m), in dependence on a ratio between a weighted sum of respective sustained signal envelope values, e.g. Ŝ(f,m) or Ŝ(n,k), or an exponentiated version thereof, e.g. Ŝ(f,m)alpha or Ŝ(n,k)alpha, and of respective transient signal envelope values, e.g. T̂(f,m) or T̂(n,k), or an exponentiated version thereof, e.g. T̂(f,m)alpha or T̂(n,k)alpha, and magnitudes of respective subband signals, e.g. X(f,m) or X(n,k), or frequency band signals, or an exponentiated version thereof, e.g. |X(f,m)|alpha or |X(n,k)|alpha.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain spectral weights Gg(n,k) according to wherein Ŝ(n,k) are sustained signal envelope values, wherein T̂(n,k) are transient signal envelope values, wherein X(n,k) are subband signal values of the input audio signal, wherein |X(n,k)| are magnitudes of subband signal values of the input audio signal, wherein n is a frequency index, wherein k is a time index, wherein α is a parameter, wherein β is a parameter, wherein γ is a parameter between 0 and 1 (and preferably, as an example, larger than or equal to 0.05), and wherein δ is a parameter between 0 and 1 (and preferably, for example, larger than or equal to 0.05).
-
Such a determination of weights may allow for a computationally inexpensive audio scene enhancement.
-
According to an embodiment of the second aspect of the invention, for applying the ambience enhancement, the audio processing system is configured to obtain the sustained signal envelopes Ŝ(n,k) using a low pass filtering of subband time trajectories of magnitudes of a time-frequency-domain representation, e.g. X(f,m) or X(n,k), of the input audio signal, and using an application of a constraint Ŝ(n,k) <|X(n,k)| or of a constraint Ŝ(n,k) <=|X(n,k)|, to obtain the transient signal envelopes T̂(n,k) according to
-
According to an embodiment of the second aspect of the invention, for applying the transient enhancement, the audio processing system is configured to obtain an estimate (e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate) of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g. X(f,m) for a given frequency index f and for a plurality time indices m) of the input audio signal (wherein subband signals may be signals associated with a single frequency bin, or signals associated with a frequency band comprising a plurality of frequency bins).
-
Furthermore, for applying the transient enhancement, such an audio processing system is configured to obtain the modified audio signal, in which transient signal components are enhanced or reduced, in dependence on the estimate of the sustained signal component.
-
According to an embodiment of the second aspect of the invention, for applying the transient enhancement, the audio processing system is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T̂(f,m) or T̂(n,k), (e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g. X(f,m) or X(n,k) for a given frequency index f or n and for a plurality time indices m or k) of the input audio signal, and to obtain the modified audio signal in dependence on the estimate of the transient signal component.
-
Hence, ambience enhancement and transient enhancement may share a common transient determination, therefore reducing the computational complexity, since intermediate results, such as the estimated transients, may be used by multiple (enhancement) modules.
-
Embodiments according to the second aspect of the invention comprise a method for obtaining a processed audio signal on the basis of an input audio signal, wherein the method comprises, in order to obtain the processed audio signal, applying a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
-
The method further comprises, in order to obtain the processed audio signal, applying a stage width enhancement, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene, e.g. having a comparatively large inter-channel level difference, relative to audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, e.g. using a side signal extraction, and/or which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference, e.g. while suppressing audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference).
-
The method further comprises, in order to obtain the processed audio signal, applying an ambience enhancement, which provides decorrelated audio signal components on the basis of the input audio signal or on the basis of a processed version of the input audio signal.
-
The method as described above is based on the same considerations as the above-described audio processing system. The method can, by the way, be completed with all features and functionalities, which are also described with regard to the audio processing system.
-
Further embodiments according to the second aspect comprise a computer program for performing a method according to an embodiment according to the second aspect, when the computer program runs on a computer.
Brief Description of the Drawings
-
The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
- Fig. 1
- shows a schematic view of an apparatus for processing an audio signal according to an embodiment;
- Fig. 2
- shows a schematic view of an apparatus for processing an audio signal with additional optional details, according to an embodiment;
- Fig. 3
- shows a schematic view of an Inter Channel Level Difference Booster according to embodiments of the invention;
- Fig. 4
- shows a schematic view of a side signal extraction unit according to embodiments;
- Fig. 5
- shows a schematic block diagram of a method according to embodiments of the first aspect of the invention;
- Fig. 6 a)-c)
- show schematic views of different variations (e.g. variants) of audio processing systems according to embodiments;
- Fig. 7
- shows a schematic view of a transient enhancement unit, e.g. a transient enhancement block, according to embodiments;
- Fig. 8
- shows a transient enhancement subunit according to an embodiment of the invention;
- Fig. 9
- shows a schematic view of a post processing unit according to embodiments of the invention;
- Fig. 10
- shows a schematic view of a post processing unit with additional, optional features, according to embodiments of the invention;
- Fig. 11
- shows a schematic view of an example of an ambience enhancement unit according to embodiments of the invention;
- Fig. 12
- shows a schematic view of an example of an ambience enhancement unit (also referred to as ambience enhancement block, with additional, optional features, according to embodiments of the invention;
- Fig. 13
- shows a schematic view of an ambience extraction unit according to embodiments of the invention;
- Fig. 14
- shows a schematic view of a detailed block diagram of an audio processing system according to embodiments of the invention;
- Fig. 15
- shows a schematic view of a module for processing an input signal according to embodiments of the invention;
- Fig. 16
- shows a schematic view of a module weight computation unit according to embodiments of the invention; and
- Fig. 17
- shows a schematic block diagram of a method for obtaining a processed audio signal on the basis of an input audio signal according to embodiments of the invention.
Detailed Description of the Embodiments
-
Equal or equivalent elements or elements with equal or equivalent functionality are denoted in the following description by equal or equivalent reference numerals even if occurring in different figures.
-
In the following description, a plurality of details is set forth to provide a more throughout explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present invention. In addition, features of the different embodiments described herein after may be combined with each other, unless specifically noted otherwise.
-
Furthermore, it is to be noted that elements having same ending numerals or same names may comprise same, according, or corresponding functionalities or details, or may be examples of each other, e.g. such as inputs 101 and 201 (corresponding numbers) or stage width enhancement unit 100 and stage width enhancement 620 (corresponding names).
-
In the following figures, different computationally efficient structures for audio signal processing according to embodiments are disclosed. Furthermore, the embodiments shown allow for good artifact cancellation due to combining processing techniques with complementary artifact behavior.
-
Fig. 1 shows a schematic view of an apparatus for processing an audio signal according to an embodiment. Fig. 1 shows apparatus 100, which may, for example, be a stage width enhancement apparatus or a stage width enhancement block, e.g. of an audio processing system.
-
The apparatus comprises an Inter Channel Level Difference Booster 110, a Side Signal Extraction unit 120 and a combiner 130. The apparatus is provided with an input audio signal 101.
-
The Inter Channel Level Difference Booster 110 is configured to provide, on the basis of an input audio signal 101 a first modified audio signal 111 in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal 101.
-
The Side Signal Extraction unit 120 is configured to provide, on the basis of the input audio signal, a second modified audio signal 121 in which side signal components are emphasized relative to centered signal components when compared to the input audio signal 101, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal 101, and/or in which only side signal components are included while centered signal components are suppressed.
-
The combiner 130 is configured to combine the first modified audio signal 111 and the second modified audio signal 121.
-
The combined signal is indicated, as an example, as a Stage Width Enhances Signal 102, as for the optional form of the apparatus 100 as a stage width enhancement block or apparatus.
-
Furthermore, as an example, combiner 130 is indicated as a summation, however, embodiments are not limited to a specific form of combination. For example, besides addition, weighted addition and/or even subtraction may be performed.
-
Fig. 2 shows a schematic view of an apparatus for processing an audio signal with additional optional details, according to an embodiment.
-
Fig. 2 shows apparatus 200 comprising an Inter Channel Level Difference Booster 210, a Side Signal Extraction unit 220, a combiner 230, a transformer 240 and inverse transformers 250. The apparatus is provided with an input audio signal 201.
-
The optional transformer 240 is shown, as an example, as a short term Fourier transformer. However, it is to be noted that other forms of transforms or for example filterbanks, see e.g. Fig. 15, may be used as well (and accordingly for corresponding inverse transformers 250 or inverse filterbanks). The input signal, which is shown, as an optional example, as a stereo signal in time domain, may be time-domain-to-spectral-domain transformed or time-domain-to-frequency-domain transformed by transformer 240. In particular, a frequency band-wise or bin-wise processing may be performed.
-
The Inter Channel Level Difference Booster 210 is configured to provide, on the basis of the transformed input audio signal 201 a first modified audio signal in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal 201 (and hence, for example, its transformed counterpart).
-
The Side Signal Extraction unit 220 is configured to provide, on the basis of the transformed input audio signal, a second modified audio signal, in which side signal components are emphasized relative to centered signal components when compared to the input audio signal 201 (and hence, for example, its transformed counterpart), and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal 201 (and hence, for example, its transformed counterpart), and/or in which only side signal components are included while centered signal components are suppressed.
-
Optionally, the Side Signal Extraction unit 220 may be configured to increase or boost energies of off-center signal components and/or to attenuate or reduce energies of centered signal components.
-
The inter-channel-level-difference boost as well as the side signal extraction may be performed in corresponding time frequency bins (or time frequency bands) of different channels of the input audio signal.
-
Results of the Inter Channel Level Difference Booster 210 and the Side Signal Extraction unit 220 are provided to inverse transformers 250 for performing an inverse transformation, e.g. frequency-domain-to-time-domain-transformation or spatial-domain-to-time-domain transformation.
-
The modified and re-transformed signals may then be combined in combiner 230 to provide an output signal 202. As an example, the output signal is shown, for the example of the apparatus 200 being a stage width enhancement block or apparatus, as a stage width enhanced stereo output signal.
-
Fig. 3 shows a schematic view of an Inter Channel Level Difference Booster according to embodiments of the invention. Fig. 3 shows Inter Channel Level Difference Booster 300 comprising an inter-channel-level-difference calculation unit 310, a gain calculation unit 320, comprising, as optional features, a subunit 320l for calculation a gain 321l for a left channel 301l of the input signal and a subunit 320r for calculation a gain 321r for a right channel 301r of the input signal and multiplication units 330.
-
The Inter Channel Level Difference Booster 300 is provided with an input signal, wherein, as shown as an optional feature in Fig. 3, said input signal may be provided in frequency domain. The input signal may referred to as X(f, k). As mentioned before, f may refer to a frequency bin. A frequency band-wise processing is possible as well. However, herein X(n,k) may be used, for example, interchangeably. It is to be noted that in general f may denote the frequency bin index, k the time index and n optionally the frequency band index or for some embodiments n may be used interchangeably for f. As shown, the input signal may comprise two (or more) input channels, hence a left and right input signal 3011 and 301r.
-
The input signals 301l and 301r are provided to inter-channel-level-difference calculation unit 310. As an optional feature, as indicated in Fig. 3, the inter-channel-level-difference calculation may be performed per frequency bin, e.g. as indicated by index f. However, embodiments, and in particular the embodiment according to Fig. 3, are not limited to such an approach.
-
As a result of inter-channel-level-difference calculation unit 310, the inter-channel-level-difference, referred to as D(f,k) (e.g. also referred to as Dx(f,k) or respectively Dx(n,k)) is obtained and provided to the gain calculation unit 320. It is to be noted that inter-channel-level differences may be determined for a plurality of corresponding time-frequency bins of the input audio signal.
-
As shown as an optional feature, the gain calculation is performed channel-wise in respective subunits (and, for example, as another optional feature per frequency bin).
-
As results of the gain unit 320, gain values 321l (also referred to as G1(f, k)) and 321r (also referred to as G2(f, k)) are provided. These gains are used to scale the respective input signals 301l and 301r (wherein, for example, signals 301l and 301r correspond to each other regarding their indices (f,k)), using multiplication units 330, so as to provide output signals 302l (e.g. Y1(f,k)) and 302r (e.g. Y2(f,k)) in frequency domain, (also referred to as Y(f,k)).
-
As indicated by indices (f, k) according to embodiments, the weight determination and scaling may, for example, be performed for a plurality of corresponding time-frequency bins of the input audio signal. Such a processing may be performed consecutively (e.g. for streaming) or in parallel (e.g. when the whole audio signal to be processed is fully available).
-
The inter channel level difference between corresponding spectral bin values of such modified audio signals 302l, 302r, e.g. Y1(f,k), Y2(f,k), may be increased and hence boosted, when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X1(f,k), X2(f,k), for example associated with a same frequency (referring to index f or n) and a same time portion, e.g. having time index k.
-
The first modified audio signal 111 may be or may comprise audio signals 302l, 302r.
-
Furthermore, the weight calculator 320 may be configured to determine the weights (here as an example G1 and G2 (embodiments are not limited to stereo signals), so as to maintain a total energy of the channels from input 301(l+r) to output 302(l+r), for example associated with a same frequency (referring to index f or n) and a same time portion, e.g. having time index k.
-
As an optional feature, gain unit 320 may be configured to determine the gains G1 and G2, so that an increase in intensity in one portion may correspond to a decrease in intensity in another. This may allow a coupled, but balanced inter-channel-level difference adaptation.
-
In particular, the weights may be calculated in a direct manner, for example without computing a panning index, and hence with good computational efficiency.
-
Inter Channel Level Difference Booster 300 may be an example for Inter Channel Level Difference Boosters 110 and/or 210.
-
Fig. 4 shows a schematic view of a side signal extraction unit according to embodiments. Fig. 4 shows side signal extraction unit 400 comprising a gain calculation unit 420 and a multiplication unit 430. Unit 400 may correspond to side signal extraction unit 210. f may denote the frequency bin index, and k the time index.
-
Input signal 401 (here shown as an example as stereo signal), referred to as X(f,k) is provided to the gain unit 420, for determining gain values 411, referred to as G(f,k). These gain values are used to scale the input signal 401 using multiplication unit 430 to obtain an output signal 402 (here accordingly shown as an example as stereo signal), referred to as Y(f, k).
-
The gain 411 may be determined so as to change a relation in energy between off-centered and centered signal components in corresponding time frequency bins of different channels of the input audio signal, e.g. as indicated by indices (f, k). The gain may hence define an attenuation of the center and/or boost of the off-center components.
-
Yet, optionally, the gains 411 may be determined so as to maintain a ratio between energies of off-center signal components and centered signal components, within a tolerance of +/-10 percent or for example +/-5 percent or for example +/-1 percent.
-
Optionally the gain unit 420 may be configured to determine the gain values 411 for scaling based on corresponding signal components in corresponding time frequency bins of different channels of the input audio signal in dependence on an inter-channel-level difference between the respective corresponding signal components, and/or in dependence on an inter-channel time difference between the respective corresponding signal components, and/or in dependence on an inter-channel phase difference between the respective corresponding signal components.
-
Hence, the gain determination may comprise a good flexibility and adaptability.
-
It is to be noted that, signal 402 may be an example for the second modified audio signal 121.
-
Fig. 5 shows a schematic block diagram of a method according to embodiments of the first aspect of the invention. Fig. 5 shows method 500 comprising providing, 510, on the basis of an input audio signal, a first modified audio signal in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal, providing, 520, on the basis of the input audio signal, a second modified audio signal in which side signal components are emphasized relative to centered signal components when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal and/or in which only side signal components are included while centered signal components are suppressed and combining, 530 the first modified audio signal and the second modified audio signal.
-
Fig. 6 a)-c) show schematic views of different variations (e.g. variants) of audio processing systems according to embodiments. Fig. 6a)-c) may show overview block diagrams shows, hence block diagrams of various ways how the basic algorithms according to embodiments may be combined.
-
Fig. 6 a)-c) show audio processing systems 600a, 600b, 600c comprising a transient enhancement unit 610, a stage width enhancement unit 620, and an ambience enhancement unit 630.
-
As further, optional features, the audio processing systems 600a, 600b, 600c comprise combination units 640 and a post processing units 650.
-
The audio processing systems 600a, 600b, 600c are configured to receive an input audio signal 601.
-
Furthermore, the audio processing systems 600a, 600b, 600c are configured to apply a transient enhancement, using transient enhancement unit 610, which emphasizes transients in the input audio signal (see Variants A and B), or in a processed version of the input audio signal (see Variant C), relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal.
-
The audio processing systems 600a, 600b, 600c are further configured to apply a stage width enhancement, using stage width enhancement unit 620, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal (see Variants A and C), or between corresponding audio signal components of different channels of a processed version of the input audio signal (see Variant B), and/or which emphasizes audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to one of the sides of an audio scene relative to audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to a center of the audio scene, and/or which selectively extracts audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to one of the sides of an audio scene.
-
The audio processing systems 600a, 600b, 600c are further configured to apply an ambience enhancement, using ambience enhancement unit 630, which provides decorrelated audio signal components on the basis of the input audio signal (see Variants A and B) or on the basis of a processed version of the input audio signal (see Variant C).
-
Furthermore, the audio processing systems 600a, 600b, 600c are configured to obtain a processed audio signal 602 on the basis of the input audio signal 601.
-
As optional features, combining units 640, as an example as adding units, are shown. However, embodiments are not limited to a specific kind of combining. As examples, summation, weighted summating and/or even subtraction may be performed.
-
As another optional feature, the audio processing systems 600a, 600b, 600c comprise a post processing unit 650, for further improvement of the enhanced and combined signal (based on enhancements 610, 620, 630).
-
As shown in Fig. 6 a) to c), embodiments are not limited to a singular arrangement of the transient enhancement 610, stage width enhancement 620, and ambience enhancement 630. Contrarily, any of the units 610, 620 and 630 may receive a version of the input signal which is preprocessed by any of the other units. The preprocessing may even comprise applying more than one enhancement and then combining the intermediate signals (see e.g. Variant B).
-
As shown in Fig. 6a) the enhancements 610, 620 and 630 may be performed in parallel or consecutively or a combination of parallel enhancement and consecutive enhancement (see Variants B and C) may be performed.
-
A parallelization of transient enhancement and ambient enhancement and subsequent combination may be beneficial with regard to canceling out artifacts. The transient enhancement may comprise emphasizing transients and the ambience enhancement may comprise suppressing transients, so as to selectively enhance a foreground of an audio scene to be rendered based on the input signal 601 (e.g. via transient enhancement 610) and to selectively enhance a background of the audio scene to be rendered based on the input signal 601 (e.g. via ambience enhancement 630), wherein a combination of the signals may allow combining the individually optimized aspects of the audio scene, wherein transient manipulation artifacts may cancel each other out because of the different kinds of manipulation, namely emphasizing vs attenuation.
-
Here, it is to be noted that the stage width enhancement unit 620 may comprise any or all of the features as the stage width enhancement units as discussed in the context of Fig. 1 to 5. In other words, apparatus 100 as well as apparatus 200 may be examples of stage width enhancement unit 620. The stage width enhancement unit may hence comprise any or all of the further optional features as disclosed in the context of Fig. 3 and 4.
-
Fig. 7 shows a schematic view of a transient enhancement unit, e.g. a transient enhancement block, according to embodiments.
-
Transient enhancement unit 700 comprises a transform unit 710, an Inter Channel Level Difference Booster 720, a Transient Enhancement subunit 730 a subtraction unit 740 and an inverse transformer 750. Transient enhancement unit 700 may be an example for the Transient enhancement unit 610. Inter Channel Level Difference Booster 300 may be an example for Inter Channel Level Difference Booster 720.
-
It is to be noted that the transformer 710 and the inverse transformer 750 are optional units. As shown, the transformation used may be a short term Fourier transform, however, as previously discussed, embodiments are not limited to the same.
-
As shown, input audio signal 701 (e.g. a preprocessed audio signal) is provided to the Inter Channel Level Difference Booster 720, which is configured to increase inter-channel-level differences between corresponding (e.g. with regard to (f, k), e.g. with regard to (n, k)) audio signal components of different channels of the input audio signal (e.g. as in Variant A, B of Fig. 6), or between corresponding audio signal components of different channels of a processed version of the input audio signal (e.g. as in Variant C of Fig. 6), in order to obtain an inter-channel-level-difference enhanced audio signal.
-
As an example, the input audio signal 701 is therefore shown as a stereo signal, hence having two channels.
-
A result of the Inter Channel Level Difference Booster 720 is provided to the transient enhancement subunit 730, which is configured to emphasize transients in the inter-channel-level-difference-enhanced audio signal, relative to non-transient audio signal components of the inter-channel-level-difference-enhanced audio signal, in order to obtain a transient enhanced audio signal.
-
As another optional feature, the output signal may be provided based on a subtraction between a result of the transient enhancement subunit 730 and the Inter Channel Level Difference Booster 720, and the inverse transformation of unit 750.
-
Fig. 8 shows a transient enhancement subunit according to an embodiment of the invention. The transient enhancement subunit may, for example be a transient enhancement / transient suppression module (e.g. depending on a parametrization for the weight computation).
-
Transient enhancement subunit 800 comprises as combiner 810, weight computation units 8201, ..., 820n, a weight calculation unit 830 and a multiplication unit 840. Transient enhancement subunit 800 may be an example for transient enhancement subunit 730.
-
Again, it is to be noted that f may denote the frequency bin index, k the time index and n optionally the frequency band index (or for example interchangeably a bin index). It is to be noted that embodiments, in general, may optionally comprise frequency bin-wise processing, and/or frequency band-wise processing, e.g. if several bins are combined to a frequency band. However, a single bin may as well be understood according to some embodiments as a frequency band. Hence, f and n may be used interchangeably according to some embodiments or f as bin index and n as band index.
-
Transient enhancement subunit 800 is provided with input signal 801. As an example, the input signal may be a frequency domain signal, e.g. referred to as X(f, k). The combiner 810 may be configured to combine frequency bins of the input signal 801 to frequency bands.
-
Band-wise combined inputs X(1, k) to X(n, k) may then be provided to respective weight computation units 8201, ..., 820n to determine band-wise weights, which may then be used to determine, based thereon, bin-wise weights, using weight calculation unit 830.
-
These bin-wise weights are then used to scale input signal 801, using multiplication unit 840, so as to obtain an output signal in frequency domain 802, referred to as Y(f, k).
-
Reference is made to Fig. 9. Fig. 9 shows a schematic view of a post processing unit according to embodiments of the invention. Fig. 9 shows post processing unit 900 comprising a dynamic range compression unit 910, an equalizer 920 and a loudness compensation unit 930. However, regarding the example of Fig. 9, it is to be noted that the post processing unit 900 may comprise an arbitrary selection of the units 910, 920, 930. Furthermore, with regard to Fig. 9, it is to be noted that the sequence of blocks can be altered.
-
Post processing unit 900 may be an example of post processing unit 650. Hence, the post-processing may be applied to a combined audio signal 901, or to a stage-width-enhanced audio signal 901, in order to obtain a post processed signal 902.
-
Reference is made to Fig. 10. Fig. 10 shows a schematic view of a post processing unit with additional, optional features, according to embodiments of the invention. Fig. 10 shows post processing unit 1000, which may be an example of post processing unit 900 or 650, with a loudness normalization unit 1010, a crossover filter bank 1020 and dynamic range compression units 1030l, 1030m, 1030h and equalizers 1040l, 1040m, 1040h, as well as a combination unit 1050.
-
As shown in Fig. 10, the post-processing unit may be configured to split-up a signal to be post-processed into a plurality of frequency ranges, as shown as an example using a filterbank 1020, to separately post-process the different frequency ranges (I=low, m=mid, h=high). Optionally, the separately processed signals may be combined using combination unit 1050, for example as a summation or weighted summation, in order to obtain the output signal 1002. As indicated in Fig. 10, the post processing may be performed in time domain, e.g. using a time domain input signal 1001 to provide a time domain output signal 1002. As an example, the signals are indicated as stereo signals.
-
With regard to Fig. 10, it is to be noted that the sequence of blocks and the number of frequency ranges can be altered. Hence, embodiments may comprise more, less or different blocks an frequency ranges. The loudness normalization 1010 is optional as well.
-
Fig. 11 shows a schematic view of an example of an ambience enhancement unit according to embodiments of the invention. Fig. 11 shows ambience enhancement unit 1100, which may be an example for ambience enhancement unit 630 comprising, as optional features, a transient suppression unit 1110, an ambience extraction unit 1120 and another optional transient suppression unit 1110, as well as a combination unit 1130 and an optional decorrelation unit 1140.
-
Hence, the input signal 1101 may be provided to the ambience extraction unit 1120, for example, in order to extract audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation. A result thereof (for example for providing an ambience signal processing), as well as input signal 1101 (for example for providing a background signal processing) are optionally provided to transient suppression units 1110. These may comprise the details and functionalities as shown in Fig. 8, however with weights chosen, so as to attenuate transients instead of boosting the same. Respective results may be combined using combination unit 1130, e.g. for performing an addition or a weighted summation. The combined result may be decorrelated using the decorrelation unit 1140.
-
In summary, Figs. 1, 9 and 11 may show detailed block diagram, hence showing more detailed block diagrams explaining the basic algorithms according to embodiments, e.g. with more optional details in comparison to Fig. 6a)-c).
-
Fig. 12 shows a schematic view of an example of an ambience enhancement unit (for example also referred to as ambience enhancement block), with additional, optional features, according to embodiments of the invention. Fig. 12 shows ambience enhancement unit 1200 comprising, as optional features, transform units 1210, as an example in the form of short term Fourier transform units (e.g. having functionalities as previously discussed), optional transient suppression units 1220 an ambience extraction unit 1230, an optional Inter Channel Level Difference Booster 1240 and optional inverse transform units 1250, as an example in the form of inverse short term Fourier transform units (e.g. having functionalities as previously discussed). Furthermore, ambience enhancement unit 1200 comprises a combination unit 1260 and as an optional feature a decorrelation unit 1270.
-
Hence, the background signal processing may, as well as the ambience signal processing, be performed in frequency or spatial domain. Furthermore, additionally inter channel level differences may be boosted for the background signal, using unit 1240.
-
Reference is made to Fig. 13. Fig. 13 shows a schematic view of an ambience extraction unit according to embodiments of the invention. Fig. 13 shows ambience extraction unit 1300 comprising a covariance computation unit 1310, a filtering unit 1320 and a direct signal estimator 1330, as well as a subtraction unit 1340. The ambience extraction unit 1300 may be an example for the ambience extraction unit 1230.
-
The input signal 1301 is shown, as an example, as a stereo input signal in frequency domain and may be referred to as X(f, k). Again, frequency band-wise processing may be performed as well.
-
As shown, using covariance computation unit 1310 a covariance matrix on the basis of the multi-channel frequency domain representation 1301 of the input audio signal may be determined, for a subsequent Wiener filtering (which may be computed using unit 1320) in order to estimate a direct signal using direct signal estimator 1330.
-
As shown, as an optional feature, the ambience extraction unit 1300 may be configured to subtract the estimated direct signal from the input audio signal 1301 (e.g. the frequency domain representation thereof) to obtain output signal 1302, e.g. in the shown form of an Ambient output signal in frequency domain.
-
Reference is made to Fig. 14, showing a schematic view of a detailed block diagram of an audio processing system (e.g. as a whole system) according to embodiments of the invention. Fig. 14 shows audio processing system 1400 comprising the subsystems 200, 700, 1000 and 1200 as previously discussed in Fig. 2, 7, 10 and 12. System 1400 may show an example for system 600a shown in Fig. 6a. The input signal is indicated as stereo input signal 1401, e.g. corresponding to input signal 601 and the output signal is indicated as stereo output signal 1402, e.g. corresponding to input signal 602. Intermediate results of the transient enhancement block 700, the stage width enhancement block 200 and the ambience enhancement block 1200 may be combined using combination unit 640 (e.g. as discussed regarding Fig. 6a). As shown the blocks 200, 700 and 1200 may share a common transform unit 1410, e.g. corresponding to units 240, 710 and 1210.
-
In this regard, it is to be noted that Fig. 3, 4, 8 and 13 may be Module Block diagrams, for example constituting portions of such a system 1400.
-
In the following a method for processing an audio signal according to embodiments are further discussed.
-
For example, as a an abstract for the following, it may be noted that one aim of methods according to embodiments is to enhance an audio signal with respect to listener preference and general sound quality. To this end (e.g. hence according to embodiments) algorithms are, for example, combined (e.g. individually or in combination) that modify spatial characteristics, dynamic characteristics, and/or tonal characteristics. These may, for example, comprise dedicated modules transient enhancement, ambience enhancement, and stage width enhancement.
-
Problems, addressed by at least some embodiments:
One aim of at least some methods according to embodiments is to process an audio signal such that it is preferred over the unprocessed signal when rating its sound quality.
-
Audio signals like musical recordings, audio books or movie sounds are created by recording, synthesizing, processing and mixing (adding) sound sources and finally processing the mixture. The processing of the audio signals is done in analog or digital domain and comprises the application of various devices (analog) or algorithms (digital), where digital processing is more often used than analog.
-
The processing of sound sources and mix aims to create a final mix of high sound quality where all sound sources are (or at least should be) audible and intelligible. The applied tools, procedures and best practices have been improved in the history of music and audio production. The preferences of human listeners are also evolving over time. Consequently, musical recordings of same genres but different eras have different characteristics.
-
Presumably, sound quality of recent music productions is often preferred over older music productions, also because listener preference is a driving force that influences music productions .
-
One aim of at least some methods according to embodiments is to modify characteristics of older music productions such that they sound as having been produced more recently.
-
The characteristics of audio signals comprise spatial characteristics, dynamic characteristics, tonal characteristics, loudness, timbral attributes and for example more.
-
Some characteristics are in general correlated with preferences of human listeners.
-
It is well-known that signals are preferred when being reproduced at slightly higher sound pressure level (SPL) as long as distinctly below the pain threshold.
-
Increasing the level results in higher loudness level where loudness denotes a perceptual attribute in contrast to the physical level. But this preference can, for example only, be achieved with restrictions (e.g. one or more of the following) that loudness should not exceed armful levels (exposure to sounds at high levels can cause hearing loss), all reproduced signals should have consistent levels, and reproduced sound should not mask other relevant sounds in the environment, e.g. conversations or alarms. Human hearing processes sound in frequency bands with non-linear (compressive) mechanisms and temporal integration on time-variant stimuli. Computational models of loudness may determine specific loudness in each frequency band and may accumulate these to compute the overall loudness. Due to the compressive effect, stimuli that are dense in the time-frequency representation may be perceived as being louder than spare signals when reproduced at same SPL.
-
Stereophonic sound reproduction aims to emulate a rich spatial sound image by a small number of signal channels and loudspeakers. This may be achieved by spatial cues of sound signals at both ears. Sound sources reproduced at different levels by each transducer may be perceived as coming from different locations and form a stereo image. Uncorrelated sounds, e.g., wind noise, reverberations, of signals reproduced with slight variations at different positions may not be localized as coming from one direction and may thereby contribute to a natural and aesthetically pleasing experience. Therefore, spatial characteristics are important for listening preference. Enhancing decorrelation and inter-channel level differences may yield richer stereophonic images and can increase listener preference.
-
Another aim of methods according to at least some embodiments is to compensate for limitations of the environment in which reproduced sound is presented. The ideal listening environment for loudspeaker reproduction is, for example, quiet (e.g. loudness of external sound sources is below or near threshold of hearing), has walls, floor and ceiling reflecting sound waves and thereby contributing uncorrelated and diffuse sounds. The size of the room should be such that early reflections reach the ears within a time span such that reflections are not heard as echoes and not merged with the direct sounds and reduce their clarity. The geometry of the room and the positions of listeners and loudspeakers should not result in large scaling of energy at different frequencies, that is, the effect of sound reproduction in the room than modelled a linear time-invariant system should have a smooth magnitude transfer function.
-
The material of walls, floor, ceiling should reflect sound to some extend without amplifying particular frequencies.
-
Vehicles, for example, have limitations in these respects, and at least some methods according to embodiments may compensate them. When listening to reproduced sound in a (optionally driving) vehicle, noise is emitted from engine, tires on roads and air turbulences. Environmental noise partially masks the reproduced program and thereby reduces its overall loudness and modifies the tonal balance. Masking follows the same principles as loudness perception, it is frequency selective and nonlinear. Since softer sounds are masked more than louder sounds, a dynamic processing that reduces the level differences between soft and loud sounds is beneficial in this respect.
-
The geometry of the cabin and material of the interior may be less appropriate for music listening as compared to a living room, for example. Therefore, the contribution of a great listening room to spatial cues of program material may be missing.
-
The same may hold for positioning of loudspeakers and this may result in spatial cues of less quality as intended during production.
-
Therefore, improvement of spatial cues is implemented in at least some methods according to embodiments.
-
A further aspect according to embodiments relates to the development of mobile sound reproduction technology, in particular mobile phones, laptops, small wireless loudspeakers. These devices are often used outside and in noisy environments.
-
The dissemination of this technology has influenced listening habits (e.g., listening more in noisy environments) and also practices in music production, in particular the use of dynamic range compression in mastering. Modern music productions are more dynamic range compressed, and presumably because this 1) improves masking of environmental noise, 2) is perceived louder, and 3) is simply preferred by artists and listeners.
-
Modern music productions sound clearer, denser, have distinct spectral and dynamic characteristics compared to older productions. Transient signals appear to be louder and bass drums in popular music are reproduced at higher levels. These are other characteristics which may be improved by application of methods according to embodiments.
-
In order to achieve some or even all of the above aims and effects, methods according to embodiments may comprise one or more of a transient processing, a stage width enhancement, an inter-channel level difference enhancement, a side signal extraction, an ambience enhancement and/or a post-processing.
-
Fig. 6a)-c) may show various ways how the above basic algorithms according to embodiments are combined, Fig. 1, 9, 11 may show more detailed block diagrams explaining such algorithms.
-
Reference is made to Fig. 15. Fig. 15 shows a schematic view of a module for processing an input signal according to embodiments of the invention. Module 1500 may be configured to boost Inter-channel level differences for audio rendering and/or to process transients.
-
Module 1500 is provided with an input signal 1501, which is processed by a filterbank 1510. As an example, the input signal is indicated as x(t) in time domain, however an optional previous transformation may be present. Filterbank 1510 may be configured, as shown to provide frequency bin-wise (or optionally band-wise e.g. subband-wise) input signals, here as an example referred to as X(1,k), ..., X(n,k).
-
Bin-wise (or band-wise) weights G may be determined, using weight computation units 1520, based on the filter outputs, for scaling the filter outputs using multiplication units 1530, in order to provide band-wise output signals Y(1,k), ..., Y(n,k). These may be retransformed to a time domain output signal y(t) using an inverse processing of filterbank 1540.
-
Fig. 15 may show a module and respectively method for boosting inter-channel level differences for audio rendering according to embodiments is disclosed. Module 1500 may correspond to Inter Channel Level Difference Booster 300, with additional filterbank 1510 for time to frequency transformation. Hence, computation shown in Fig. 3 may correspond to the computation shown in Fig. 15.
-
In summary or for example as an abstract for the following, such a method for manipulating inter-channel level differences of an audio signal having two channels may have the aim to widen the stereo image of the signal, which may be useful for many audio rendering systems, e.g. upmixing and surround sound.
-
The following section may be split in subsections Aim (1), Algorithmic Description (2).
-
Aim (1) of at least some embodiments for manipulating inter-channel level differences of an audio signal: Such an inventive module, e.g. module 1500, may, for example, increase inter-channel level differences of an audio signal in time-frequency domain. This may, for example, result in the audible effect that the perceived position in the stereo image of sources panned off center are further moved away from the center of the stereo image. This may, for example, be achieved without modifying the loudness of the sources and/or without modifying the timbre, loudness and/or position of sources panned to the center.
-
This may, for example, be applied for upmixing of two-channel stereo signals to surround signals, and/or stereophonic enhancement for rendering applications.
-
Algorithmic description (2) of at least some embodiments for manipulating inter-channel level differences of an audio signal: The processing may, for example, work on frequency bands that are processed independently of each other. Therefore, a filterbank (or, alternatively, frequency transform) may, for example, be used or even required, optionally together with its inverse processing for computing the broad-band output signal (or time signal). This can, for example, be implemented by means of short-term Fourier transform (STFT) and its inverse. The block diagram of Fig. 15 may illustrate the spectral weighting.
-
More specifically, the module may, for example, be implemented as a spectral weighting of STFT coefficients with real-valued weights. where X1(n,k), X2(n,k) are the STFT coefficients of the first and second channel, respectively, of the input signal, Y1(n,k), Y2(n,k) are the STFT coefficients of the output signal, G1(n,k), G2(n,k) are the spectral weights, n is the frequency bin index and k is the time index. The approach may hence be performed for all frequency bins, e.g. 1, ..., n.
-
Spectral weights G1(n,k), G2(n,k) may, for example, be computed as function of inter-channel level differences (ICLD), i.e. differences of levels of both channels in each time-frequency bin. Spectral weights G1(n,k), G2(n,k) may, for example, be computed such that ICLD of output is larger than ICLD of input for all coefficient pairs with non-zero ICLD. Spectral weights G1(n,k), G2(n,k) may, for example, be computed such that total level in each time-frequency bin of output equals the total level of input.
-
We compute power per channel where X* denotes the complex conjugate of a complex-valued STFT coefficient X.
-
Referring to Fig. 15, module 1500 may show a schematic view of a block diagram of the spectral weighting. Hence, the computation of the spectral weights may be performed using weight computation units 1520 using the above and following formulas.
-
The ICLD of the input may, for example, be obtained as
-
The spectral weights may, for example, be obtained as otherwise 0, otherwise 0, where α is a user parameter that modifies the intensity of the processing, and finally to ensure that the total energy of the output equals the total energy of the input the spectral weights may, for example, be scaled with the ratio of input power to output power (e.g. before level correction)
-
In contrast to conventional methods, this processing is very efficient because it does not require computing intermediate steps.
-
Referring again to Fig. 15, a module and respectively method for transient processing according to embodiments is disclosed. Hence, the structure as shown in Fig. 15 may be used for boosting inter-channel level differences as well as transient processing. Module 1500 may correspond to transient enhancement 800, with additional filterbank 1510 for time to frequency transformation. Hence, the shown in Fig. 8 may correspond to the computation shown in Fig. 15. Accordingly, module 1500 may correspond to transient suppression 1110, 1220, with additional filterbank 1510 for time to frequency transformation. Depending on a parametrization of module 1500 transient boost or transient suppression may be achieved.
-
In summary or for example as an abstract for the following, according to a method according to embodiments for manipulating the transient signal portions in an audio signal, the input signal may be assumed to be an additive mixture of a sustained signal and a transient signal. The spectral envelope of the sustained signal may, for example, be estimated first, and the spectral envelope of the transient signal may, for example, be obtained using the first result and the signal model. Spectral weights may, for example, be computed from the estimated spectral envelopes of the sustained signal and the transient signal, for example, by means of Wiener filtering and/or a generalized spectral weighting method, such that the sustained signal or the transient signal are attenuated or amplified.
-
The following section may be split in subsections Aim (1), Algorithmic Description (2) Signal model (2.1), Overview (2.2), Signal estimation (2.3), Spectral weighting (2.4), Generalized spectral weighting (2.4.1).
-
Aim (1) of at least some embodiments for transient processing: The module may, for example, attenuate or amplify transient signal components. Transient signal components (or transients, in short) are, for example, the attack part of a note, a percussive hit, a clicking sound or the first 100 ms of a gun shot recording.
-
Why is this useful? The attenuation of transients may, for example, be applied for processing surround signals, artificial reverberation and decorrelation and/or stereophonic enhancement.
-
It may help to keep the localization of sound in the front, may prevent artifacts (for example because transients are critical signals for reverberation and decorrelation), and can be desired for artistic reasons.
-
In contrast, the amplification of transients can be desired when it is mixed with a surround signal in order to maintain the sonic characteristic of the input signal (e.g. punch) or as a tool for sound design.
-
The counterpart of the transient may be a stationary or steady-state and/or sustained signal component. Sustained signal components may, for example, relate to sound sources with slowly evolving temporal envelopes.
-
Algorithmic description (2) of at least some embodiments for transient processing: The processing works (may hence, according to embodiments, optionally be performed) on frequency bands that are processed (e.g. at least mainly) independently of each other. Therefore, a filterbank (or, alternatively, frequency transform) may be used or even required, for example, together with its inverse processing for computing the broad-band output signal (or time signal). More specifically, the module may, for example, be implemented as a spectral weighting processing, e.g. with real-valued weights. The block diagram of Fig. 15 illustrates an example for the spectral weighting.
-
In principle, two basic approaches to transient processing are feasible and may hence be performed according to embodiments. One is to detect first the occurrence of a transient sound and then to manipulate it. The second is to replace the binary detection of a transient by an estimation of the "transientness" or transient presence probability (TPP) which is then used to control the subsequent processing. If, for example, in the simplest case the processing is a scaling of the signal which is inversely proportional to the TPP, transients can be suppressed. The second principle is focused on here.
-
Signal model (2.1) according to embodiments:
The input time domain signal x(t) may, for example, be assumed to be an additive mixture of a sustained signal s(t) and a transient signal t(t).
-
The input signal may, for example, be processed in the frequency domain, for example by using the STFT or another transform or a filterbank, leading to sub-band signals for each frequency band. The matrix X(n, k) may, for example, be represent the sub-band signals of the input signal xm or x(t), with subsampled time index k and frequency band index n. In the following the matrices of sub-band signals are denoted by uppercase letters and are represented by real-valued non-negative quantities, i.e. in case of complex-valued sub-band signals only their magnitudes respectively powers may, for example, be modified whereas their phases may, for example, not be modified.
-
Overview (2.2) according to embodiments: The transient signals and/or sustained signals may, for example, be modified by means of a spectral weighting method, for example each sub-band signal may, for example, be multiplied with a gain factor G(n,k).
-
The output time signal y(t) (e.g. also referred to as ym) may, for example, be obtained from the sub-band signals Y(n, k) by using the inverse processing corresponding to the STFT or transform or filterbank used for computing the input sub-band signals.
-
The spectral weights G(n,k) may, for example, be computed using estimates of the sub-band signals of sustained signal S(n,k) and transient signal T(n,k) (e.g. using weigh computation units 1520).
-
Signal estimation according to embodiments (2.3): The sustained signal envelopes Ŝ(n,k) may, for example, be estimated by means of low-pass filtering the sub-band time trajectories of magnitudes of X(n,k) and optionally applying the constraint Ŝ(n,k) <= |X(n,k)|. The transient signal envelopes T̂(n,k) can, for example, then be estimated by taking the signal model in Equation (1) into account.
-
Spectral weighting (2.4) according to embodiments: Generalized spectral weighting (2.4.1): The spectral weights may, for example, be determined as:
-
This gaining rule is based on a spectral subtraction and Wiener filtering by choosing α and β accordingly. The parameters γ and δ may control the characteristics of the attenuation, 0 <= γ <= 1 and 0 <= δ <= 1. Reducing γ may, for example attenuate the sustained signal components and reducing δ may, for example attenuate the transient signal components.
-
Reference is made to Fig. 16. Fig. 16 shows a schematic view of a weight calculation unit according to embodiments of the invention. As an example, weight calculation unit 1600 is shown for transient suppression and/or transient enhancement, and may hence be an example of weight computation units 1520.
-
Weight calculation unit 1600 comprises a sustained signal estimator 1610, which may be configured to receive an input signal 1601, in order to determine, e.g. as explained above, based on a filtering, a sustained signal component 1611 (e.g. an estimate thereof) of the input signal.
-
Based on a subtraction 1620 of the input signal and the sustained signal component 1611 a transient signal component 1621 of the input signal (e.g. an estimate thereof) may be determined. Input signal 1601, sustained signal component 1611 and transient signal component 1621 are provided to the weight calculator 1630 in order to determine the weights for the transient enhancement or suppression, e.g. according to eqn. (4).
-
Reference is made to Fig. 17. Fig. 17 shows a schematic block diagram of a method for obtaining a processed audio signal on the basis of an input audio signal according to embodiments of the invention. Method (1700) comprises applying (1710) a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal, applying (1720) a stage width enhancement, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene relative to audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, and/or which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene and applying (1730) an ambience enhancement, which provides decorrelated audio signal components on the basis of the input audio signal or on the basis of a processed version of the input audio signal, in order to obtain the processed audio signal.
Further embodiments
-
A first further embodiment comprises an apparatus for processing an audio signal, e.g. an ambience enhancement block, wherein the apparatus is configured to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
-
The apparatus is further configured to obtain, on the basis of the input audio signal, a second modified audio signal using an ambience extraction (which, for example, prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation; and/or which, for example, provides the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal).
-
The apparatus is further configured to combine, e.g. using a weighted combination, the first modified audio signal, or a post-processed version of the first modified audio signal, and the second modified audio signal, or a post-processed version of the second modified audio signal, in order to obtain a combined audio signal; and the apparatus is configured to apply a decorrelation to the combined audio signal, in order to obtain a processed audio signal, e.g. an ambience-enhanced signal.
-
A second further embodiment comprises the apparatus according to the first further embodiment, wherein the apparatus, or for example the ambience extraction, is configured to extract from the input audio signal ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation (e.g. left-right correlation; e.g. an inter-channel correlation which is below a predetermined threshold value, or an inter-channel correlation which is smaller than an inter-channel correlation of direct signal components; e.g. such that direct signal components (and also non-transient direct signal components) are at least partially suppressed in the second modified audio signal, while ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation are non-suppressed or enhanced in the second modified audio signal).
-
Alternatively or in addition, the apparatus, or for example the ambience extraction, is configured to emphasize ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation over, e.g. relative to, direct signal components and/or signal components having a comparatively high, or for example higher, inter-channel correlation.
-
Alternatively or in addition, the apparatus, or for example the ambience extraction, is configured to attenuate direct signal components and/or signal components having a comparatively high inter-channel correlation relative to ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation.
-
A third further embodiment comprises the apparatus according to the first or second further embodiment, wherein the apparatus is configured to compute a covariance matrix on the basis of a frequency domain representation of the input audio signal, e.g. on the basis of a stereo input signal in a frequency domain, e.g. on the basis of X(f,k) or X(f,m) or X(n,k) (which may be used herein, in general optionally interchangeably). Furthermore the apparatus according to the third further embodiment is configured to compute a multichannel parametric Wiener filter on the basis of the covariance matrix, to estimate a direct signal using the multichannel parametric Wiener filter, and to subtract the estimated direct signal from the input audio signal, in order to obtain the second modified audio signal, e.g. Y(f,k) or Y(f,m)or Y(n,k), (e.g. to thereby . obtain the second modified audio signal such that signal components of the input audio signal having a comparatively higher inter-channel correlation are reduced when compared to signal components having a comparatively lower inter-channel correlation).
-
A fourth further embodiment comprises the apparatus according to one of the first to third further embodiments, wherein the apparatus is configured to apply a transient suppression to the second modified audio signal, in order to obtain a post-processed version of the second modified audio signal for combination with the first modified audio signal.
-
A fifth further embodiment comprises the apparatus according to one of the first to fourth further embodiments, wherein the apparatus is configured to obtain the first modified audio signal using a spectral weighting process with real weights.
-
A sixth further embodiment comprises the apparatus according to one of the first to fourth further embodiments, wherein the apparatus is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate, of a, non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g. X(f,m) for a given frequency index f and for a plurality time indices m) of the input audio signal (wherein subband signals may be signals associated with a single frequency bin, or signals associated with a frequency band comprising a plurality of frequency bins). Furthermore, the apparatus according to the sixth further embodiment is further configured to obtain the first modified audio signal in dependence on the estimate of the sustained signal component.
-
A seventh further embodiment comprises the apparatus according to one of the first to sixth further embodiments, wherein the apparatus is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component (e.g. T̂(f,m) or T̂(n,k); e.g. an estimate of a temporal evolution of the transient signal component; e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g. X(f,m) or X(n,k) for a given frequency index f or n and for a plurality time indices m or k) of the input audio signal, and the apparatus is configured to obtain the first modified audio signal in dependence on the estimate of the transient signal component.
-
An eighth further embodiment comprises the apparatus according to one of the first to seventh further embodiments, wherein the apparatus is configured to combine a plurality of frequency bins of a time-frequency-domain representation of the input audio signal, in order to obtain frequency band values associated with a plurality of frequency bands; (wherein the frequency band values are optionally examples of the subband values), and wherein the apparatus is configured to obtain the estimate of the sustained signal component and/or the estimate of the transient signal component in dependence on the frequency band values.
-
A ninth further embodiment comprises the apparatus according to one of the first to eighth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. Gg(f,m) or Gg(n,k), associated with frequency bins, e.g. generally speaking with frequency subbands, on the basis of the frequency band values (e.g. using frequency-band-wise estimates of sustained signal components and/or of transient signal components derived from the frequency band values).
-
A tenth further embodiment comprises the apparatus according to one of the first to ninth further embodiments, wherein the apparatus is configured to obtain temporally smoothened (e.g. low-pass filtered subband signals and/or subband signals purged from peaks) subband signals or temporally smoothened frequency-band signals (e.g. using a combination of frequency bins (frequency subbands) to frequency bands) or temporally smoothened subband signal envelopes or temporally smoothened frequency band signal envelopes (e.g. low-pass filtered subband signal envelopes and/or subband signal envelopes purged from peaks; e.g. low-pass filtered frequency-band signal envelopes or frequency band signal envelopes purged from peaks) on the basis of subband signals representing the input audio signal.
-
As an example, frequency-band signals may be a special example of subband signals, wherein the term "subband signals" comprises both signals associated with a single frequency bins and signals associated with a frequency band comprising a plurality of frequency bins.
-
Furthermore, the apparatus according to the tenth further embodiments is configured to obtain the first modified audio signal in dependence on the temporally smoothened (e.g. low-pass filtered and/or purged from peaks) subband signals or the temporally smoothened frequency band signals or the temporally smoothened subband signal envelopes or the temporally smoothened frequency band signal envelopes.
-
An eleventh further embodiment comprises the apparatus according to one of the first to tenth further embodiments, wherein the apparatus is configured to obtain sustained signal envelopes, e.g. Ŝ(f,m) or Ŝ(n,k), using a temporal smoothing, e.g. low pass filtering, of subband time trajectories or frequency band time trajectories (wherein frequency band time trajectories are a special case of subband time trajectories, wherein the term "subband time trajectories" comprises time trajectories comprising only a single frequency bin and time trajectories comprising a plurality of frequency bins) of the input audio signal (e.g. of a subband representation X(f,m) or X(n,k) of the input audio signal) and using an application of a constraint (e.g. Ŝ(f,m)<X(f,m) or Ŝ(n,k)<=magnitude(X(n,k))) requiring that respective sustained signal envelopes (e.g. Ŝ(f,m) or S^(n,k)) do not exceed magnitudes of respective, for example associated, subband signals, e.g. |X(f,m)|, or magnitudes of respective frequency band signals of the input audio signal or subband time trajectories of magnitudes of the input audio signal or frequency band time trajectories of magnitudes of the input audio signal.
-
As an example, the subband time trajectories may, for example, represent a temporal evolution of energies in respective subbands of the input audio signal, wherein subbands may, for example, be frequency bands of a filterbank or may, for example, be frequency bins of a time-domain-to-spectral-domain transform or of a time-domain-to-frequency-domain transform.
-
A twelfth further embodiment comprises the apparatus according to one of the first to eleventh further embodiments, wherein the apparatus is configured to obtain transient signal envelopes, e.g. T̂(f,m) or T̂(n,k), using a subtraction of sustained signal envelopes, e.g. Ŝ(f,m) or Ŝ(n,k), from respective subband signals, e.g. X(f,m), or from magnitudes or respective subband signals or from frequency band signals of the input audio signal or from magnitudes of frequency band signals of the input audio signal or from or the subband time trajectories, e.g. of magnitudes, of the input audio signal or from frequency band time trajectories, e.g. of magnitudes, of the input audio signal.
-
A thirteenth further embodiment comprises the apparatus according to one of the first to twelfth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. Gg(f,m)or Gg(n,k), in dependence on sustained signal envelope values (e.g. in dependence on the sustained signal envelope values mentioned before, e.g. Ŝ(n,k)).
-
A thirteenth further embodiment comprises the apparatus according to one of the first to twelfth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. Gg(f,m) or Gg(n,k), in dependence on transient signal envelope values (e.g. in dependence on the transient signal envelope values mentioned before, e.g. T̂(n,k)).
-
A fourteenth further embodiment comprises the apparatus according to one of the first to thirteenth further embodiments, wherein the apparatus is configured to obtain respective spectral weights, e.g. Gg(f,m), in dependence on a ratio between a weighted sum of respective sustained signal envelope values, e.g. Ŝ(f,m) or Ŝ(n,k), or an exponentiated version thereof, e.g. Ŝ(f,m)alpha or Ŝ(n,k)alpha, and of respective transient signal envelope values, e.g. T̂(f,m) or T̂(n,k), or an exponentiated version thereof, e.g. T̂(f,m)alpha or T̂(n,k)alpha, and magnitudes of respective subband signals, e.g. X(f,m) or X(n,k), or frequency band signals, or an exponentiated version thereof, e.g. |X(f,m)|alphaor |X(n,k)|alpha.
-
A fifteenth further embodiment comprises the apparatus according to one of the first to fourteenth further embodiments, wherein the apparatus is configured to obtain spectral weights Gg(n,k) according to wherein Ŝ(n,k) are sustained signal envelope values, wherein T̂(n,k) are transient signal envelope values, wherein X(f,m) are subband signal values of the input audio signal, wherein |X(f,m)| are magnitudes of subband signal values of the input audio signal, wherein n is a frequency index, wherein k is a time index, wherein αis a parameter, wherein β is a parameter, wherein γ is a parameter between 0 and 1, and for example preferably larger than or equal to 0.05, and wherein δ is a parameter between 0 and 1, and for example preferably larger than or equal to 0.05.
-
A sixteenth further embodiment comprises the apparatus according to one of the first to fifteenth further embodiments, wherein the apparatus is configured to obtain the sustained signal envelopes Ŝ(n,k) using a low pass filtering of subband time trajectories of magnitudes of a time-frequency-domain representation, e.g. X(f,m) or X(n,k), of the input audio signal, and using an application of a constraint Ŝ(n,k) <|X(n,k)| or of a constraint Ŝ(n,k) <=|X(n,k)|, and wherein the apparatus is configured to obtain the transient signal envelopes T̂(n,k) according to
-
A first additional embodiment comprises an transient signal processor, e.g. a transient enhancement or a transient suppression, wherein the transient signal processor is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate, of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g. X(f,m) for a given frequency index f and for a plurality time indices m) of the input audio signal (wherein subband signals may be signals associated with a single frequency bin, or signals associated with a frequency band comprising a plurality of frequency bins). Furthermore, the transient signal processor is configured to obtain the modified audio signal, in which transient signal components are enhanced or reduced, in dependence on the estimate of the sustained signal component.
-
A second additional embodiment comprises the transient signal processor according to one of the first or second additional embodiments, wherein the transient signal processor is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T̂(f,m) or T̂(n,k), e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component, on the basis of one or more subband signals, (e.g. X(f,m) or X(n,k) for a given frequency index f or n and for a plurality time indices m or k) of the input audio signal, wherein the transient signal processor is configured to obtain the modified audio signal in dependence on the estimate of the transient signal component.
-
Another embodiment according to the invention comprises an inter-channel level difference modifier, wherein the inter-channel level difference modifier is configured to determine inter-channel-level differences (e.g. inter-channel level difference values Dx(n,k); wherein is should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal). The inter-channel level difference modifier is configured to scale spectral bin values, e.g. X1(n,k), X2(n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of a modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g. Y1(n,k), Y2(n,k), is modified when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X1(n,k), X2(n,k), associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
-
Furthermore the inter-channel level difference modifiers is configured to map (e.g. directly, e.g. without a computation of a panning index or the like) a value representing the inter-channel level difference, e.g. Dx(n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain, using a linear or piecewise-linear mapping.
-
This may, for example, be performed such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign. Optionally the above may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
-
Another embodiment according to the invention comprises a consumer-side audio signal processor, for processing an input audio signal for a playback to a listener, wherein the audio signal processor comprises a transient enhancement configured to emphasize transients in an input audio signal, in order to obtain a processed audio signal for playback to the listener.
Implementation alternatives:
-
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
-
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
-
Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
-
Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
-
Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
-
In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
-
A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary.
-
A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
-
A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
-
A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
-
A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
-
In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
-
The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
-
The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and/or in software.
-
The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
-
The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and/or by software.
-
The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.