EP4732276A1 - Enhancing audio content - Google Patents

Enhancing audio content

Info

Publication number
EP4732276A1
EP4732276A1 EP24739951.2A EP24739951A EP4732276A1 EP 4732276 A1 EP4732276 A1 EP 4732276A1 EP 24739951 A EP24739951 A EP 24739951A EP 4732276 A1 EP4732276 A1 EP 4732276A1
Authority
EP
European Patent Office
Prior art keywords
denoising
feature extraction
audio
trained
audio segment
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24739951.2A
Other languages
German (de)
French (fr)
Inventor
Jundai SUN
Zhiwei Shuang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Dolby Laboratories Licensing Corp
Original Assignee
Dolby Laboratories Licensing Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Dolby Laboratories Licensing Corp filed Critical Dolby Laboratories Licensing Corp
Publication of EP4732276A1 publication Critical patent/EP4732276A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/044Recurrent networks, e.g. Hopfield networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Auxiliary Devices For Music (AREA)
  • Soundproofing, Sound Blocking, And Sound Damping (AREA)

Abstract

Techniques for enhancing audio signals are provided herein. In some embodiments, a method for enhancing audio segments may involve receiving an input audio segment to be enhanced. The method may further involve generating features indicative of a type of audio content of the input audio segment. The method may further involve generating features indicative of noise present in the input audio segment by providing at least the features indicative of the type of audio content to a trained denoising feature extraction model. The method may further involve generating a denoising mask based on the features indicative of noise present in the input audio segment. The method may further involve generating an enhanced audio segment at least by applying the denoising mask to the input audio segment.

Description

ENHANCING AUDIO CONTENT
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from PCT Application No. PCT/CN2023/102113 filed on 25 June 2023 and U.S. Provisional Application. No. 63/625,184 filed on 25 January 2024, each of which is incorporated by reference herein in its entirety.
TECHNICAL FIELD
[0002] This disclosure pertains to systems, methods, and media for enhancing audio content.
BACKGROUND
[0003] Enhancement of audio content may be performed by audio content creators, who may capture audio content that includes speech and/or music, as well as noise. The noise may include environmental noise (e.g., birds chirping, wind, etc.), background talkers, etc. Removal of such noise may be desirable, e.g., to enhance speech content or other primary content in an audio signal. However, removing noise, particularly while not degrading the primary content of the audio signal, may be difficult.
NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon). [0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and/or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.
SUMMARY
[0008] In accordance with some embodiments, techniques for enhancing audio signals are provided herein. In some embodiments, a method for enhancing audio signals may involve receiving an input audio segment to be enhanced. The method may further involve generating features indicative of a ty pe of audio content of the input audio segment. The method may further involve generating features indicative of noise present in the input audio segment by providing at least the features indicative of the type of audio content to a trained denoising feature extraction model. The method may further involve generating a denoising mask based on the features indicative of noise present in the input audio segment. The method may further involve generating an enhanced audio segment at least by applying the denoising mask to the input audio segment.
[0009] In some examples, the method may further involve obtaining initial features associated with the input audio segment by providing the input audio segment to a trained initial feature extraction model, wherein the features indicative of the ty pe of audio content of the input audio segment are generated using the initial features. In some examples, the trained denoising feature extraction model receives the initial features as an input.
[0010] In some examples, the features indicative of the ty pe of audio content are extracted by a trained discriminative feature model. In some examples, the trained discriminative feature model and the trained denoising feature extraction model are trained in a multi-stage training process that comprises training both the discriminative feature model and the denoising feature extraction model concurrently in a first stage, freezing weights of the denoising feature extraction model while fine-tuning weights of the discriminative feature model in a second stage, and freezing weights of the discriminative feature model while fine-tuning weights of the denoising feature extraction model in a third stage. In some examples, at least the trained discriminative feature model is trained using a training set that includes audio content samples labeled as including or not including music content, wherein a label is automatically generated based on an energy-based rule. In some examples, the discriminative feature model and the denoising feature extraction model are trained using training samples that include at least speech in noise audio segments and music in noise audio segments, and wherein the lowest signal to noise ratio (SNR) for the speech in noise audio segments is lower than the lowest SNR for the music in noise audio segments. In some examples, the trained discriminative feature model comprises a gated recurrent unit (GRU).
[0011] In some examples, the trained denoising feature extraction model comprises a gated recurrent unit (GRU).
[0012] In some examples, the features indicative of the type of audio content comprise an indication of whether the input audio segment includes music content. In some examples, the method further involves smoothing indications of whether the input audio segment includes music content over multiple frames of the input audio segment.
[0013] In some examples, the method further involves determining a degree to which harmonics of the input audio segment are to be enhanced in the enhanced audio segment by providing at least the features indicative of the type of audio content to a trained pitch strength feature extraction block. In some examples, the degree to which the harmonics of the input audio segment are enhanced is inversely proportional to a confidence that the input audio segment includes music content.
[0014] In some examples, the denoising mask is generated such that an aggressiveness of denoising is inversely proportional to a confidence that the input audio segment includes music content.
[0015] In some examples, at least the denoising feature extraction model is trained using a training set that includes labeled ambient sounds. [0016] Some or all of the operations, functions and/or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.
[0017] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0018] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.
BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a diagram that illustrates an example system for enhancing audio content based on content type in accordance with some embodiments.
[0020] Figure 2 shows an example implementation of the system shown in Figure 1 in accordance with some embodiments.
[0021] Figure 3 shows an example implementation of a system for enhancing audio content that includes pitch strengthening based on content type in accordance with some embodiments.
[0022] Figure 4 is a plot that relates gain parameters for noise attenuation based on a likelihood that audio content includes music in accordance with some embodiments. [0023] Figure 5 is another example plot that relates gain parameters for noise attenuation based on a likelihood that audio content includes music in accordance with some embodiments.
[0024] Figure 6 is an example of a plot that relates pitch strengthening parameters based on a likelihood that audio content includes music in accordance with some embodiments.
[0025] Figure 7 is a flowchart of an example process for using a trained model for performing audio content enhancement in accordance with some embodiments.
[0026] Figure 8 is a flowchart of an example process for training a model using a multi-stage process in accordance with some embodiments.
[0027] Figure 9 is a flowchart of an example process for using a trained model for performing audio content enhancement in accordance with some embodiments.
[0028] Figure 10 shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0029] Like reference numbers and designations in the various drawings indicate like elements.
DETAILED DESCRIPTION OF EMBODIMENTS
[0030] Enhancement of audio content may be performed by audio content creators, who may capture audio content that includes speech and/or music, as well as noise. The noise may include environmental noise (e.g., birds chirping, wind, etc ), background talkers, etc. Removal of such noise may be desirable, e.g., to enhance speech content or other primary content of interest in an audio signal. However, removing noise, particularly while not degrading the primary content of the audio signal, may be difficult. For example, when noise is attenuated, environmental sounds and music may also be attenuated, which may alter the audio signal in an undesirable manner.
[0031] Disclosed herein are techniques for enhancing audio signals. For example, an enhanced audio signal may have noise attenuated in a manner that considers the type of audio content in the audio signal. As a more particular example, noise may be attenuated more strongly for audio signals that do not include music or have a low likelihood of including music. As another example, an enhanced audio signal may have enhanced pitch strength, sometimes referred to as enhanced harmonics. The harmonics may be enhanced in a manner that is dependent on the likelihood the audio signal includes music. For example, harmonics may be more strongly enhanced for audio signals less likely to include music, and harmonic enhancement may be generally inhibited for audio signals likely to include music.
[0032] Enhancement of audio signals based on audio content type may be implemented by generating a denoising mask and/or pitch strength metric based on features of the audio signal indicative of content type. For example, a discriminative feature extraction model may extract features usable to determine a likelihood the audio signal includes music. These features may be provided to a trained model (e.g., a trained denoise feature extraction model) configured to generate features indicative of noise, which may in turn be used to generate a denoising mask. Similarly, the features may be provided to a trained model configured to generate a pitch strength metric that indicates a degree to which harmonics are to be enhanced. By utilizing the features indicative of audio content type to generate a denoising mask and/or pitch strength metric, enhancement of the audio signal may in turn be dependent on audio content type such that enhancement of the audio signal differs for audio signals likely to include music compared to audio signals unlikely to include music. Additionally, in some embodiments, a denoising gain and/or a pitch strength gain may be modified based on a likelihood that the audio signal includes music, thereby providing an additional avenue by which enhancement of audio signals may be content type dependent. By enhancing audio signals based on content type, noise of particular types, e g., “annoying noise” may be attenuated or removed while allowing some types of environmental sounds (e.g., birds chirping) to remain to provide context. Additionally, by enhancing audio signals based on content type, e.g., whether the audio signal includes music or not (or a probability the audio signal includes music), audio signals that includes music may be enhanced in a manner the leaves the music generally unaltered.
[0033] Figure 1 illustrates an example system usable for enhancing audio signals in accordance with some embodiments. As illustrated, a shared feature extraction block 102 may be configured to receive, as an input, an audio signal or a representation of an audio signal (e.g., a frequency domain representation of an audio signal). Shared feature extraction block 102 may be configured to generate, or extract, features associated with the input audio signal. In some embodiments, the extracted features may be considered coarse features usable by multiple other downstream models or feature extraction blocks.
[0034] The extracted features may be provided to both a discriminative feature extraction block 104 and a denoise feature extraction block 106. Discrimination feature extraction block 104 may be configured to generate, or extract, features that may be usable to determine content type information associated with the input audio signal. For example, the features may be usable to determine a likelihood or a confidence level that the audio signal includes music. Denoise feature extraction block 106 may be configured to generate, or extract, features usable to generate a denoising mask, which, when applied to the input audio signal (or a frequency domain representation of the audio signal), enhances the audio signal by removing or attenuating noise. Note that, as shown in Figure 1, denoise feature extraction block 106 is configured to receive, as input, features generated by discriminative feature extraction block 104. By considering features usable to determine content type information, the denoising features may be dependent on content type. For example, the denoising features may allow a different denoising mask to be generated for audio signals that include music relative to those that do not, e.g., to less aggressively attenuating noise in audio signals that includes music relative to audio signals that do not include music.
[0035] As illustrated in Figure 1, features generated by discriminative feature extraction block 104 may be provided to classification layer 108, which may be configured to generate a music flag. The music flag may indicate whether or not the audio signal includes music. In some embodiments, the music flag may be a Boolean value that indicates a prediction of whether or not the audio signal includes music. Alternatively, in some embodiments, the music flag may be a continuous value (e.g., from 0 to 1) that indicates a probability7 the audio signal includes music.
[0036] As illustrated in Figure 1. features generated by denoise feature extraction block 106 may be provided to denoising block 110, which is configured to generate a denoising mask based on the denoising features. As described above, because the denoising features are generated based on features indicative of content ty pe, the denoising mask is also generated based on content type, and in particular, may be generated based on a likelihood the audio signal includes music.
[0037] In some embodiments, a feature extraction block (e.g., a shared feature extraction block, a discriminative feature extraction block, and/or a denoise feature extraction block) may be implemented using one or more trained machine learning model architectures. The various feature extraction blocks may utilize the same type of architecture or different architectures. In some embodiments, a feature extraction block may utilize a gated recurrent unit (GRU) architecture. A GRU architecture may include a gating mechanism that allows particular features or parameters to be utilized or not utilized (e g., forgotten). In some embodiments, other architectures such as a long short-term memory (LSTM) netw ork, a transformer network, or the like may be utilized. [0038] Figure 2 illustrates an example implementation of the system shown in Figure 1 that utilizes GRU networks to implement discrimination feature extraction block 104 and denoise feature extraction block 106. Note that, at 202. output from discrimination feature extraction block 104 is provided as input to denoise feature extraction block 106 at an initial concatenation layer of denoise feature extraction block 106 that allows the discrimination features to be concatenated with shared extracted features which are used by denoise feature extraction block 106.
[0039] In some embodiments, enhancing the audio signal may involve enhancing harmonic structures. Harmonic structure enhancement may be implemented using a comb filter. Note that the degree of comb filtering may affect the perception of the enhanced audio signal; for example, not enough filtering may result in the enhanced audio signal sounding rough, whereas too much filtering may result in the enhanced audio signal sounding robotic. Accordingly, the degree of filtering to be applied to enhance harmonic structures may be determined and represented by a pitch strength metric, which may be a value which indicates a degree to which comb filtering is to be applied. In some embodiments, the pitch strength may be determined based on an output of a strength feature extraction block. In some embodiments, the strength feature extraction block maytake, as input, features generated or extracted by a discriminative feature extraction block that may be used to determine a type of content associated with the input signal. In this way, the pitch strength metric that determines a degree to which harmonic enhancement is applied, may in turn be dependent on the content ty pe of the audio signal. For example, harmonic enhancement maybe inhibited or applied to a lower degree for audio signals with a high likelihood of containing music.
[0040] Figure 3 illustrates an example implementation that utilizes a pitch strength feature extraction block 302 and a strength layer 304 to generate a pitch strength metric. As described above, the pitch strength metric may be a value (e.g., between 0 and 1) that indicates a degree to which harmonic enhancement (e.g., comb filtering) is to be applied to the audio signal. Note that, at 308, shared features (e.g., coarse features) extracted by the shared feature extraction block maybe provided to pitch strength feature extraction block 302, and, at 310. discriminative features extracted by discriminative feature extraction block 104 may be provided to strength feature extraction block 302. Use of the discriminative features may allow the pitch strength metric to be determined based on the content ty pe of the audio signal. For example, in some embodiments, the pitch strength metric may be determined based on a likelihood the audio signal includes music. [0041] As described above, the features extracted by a discriminative feature extraction block may be used to generate a music flag. The music flag may indicate a likelihood that the audio signal includes music content. For example, in some embodiments, the music flag may be a continuous value that indicates a probability that the audio signal includes music. In some embodiments, the music flag may be used to modify a denoising gain and/or to modify a pitch strengthening gain such that gains are applied based on the likelihood that the content includes music. For example, denoising may be attenuated for content that is likely to include music to avoid altering the music content. As another example, pitch strengthening (e.g.. harmonic enhancement) may be attenuated or inhibited for content that is likely to include music to avoid altering the harmonic structure of the music.
[0042] In some embodiments, an initial denoising gain may be modified based on the likelihood that the audio signal includes music. For example, the initial denoising gain may be modified using a power-law function such that the initial denoising gain is raised to a power that is dependent on the likelihood that the audio signal includes music. In some embodiments, the modified denoising gain may be determined by: gainmodified = gain
[0043] In the equation given above, "‘gain7’ represents an initial value of a denoising gain and y is a metric based on the likelihood the audio signal includes music. For example, in some embodiments, y may be determined by: y = a{-{-musicfiag-musicTH)')
[0044] In the equation given above, a is a scalar value. In some embodiments, a may be 10, 15, 20, 25, etc. As described above, “musicfiag" is a value that indicates a likelihood the audio signal includes music, and may be a continuous value between 0 and 1. In some embodiments, “musicra” is a threshold value that is used to distinguish between speech and music. Example values include 0.5, 0.7, etc. Note that the value of musiciH may be tuned based on a false-alarm rate. In some embodiments, y may be limited to have a maximum value of 1.
[0045] Figure 4 illustrates an example graph of y as a function of the music flag value. Note that for values of the music flag less than about 0.4, y is 1. Accordingly, the initial denoising gain is not modified. However, with increasing confidence that the audio signal includes music (e.g., for values of the music flag greater than 0.4), y decreases in value. Lower values of y cause the initial denoising gain to be reduced due to the increasing likelihood that the audio signal includes music content.
[0046] Figure 5 depicts a two-dimensional plot that illustrates processed gain as a function of initial gain and music flag (or confidence that the audio signal includes music content). Note that for 0 confidence that the audio signal includes music content, the initial denoising gain is not modified and remains the same. However, for increasing confidence that the audio signal includes music content, the initial denoising gain is increasingly modified such that the gain is reduced, thereby attenuating the denoising effect. In other words, denoising aggressiveness may be inversely proportional to the likelihood the audio signal includes music.
[0047] In some embodiments, a pitch strength metric that indicates a degree to which harmonic enhancement is to be performed may be modified based on the music flag. This may allow inhibition of harmonic enhancement on audio signals that are likely to include music content to avoid alteration of the harmonic structure of the music. In some embodiments, an initial pitch strength may be modified by multiplying the initial pitch strength by a parameter 0. The parameter 0 may a function (e.g.. a non-linear function) of the likelihood music flag, or. in other words, a function of the likelihood that the audio signal includes music. In one example, the initial pitch strength may be modified by: pitch strengthmodified = pitch strength * 6
[0048] The parameter 0 may be a non-linear function of the music flag, such as an exponential function. In some embodiments, the parameter 0 may be a based on a polynomial function of the music flag. In one example, 0 may be determined by:
[0049] As described above, music™ is a predetermined threshold value. In the equation given above, k may be an even number greater than or equal to 2 (e.g., 2, 4, 6, etc.).
[0050] Figure 6 is a two-dimensional plot that illustrates modified pitch gain (or modified pitch strength) as a function of initial pitch gain (or initial pitch strength) and confidence or likelihood that the audio signal includes music. Note that, when the audio signal is determined to not include music (e.g., a music confidence value of 0), the initial pitch strength is not modified. However, for higher levels of confidence that the audio signal includes music, the initial pitch strength is reduced, which allows inhibition of harmonic enhancement for audio signals likely to include music. In other words, the degree to which harmonics are enhanced may be inversely proportional to the likelihood the audio signal includes music.
[0051] To avoid discontinuous jumps in enhancement of the audio signal from frame to frame based on the likelihood the frame includes music, the values of the music flag may be smoothed. The smoothed value of the music flag may then be used as described above to determine denoising gain and/or pitch enhancement gain. The music flag may be smoothed using an attack and release rule such that the flag may increase or decrease in value at a rate that is controlled by a release and/or attack factor. For example, in some embodiments, the music flag may be smoothed by:
[0052] In the equation given above, a and P are parameters which may be tuned. In some embodiments, a may be 0.02, 0.05, 0.08, or the like. In some embodiments. may be 0.85, 0.9, 0.95, 0.98, or the like.
[0053] Figure 7 is a flowchart of an example process for enhancing audio signals based on content type in accordance with some embodiments. In some implementations, blocks of process 700 may be performed by one or more control systems or processors of a computing device, such as a server, a laptop computer, a desktop computer, a tablet computer, a mobile phone, etc. An example of such a control system is control system 1010 shown in and described below in connection with Figure 10. In some embodiments, blocks of process 700 may be executed in an order other than what is shown in Figure 7. In some embodiments, two or more blocks of process 700 may be executed substantially in parallel. In some embodiments, one or more blocks of process 700 may be omitted.
[0054] Process 700 can begin at 702 by receiving an input audio segment to be enhanced. The input audio segment may be, e.g., a frame of an audio signal. The input audio segment may correspond to user-captured audio content, e.g., via one or more microphones of a user device such as a mobile phone or tablet computer.
[0055] At 704, process 700 can provide a representation of the input audio segment to a trained feature extraction model configured to extract an initial set of features. The representation of the input audio segment may be a frequency domain representation of the input audio segment. The trained feature extraction model may be the shared feature extraction model shown in and described above in with Figure 1. In some embodiments, the initial set of features may be considered coarse features of the input audio segment that may be used by downstream feature extractors to extract more fine-grained features.
[0056] At 706, the initial set of features may be provided to a trained discriminative feature extraction model configured to extract features indicative of a t pe of audio content. An example of a trained discriminative feature extraction model is discriminative feature extraction block 104 shown in Figure 1. In some embodiments, the trained discriminative feature extraction model may utilize a GRU architecture. The extracted features may be usable to classify the input audio segment as including music, and/or to predict a probability that the input audio segment includes music.
[0057] At 708, the initial set of features and the features indicative of the type of audio content may be provided to a trained denoising feature extraction model configured to extract features indicative of noise. In other words, the features indicative of noise may be generated based at least in part on the features indicative of the type of audio content. An example of a trained denoising feature extraction model is denoise feature extraction block 106 shown in Figure 1. The trained denoising feature extraction model may utilize a GRU architecture.
[0058] At 710, process 700 can optionally generate a prediction of whether the input audio segment includes music content based on the features indicative of the type of audio content. For example, at 710, process 700 can generate a music flag, as described above, which may indicate a probability or a confidence level that the input audio segment includes music. Note that, as described above, in some embodiments, the music flag may be smoothed from frame-to-frame to allow7 stability in processing performed using the music flag (as described above in connection with Figures 4-6).
[0059] At 712, process 700 can generate a denoising mask based on the features indicative of noise. For example, as described above in connection with Figure 1, the features indicative of noise may be provided to a denoising block configured to generate the denoising mask. In some embodiments, the denoising mask may further be generated based on the music flag, as determined above at block 710. For example, in some embodiments, an initial denoising gain may be modified based on the music flag such that the gain is reduced with increasing likelihood of the audio segment containing music. [0060] At 714, process 700 can apply the denoising mask to the representation of the input audio segment. For example, process 700 can multiply the denoising mask by the representation of the input audio segment in the frequency domain to generate a denoised version of the input audio segment. Note that because the denoising mask was generated by taking into consideration the content type of the audio content in the input audio segment, denoising may be performed in a manner that controls denoising aggressiveness based on content type.
[0061] At 716. process 700 can optionally perform harmonic enhancement based on the features indicative of the type of audio content. For example, in some embodiments, process 700 can determine a pitch strength metric that represents a degree to which comb filtering is to be performed based on the features indicative of the type of audio content. In some embodiments, the pitch strength metric may be modified based on a value of the music flag (e.g., as determined at block 710), such that harmonic enhancement is attenuated or inhibited for audio segments likely to include music content.
[0062] In some embodiments, process 700 may loop through blocks 702-716 for multiple frames of an audio signal to generate an enhanced audio signal. In some embodiments, an enhanced audio signal may be stored (e.g., memory of a user device or on a server device) and/or may be played back.
[0063] In some embodiments, a shared feature extraction block, a discriminative feature extraction block and classification layer, and a denoise feature extraction block and denoising layer may be trained using a training set. In some embodiments, the training set may include training samples that include speech in noise audio segments and music in noise audio segments. In some implementations, the lowest SNR for the speech in noise audio segments may be lower than the lowest SNR for the music in noise audio segments. For example, in some embodiments, the lowest SNR for the speech in noise audio segments may be -10 dB, -5 dB, -3 dB, or 0 dB, while the lowest SNR for the music in noise audio segments may be 0 dB, 5 dB, 10 dB, etc. In one example, the lowest SNR for the speech in noise audio segments may be 15 dB lower than the lowest SNR for the music in noise audio segments. In some embodiments, an audio segment included in a training sample may be labeled as including or not including music. In some implementations, a label may be automatically generated based on, e.g., an energy-based rule. In some embodiments, an audio segment in a training set may be labeled as including or not including ambient sounds, e.g., birds chirping, car horns, etc. For example, particular audio objects in the audio segment may be labeled. In some embodiments, at least the denoising feature extraction model or block (and optionally, other models or blocks, such as a shared feature extraction block) may be trained using audio segments labeled as including or not including ambient sounds.
[0064] In some embodiments, a shared feature extraction block, a discriminative feature extraction block and classification layer, and a denoise feature extraction block and denoising layer may be trained in multiple stages. For example, in some embodiments all blocks and layers may be trained concurrently in an initial training stage. Continuing with this example, in some embodiments, a first fine-tuning stage can be performed by freezing weights associated with the shared feature extraction model, the denoising feature extraction block and denoising layer, while continuing to update weights associated with the discriminative feature extraction block and the classification layer. Continuing still further, in a second fine-tuning stage, weights associated with the shared feature extraction model, the discriminative feature extraction block, and the classification layer can be frozen while continuing to train and update the weights associated with the denoising feature extraction block and the denoising layer. Each fine-tuning stage may allow training to be focused on layers of inadequate quality using a new training data set that has yet to be used to train any portion of the models. Note that, in some implementations, weights associated with each layer may be tracked in a dataset, or via metadata associated with that layer, which may allow weights associated with a subset of layers to be frozen while continuing to update weights associated with other layers.
[0065] Figure 8 is a flowchart of an example process for training blocks of an audio signal enhancement system in accordance with some embodiments. In some implementations, blocks of process 800 may be performed by one or more control systems or processors of a computing device, such as a server device or other computing device. An example of such a control system is control system 1010 shown in and described below in connection with Figure 10. In some embodiments, blocks of process 800 may be executed in an order other than what is shown in Figure 8. In some embodiments, two or more blocks of process 800 may be executed substantially in parallel. In some embodiments, one or more blocks of process 800 may be omitted.
[0066] Process 800 can begin at 802 by obtaining a training set, where each training sample comprises an audio segment, a music content indication, and a target denoising mask. The music content indication may be the result of a manual annotation. The target denoising mask may be one that was used to generate the audio segment that includes noise. For example, each audio segment may be one that was generated by applying the target denoising mask to a clean audio segment to generate the audio segment included in the training sample. Note that the audio segments may include a range of signal to noise ratios (SNRs). In some embodiments, the training set may include training samples that include speech in noise audio segments and music in noise audio segments. In some implementations, the lowest SNR for the speech in noise audio segments may be lower than the lowest SNR for the music in noise audio segments. For example, in some embodiments, the lowest SNR for the speech in noise audio segments may be -10 dB, -5 dB, -3 dB, or 0 dB, while the lowest SNR for the music in noise audio segments may be 0 dB, 5 dB, 10 dB, etc. In one example, the lowest SNR for the speech in noise audio segments may be 15 dB lower than the lowest SNR for the music in noise audio segments.
[0067] At 804, process 800 can perform a first stage of training by utilizing a first subset of the training set to train a shared feature extraction block, a discriminative feature extraction block and classification layer, and a denoising feature extraction and denoising layer. During the first stage of training, weights associated with each block and/or layer may be updated.
[0068] At 806, process 800 can perform a second stage of training by freezing weights associated with the shared feature extraction block and weights associated with the denoising feature extraction block and denoising layer while continuing to train to discriminative feature extraction block and classification layer using a second subset of the training set. Note that the second stage of training may continue until a target classification performance is achieved.
[0069] At 808. process 800 can perform a third stage of training by freezing weights associated with the shared feature extraction block and weights associated with the discriminative feature extraction block and the classification layer while training the denoising feature extraction block and the denoising layer using a third subset of the training set. Note that the third stage of training may continue until a target error associated with the generated denoising mask has been achieved.
[0070] Figure 9 is a flowchart of an example process for enhancing audio signals based on content type in accordance with some embodiments. In some implementations, blocks of process 900 may be performed by one or more control systems or processors of a computing device, such as a server, a laptop computer, a desktop computer, a tablet computer, a mobile phone, etc. An example of such a control system is control system 1010 shown in and described below in connection with Figure 10. In some embodiments, blocks of process 900 may be executed in an order other than what is shown in Figure 9. In some embodiments, two or more blocks of process 900 may be executed substantially in parallel. In some embodiments, one or more blocks of process 900 may be omitted. [0071] Process 900 can begin at 902 by receiving an input audio segment to be enhanced. The input audio segment may be, e.g., a frame of an audio signal. The input audio segment may correspond to user-captured audio content, e.g., via one or more microphones of a user device such as a mobile phone or tablet computer.
[0072] At 904, process 900 can generate features indicative of a type of audio content of the input audio segment. For example, the ty pe of audio content may be music, speech, etc. The features indicative of the type of audio content may’ be generated by a discriminative feature extraction model, e.g., as shown in and described above in connection with Figure 1. In some embodiments, the discriminative feature extraction model may take, as input, coarse features, e.g., generated by a shared feature extraction model as shown in and described above in connection with Figure 1. Alternatively, in some embodiments, the discriminative feature extraction model may take, as input, a representation of the audio segment (e.g., a frequency domain representation of the audio segment) or the audio segment itself.
[0073] At 906. process 900 can generate features indicative of noise present in the input audio segment by providing at least the features indicative of the type of audio content to a trained denoising feature extraction model. An example of the trained denoising feature extraction model is denoise feature extraction block 106 of Figure 1. In some embodiments, the trained denoising feature extraction model may additionally take as input coarse features, e.g.. generated by a shared feature extraction model as shown in and described above in connection with Figure 1.
[0074] At 908, process 900 can generate a denoising mask based on the features indicative of noise present in the input audio segment. For example, the output of the trained denoising feature extraction model may be provided to a denoising layer configured to generate the denoising mask based on the features indicative of noise.
[0075] At 910. process 900 can generate an enhanced audio segment at least by applying the denoising mask to the input audio segment. For example, in some embodiments, process 900 can multiply the denoising mask by a frequency domain representation of the input audio segment to generate the enhanced audio segment. Note that, in some embodiments, as described above in connection with Figures 4-6, a denoising gain may be modified based on a music flag that indicates the probability that the input audio segment includes music. Accordingly, denoising may be performed based on content of the audio segment using the features indicative of the type of audio content (e.g., as described above in connection with block 906), and, in some embodiments, additionally based on the music flag by modifying an initial denoising gain based on the music flag.
[0076] Note that, in some embodiments, additional enhancement may be performed on the input audio segment by performing pitch strengthening, or enhancing harmonics of the audio segment. As described above in connection wi th Figures 3 and 6, a pitch strength metric may be determined based on the features indicative of the type of audio content, which may optionally be modified based on a music flag. Accordingly, harmonic enhancement may be performed based on the t pe of audio content such that the degree of harmonic enhancement is dependent on the likelihood the audio segment includes music.
[0077] In some embodiments, blocks 902-910 may be looped through, e.g., for multiple frames of an audio signal. The enhanced audio segments may be stored (e.g., in local memory and/or on a server device or other remote device), and/or played back.
[0078] Figure 10 is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 10 are merely provided by way of example. Other implementations may include more, fewer and/or different types and numbers of elements. According to some examples, the apparatus 1000 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 1000 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0079] According to some alternative implementations the apparatus 1000 may be. or may include, a server. In some such examples, the apparatus 1000 may be, or may include, an encoder. Accordingly, in some instances the apparatus 1000 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 1000 may be a device that is configured for use in "the cloud/’ e.g., a server.
[0080] In this example, the apparatus 1000 includes an interface system 1005 and a control system 1010. The interface system 1005 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 1005 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may. in some examples, pertain to one or more software applications that the apparatus 1000 is executing.
[0081] The interface system 1005 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and/or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0082] The interface system 1005 may include one or more network interfaces and/or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 1005 may include one or more wireless interfaces. The interface system 1005 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and/or a gesture sensor system. In some examples, the interface system 1005 may include one or more interfaces between the control system 1010 and a memory system, such as the optional memory system 1015 shown in Figure 10. However, the control system 1010 may include a memory' system in some instances. The interface system 1005 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0083] The control system 1010 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and/or discrete hardware components.
[0084] In some implementations, the control system 1010 may reside in more than one device. For example, in some implementations a portion of the control system 1010 may reside in a device within one of the environments depicted herein and another portion of the control system 1010 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 1010 may reside in a device within one environment and another portion of the control system 1010 may reside in one or more other devices of the environment. For example, a portion of the control system 1010 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 1010 may reside in another device that is implementing the cloud-based service, such as another server, a memory device, etc. The interface system 1005 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0085] In some implementations, the control system 1010 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 1010 may be configured for training one or more models, generating enhanced audio content using one or more models, or the like.
[0086] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 1015 shown in Figure 10 and/or in the control system 1010. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, be executable by one or more components of a control system such as the control system 1010 of Figure 10.
[0087] In some examples, the apparatus 1000 may include the optional microphone system 920 shown in Figure 10. The optional microphone system 1020 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 1000 may not include a microphone system 1020. However, in some such implementations the apparatus 1000 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 1010. In some such implementations, a cloud-based implementation of the apparatus 1000 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 1010.
[0088] According to some implementations, the apparatus 1000 may include the optional loudspeaker system 1025 shown in Figure 10. The optional loudspeaker system 1025 may include one or more loudspeakers, which also may be referred to herein as "‘speakers” or. more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 1000 may not include a loudspeaker system 1025. In some implementations, the apparatus 1000 may include headphones. Headphones may be connected or coupled to the apparatus 1000 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0089] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and/or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and/or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0090] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g.. programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and/or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and/or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and/or a keyboard), a memory, and a display device.
[0091] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.
[0092] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.

Claims

1. A method for enhancing audio signals, the method comprising: receiving an input audio segment to be enhanced; generating features indicative of a type of audio content of the input audio segment; generating features indicative of noise present in the input audio segment by providing at least the features indicative of the t pe of audio content to a trained denoising feature extraction model; generating a denoising mask based on the features indicative of noise present in the input audio segment; and generating an enhanced audio segment at least by applying the denoising mask to the input audio segment.
2. The method of claim 1, further comprising obtaining initial features associated with the input audio segment by providing the input audio segment to a trained initial feature extraction model, wherein the features indicative of the type of audio content of the input audio segment are generated using the initial features.
3. The method of claim 2, wherein the trained denoising feature extraction model receives the initial features as an input.
4. The method of any one of claims 1-3, wherein the features indicative of the type of audio content are extracted by a trained discriminative feature model.
5. The method of claim 4, wherein the trained discriminative feature model and the trained denoising feature extraction model are trained in a multi-stage training process that comprises training both the discriminative feature model and the denoising feature extraction model concurrently in a first stage, freezing weights of the denoising feature extraction model while fine-tuning weights of the discriminative feature model in a second stage, and freezing weights of the discriminative feature model while finetuning weights of the denoising feature extraction model in a third stage.
6. The method of any one of claims 4 or 5, wherein at least the trained discriminative feature model is trained using a training set that includes audio content samples labeled as including or not including music content, wherein a label is automatically generated based on an energy-based rule.
7. The method of any one of claims 4-6, wherein the discriminative feature model and the denoising feature extraction model are trained using training samples that include at least speech in noise audio segments and music in noise audio segments, and wherein the lowest signal to noise ratio (SNR) for the speech in noise audio segments is lower than the lowest SNR for the music in noise audio segments.
8. The method of any one of claims 4-7, wherein the trained discriminative feature model comprises a gated recunent unit (GRU).
9. The method of any one of claims 1-8, wherein the trained denoising feature extraction model comprises a gated recurrent unit (GRU).
10. The method of any one of claims 1-9, wherein the features indicative of the type of audio content comprise an indication of whether the input audio segment includes music content.
11. The method of claim 10, further comprising smoothing indications of whether the input audio segment includes music content over multiple frames of the input audio segment.
12. The method of any one of claims 1-11, further comprising determining a degree to which harmonics of the input audio segment are to be enhanced in the enhanced audio segment by providing at least the features indicative of the type of audio content to a trained pitch strength feature extraction block.
13. The method of claim 12, wherein the degree to which the harmonics of the input audio segment are enhanced is inversely proportional to a confidence that the input audio segment includes music content.
14. The method of any one of claims 1-13, wherein the denoising mask is generated such that an aggressiveness of denoising is inversely proportional to a confidence that the input audio segment includes music content.
15. The method of any one of claims 1-14, wherein at least the denoising feature extraction model is trained using a training set that includes labeled ambient sounds.
16. A system comprising: one or more processors; and a non-transilory computer-readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of claims 1-15.
17. A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of claims 1-15.
EP24739951.2A 2023-06-25 2024-06-18 Enhancing audio content Pending EP4732276A1 (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
CN2023102113 2023-06-25
US202463625184P 2024-01-25 2024-01-25
PCT/US2024/034472 WO2025006266A1 (en) 2023-06-25 2024-06-18 Enhancing audio content

Publications (1)

Publication Number Publication Date
EP4732276A1 true EP4732276A1 (en) 2026-04-29

Family

ID=91853752

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24739951.2A Pending EP4732276A1 (en) 2023-06-25 2024-06-18 Enhancing audio content

Country Status (3)

Country Link
EP (1) EP4732276A1 (en)
CN (1) CN121713237A (en)
WO (1) WO2025006266A1 (en)

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10937443B2 (en) * 2018-09-04 2021-03-02 Babblelabs Llc Data driven radio enhancement
US11562761B2 (en) * 2020-07-31 2023-01-24 Zoom Video Communications, Inc. Methods and apparatus for enhancing musical sound during a networked conference
CN113539283B (en) * 2020-12-03 2024-07-16 腾讯科技(深圳)有限公司 Audio processing method, device, electronic device and storage medium based on artificial intelligence
CN112908352B (en) * 2021-03-01 2024-04-16 百果园技术(新加坡)有限公司 Audio denoising method and device, electronic equipment and storage medium

Also Published As

Publication number Publication date
WO2025006266A1 (en) 2025-01-02
CN121713237A (en) 2026-03-20

Similar Documents

Publication Publication Date Title
US20240177726A1 (en) Speech enhancement
US12597434B2 (en) Control of speech preservation in speech enhancement
US20160071527A1 (en) Method and System for Scaling Ducking of Speech-Relevant Channels in Multi-Channel Audio
CN108806707B (en) Voice processing method, device, equipment and storage medium
JP2015050685A (en) Audio signal processing apparatus and method, and program
CN118800268B (en) Voice signal processing method, voice signal processing device and storage medium
CN104364845B (en) Processing meanss, processing method, program, computer-readable information recording medium and processing system
JP2026035611A (en) Data Augmentation for Speech Improvement
CN111048118A (en) A voice signal processing method, device and terminal
CN110992975A (en) A voice signal processing method, device and terminal
WO2025006266A1 (en) Enhancing audio content
JP7735537B2 (en) Detecting Environmental Noise in User-Generated Content
JP7826507B2 (en) Method and audio processing system for wind noise suppression
CN114974279A (en) Sound quality control method, device, equipment and storage medium
US20260024541A1 (en) Speech enhancement and interference suppression
CN115910094B (en) Audio frame processing method, device, electronic device and storage medium
CN111048096A (en) A voice signal processing method, device and terminal
CN118215961A (en) Control of speech retention in speech enhancement
WO2025264909A1 (en) Echo cancellation
WO2025030069A1 (en) Generation of audio content
CN121237116A (en) Sound pickup methods, devices, electronic equipment, and software products
CN117859176A (en) Detecting ambient noise in user-generated content
Narangale et al. Effective prototype algorithm for noise removal through gain and range change
CN115240700A (en) A kind of acoustic equipment and its sound processing method
CN116189701A (en) Call noise reduction method, terminal and computer-readable storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20260115

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR