WO2025240232A1 - Decoder silence generation without coded silence description - Google Patents

Decoder silence generation without coded silence description

Info

Publication number
WO2025240232A1
WO2025240232A1 PCT/US2025/028522 US2025028522W WO2025240232A1 WO 2025240232 A1 WO2025240232 A1 WO 2025240232A1 US 2025028522 W US2025028522 W US 2025028522W WO 2025240232 A1 WO2025240232 A1 WO 2025240232A1
Authority
WO
WIPO (PCT)
Prior art keywords
background noise
audio frame
speech
decoded audio
generate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/028522
Other languages
French (fr)
Inventor
Pravin Kumar RAMADAS
Vivek Rajendran
Toru Uchino
Zisis Iason Skordilis
Hirak DASGUPTA
Nikolai Konrad Leung
Masato Kitazoe
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Qualcomm Inc
Original Assignee
Qualcomm Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Qualcomm Inc filed Critical Qualcomm Inc
Publication of WO2025240232A1 publication Critical patent/WO2025240232A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/012Comfort noise or silence coding
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals
    • G10L25/84Detection of presence or absence of voice signals for discriminating voice from noise

Definitions

  • the present disclosure generally relates to audio coding (e.g., audio encoding and/or decoding). For example, aspects of the present disclosure relate to systems and techniques for decoder silence generation without coded silence description.
  • Audio coding also referred to as voice coding and/or speech coding
  • voice coding and/or speech coding is a technique used to represent a digitized audio signal using as few bits as possible (thus compressing the speech data), while attempting to maintain a certain level of audio quality.
  • An audio or voice encoder is used to encode (or compress) the digitized audio (e.g., speech, music, etc.) signal to a lower bit- rate stream of data.
  • the lower bit-rate stream of data can be input to an audio or voice decoder, which decodes the stream of data and constructs an approximation or reconstruction of the original signal.
  • the audio or voice encoder-decoder structure can be referred to as an audio coder (or voice coder or speech coder) or an audio/voice/speech coder-decoder (codec).
  • Audio coders exploit the fact that speech signals are highly correlated waveforms.
  • Some speech coding techniques are based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech.
  • the different phonemes e.g., vowels, fricatives, and voice fricatives
  • source excitation
  • spectral shape filter
  • the first device includes: one or more memories comprising instructions; and one or more processors coupled to the one or more memories and configured to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • a method for audio playback is provided.
  • the method includes receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • a non-transitory computer-readable medium having stored thereon instructions is provided.
  • the instructions when executed by one or more processors, cause the one or more processors to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • an apparatus for audio playback is provided.
  • the apparatus includes means for receiving a first decoded audio frame, the first decoded audio frame including active speech; means for generating information associated with background noise in the first decoded audio frame; means for detecting a transition to an inactive region after the first decoded audio frame; and means for synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • Aspects generally include a method, apparatus, system, computer program product, non- transitory computer-readable medium, user equipment, base station, wireless communication Qualcomm Ref. No.2400252WO device, and/or processing system as substantially described herein with reference to and as illustrated by the drawings and specification.
  • aspects are described in the present disclosure by illustration to some examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios.
  • Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and/or packaging arrangements.
  • some aspects may be implemented via integrated chip implementations or other non-module- component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail/purchasing devices, medical devices, and/or artificial intelligence devices).
  • Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and/or system-level components.
  • Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects.
  • transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and/or summers).
  • RF radio frequency
  • aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and/or end-user devices of varying size, shape, and constitution.
  • Qualcomm Ref. No.2400252WO
  • FIG. 1 is a block diagram illustrating an example speech processing system, in accordance with some examples
  • FIG.2 is a block diagram illustrating an example feature generator, in accordance with some examples
  • FIG.3 is a block diagram illustrating an example of a voice coding system, in accordance with some examples
  • FIG. 4 is a block diagram illustrating an example of a code-excited linear prediction (CELP)-based voice coding system utilizing a fixed codebook (FCB), in accordance with some examples;
  • FIG. 5 is a block diagram illustrating an example of a voice coding signal synthesis system utilizing a linear time-varying filter generated using a neural network model and a separate linear predictive coding (LPC) filter, in accordance with some examples;
  • FIG. 6 illustrates an audio waveform representing audio information received by a wireless device, in accordance with aspects of the present disclosure; [0021] FIG.
  • FIG. 7 is a block diagram providing an overview of a technique for decoder silence generation without, or with minimal, coded SIDs, in accordance with aspects of the present disclosure; Qualcomm Ref. No.2400252WO [0022]
  • FIG. 8A is a block diagram illustrating an example inactive synthesizer, in accordance with aspects of the present disclosure; [0023]
  • FIG. 8B is a is a block diagram illustrating another example inactive synthesizer, in accordance with aspects of the present disclosure;
  • FIG. 9A is a block diagram illustrating a denoiser, in accordance with aspects of the present disclosure; [0025] FIG.
  • FIG. 9B is a block diagram illustrating non-negative matrix factorization (NMF) denoising, in accordance with aspects of the present disclosure
  • NMF non-negative matrix factorization
  • FIG.10 is a block diagram illustrating a technique for location aware background noise synthetization, in accordance with aspects of the present disclosure
  • FIG. 11 is a block diagram illustrating a technique for random background noise synthetization, in accordance with aspects of the present disclosure
  • FIG.12 illustrates a spectrum estimator for measuring additional spectral characteristics, in accordance with aspects of the present disclosure
  • FIG.13 is a block diagram illustrating another technique for decoder silence generation without, or with minimal, coded SIDs, in accordance with aspects of the present disclosure
  • FIG.10 is a block diagram illustrating a technique for location aware background noise synthetization, in accordance with aspects of the present disclosure
  • FIG. 11 is a block diagram illustrating a technique for random background noise synthetization, in accordance with aspects of the present disclosure
  • FIG.12 illustrate
  • FIG. 14 is signal diagram illustrating signals for decoder silence generation without coded silence descriptions, in accordance with aspects of the present disclosure
  • FIG.15 is a block diagram illustrating an example of an audio codec system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer
  • FIG. 15 is a block diagram illustrating an example of an audio codec system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer
  • FIG. 16 illustrates an example of an audio codec system (e.g., generative voice codec) that includes a decoder configured to implement an NHV-based neural speech synthesizer with a technique for random background noise synthesis, in accordance with aspects of the present disclosure
  • FIG.17 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using Qualcomm Ref. No.2400252WO a neural speech synthesizer comprising a linear prediction coding (LPC) network with a technique for random background noise synthesis, in accordance with aspects of the present disclosure
  • LPC linear prediction coding
  • FIG. 18 is a flow diagram illustrating an example of a process for audio playback, in accordance with aspects of the present disclosure.
  • FIG.19 is a diagram illustrating an example of a computing system, according to aspects of the present disclosure.
  • DETAILED DESCRIPTION [0036] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
  • SID frames may be used to describe the background noise during breaks.
  • the SID may include information about background noise, and this information may be used to generate comfort noise by a decoding device.
  • silence may refer to audio frames without speech, but with background noise, while absolute silence may refer to an absence of speech and background noise.
  • radio resource and/or encoder device power may be limited and techniques to reduce encoding and transmitting SID frames may be useful.
  • Qualcomm Ref. No.2400252WO Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for decoder silence generation without coded silence description.
  • SID frames may be minimized or eliminated by predicting and synthesizing comfort noise at the decoder (e.g., decoding device).
  • speech may include regions, such as an active speech region, a hangover region, and an inactive (e.g., silent) region.
  • an active speech region speech is present along with background noise.
  • background noise may be present.
  • it may be useful to generate information associated with the background noise while in the active speech region. This information associated with the background noise may be used to synthesize background noise when a transition to the inactive region is detected.
  • the information associated with the background noise may be obtained by applying a denoise filter to an audio frame from the active speech region and subtracting the denoised version of the audio frame from the audio frame to obtain a background noise frame. Information associated with the background noise may then be generated based on the background noise frame.
  • a machine learning (ML) model may be used to predict the background noise from the audio frame.
  • NMF non-negative matrix factorization
  • an audio frame from the handover region may be received. As indicated above, audio frames from the handover region may include primarily background noise.
  • the information about the background noise may be generated using audio frames from the active speech region and audio frames from the handover region.
  • the synthesized background noise may be synthesized based on location information from the sending device. In some cases, the synthesized background noise may be synthesized in part by a ML model. In some cases, the synthesized background noise may be synthesized based on random noise. In some cases, a transition to an inactive region may be signaled by a radio access network message, such as an RRC message.
  • a radio access network message such as an RRC message.
  • Autoencoders perform dimensionality reduction (e.g., an N-dimensional vector input vector is passed into the encoder of the autoencoder, and lossy compression is performed to represent the important aspects of the input vector in an M- dimensional vector, where M is smaller than N, e.g. by an order of magnitude).
  • a drawback of using an autoencoder for speech coding is that it cannot necessarily exploit the temporal relationship between sets of input data.
  • U.S. Patent No.11,526,734 (“the ‘734 patent”, assigned to Qualcomm, Inc.), "Method and Apparatus for recurrent auto encoding", Yang et. al.
  • the FRAE in the '734 patent can be used for training and application of compression of sequential data with temporal correlation.
  • the recurrent structure of the FRAE can be used to efficiently extract the redundancy embedded along the time-dimension of sequential data and enables compact discrete representation of the data at the bottleneck in a sequential fashion.
  • Table 1 of the '734 patent illustrates the MSE (Mean Squared Error) used as Mel-scale mean-square-error configured as a reconstruction loss for training of the FRAE, where the MSE of each frequency bin is scaled according to its weight at Mel-frequency both for latent feedback and output feedback.
  • the FRAE described in the ‘734 patent has two advantageous features not present in autoencoders: (1) recurrent layers, e.g. LSTM or GRU layers, have memory of the past; and (2) feedback from the decoder of the autoencoder to the encoder of the autoencoder.
  • the feedback connection 150 in Fig. 2, Fig. 3 and Fig. 7 of the ‘734 patent provides additional historical information from a state (h t ) in the decoder to the encoder indicative of how reconstruction of a prior input has fared. This feedback loop is present during training and inference, which allows the encoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder.
  • the feedback loop is analogous to a mode switch input in that it is not encoded by the encoder, but it is used as an input that influences how the encoder operates on the next input vector.
  • the ‘734 patent describes a second feedback connection (152) from the state (h t ) of the decoder for a first iteration of series of inputs X t , to the next iteration of series of inputs Qualcomm Ref. No.2400252WO Xt+1.
  • This second feedback connection (152) enables the decoder to learn from its previous reconstruction attempts, providing additional historical context about how reconstruction of a prior input has fared.
  • This second feedback loop is also present during both training and inference, which allows the decoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder.
  • the ‘734 patent also describes other optional feedback connections. For example, an embedding vector (z) may be fed back via a third feedback connection (356 in Fig.3 of the ‘734 patent) to the encoder. As another optional example, the '734 patent describes that the FRAE can be a variational autoencoder.
  • FIG. 1 illustrates an example implementation of a system-on-a-chip (SoC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein.
  • SoC system-on-a-chip
  • Parameters or variables e.g., neural signals and synaptic weights
  • system parameters associated with a computational device e.g., neural network with weights
  • delays e.g., frequency bin information, task information, among other information
  • NPU neural processing unit
  • GPU graphics processing unit
  • DSP digital signal processor
  • Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118.
  • the SoC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and the like, and a multimedia processor 112 that may, Qualcomm Ref. No.2400252WO for example, detect and recognize gestures, speech, and/or other interactive user action(s) or input(s).
  • the NPU 108 is implemented in the CPU 102, DSP 106, and/or GPU 104.
  • the SoC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and/or a signal synthesis system 120.
  • ISPs image signal processors
  • the signal synthesis system 120 can be implemented or configured as a speech synthesis system, including a neural speech decoder and/or a neural homomorphic vocoder (NHV) system, which can be used to generate speech (e.g., perform speech synthesis), can be implemented in a text-to-speech (TTS) system, etc.
  • the sensor processor 114 can be associated with or connected to one or more sensors for providing sensor input(s) to sensor processor 114.
  • the one or more sensors and the sensor processor 114 can be provided in, coupled to, or otherwise associated with a same computing device.
  • the one or more sensors can include one or more microphones for receiving sound (e.g., an audio input), including sound or audio inputs that can be used to perform various speech synthesis tasks and/or to generate a reconstructed speech signal from the audio input, etc.
  • the sound or audio input received by the one or more microphones (and/or other sensors) may be digitized into data packets for analysis and/or transmission.
  • the audio input may include ambient sounds in the vicinity of a computing device associated with the SoC 100 and/or may include speech from a user of the computing device associated with the SoC 100.
  • a computing device associated with the SoC 100 can additionally, or alternatively, be communicatively coupled to one or more peripheral devices (not shown) and/or configured to communicate with one or more remote computing devices or external resources, for example using a wireless transceiver and a communication network, such as a cellular communication network.
  • SoC 100, DSP 106, NPU 108 and/or signal synthesis (e.g., neural speech decoder, NHV, etc.) system 120 may be configured to perform audio signal processing.
  • the signal synthesis system 120 may be configured to perform steps for speech synthesis and/or neural homomorphic vocoding, etc.
  • FIG.2 depicts an example of a feature generator 200, in accordance with aspects of the present disclosure. It should be understood that many techniques may be used to generate feature vectors for an audio input and that feature generator 200 is just a single example of a technique that may be used to generate feature vectors.
  • Feature generator 200 receives an audio signal at signal pre-processor 202.
  • the audio signal may be from an audio source of an electronic device, such a microphone.
  • Signal pre- processor 202 may perform various pre-processing steps on the received audio signal. For example, signal pre-processor 202 may split the audio signal into parallel audio signals and delay one of the signals by a predetermined amount of time to prepare the audio signals for input into a Fast-Fourier Transform (FFT) circuit.
  • FFT Fast-Fourier Transform
  • signal pre-processor 202 may perform a windowing function, such as a Hamming, Hann, Blackman-Harris, Kaiser-Bessel window function, or other sine-based window function, which may improve the performance of further processing stages, such as signal domain transformer 204.
  • a windowing (or window) function in may be used to reduce the amplitude of discontinuities at the boundaries of each finite sequence of received audio signal data to improve further processing.
  • signal pre-processor 202 may convert the audio signal data from parallel to serial, or vice versa, for further processing.
  • the pre-processed audio signal from the signal pre-processor 202 may be provided to signal domain transformer 204, which may transform the pre-processed audio signal from a first domain into a second domain, such as from a time domain into a frequency domain.
  • signal domain transformer 204 implements a Fourier transform, such as a Fast-Fourier transform (FFT).
  • FFT Fast-Fourier transform
  • the Fast Fourier transform may be a 16-band (or bin, channel, or point) FFT, which generates a compact feature set that may be efficiency processed by a model.
  • a Fourier transform provides fine spectral domain information about the incoming audio signal as compared to conventional single channel processing, such as conventional hardware SNR threshold detection.
  • the result of signal domain transformer 204 is a set of audio features, such as a set of voltages, powers, or energies per frequency band in the transformed data.
  • the set of audio features may then be provided to signal feature filter 206, which may reduce the size of or compress the feature set in the audio feature data.
  • signal feature filter 206 may discard certain features from the audio feature set, such as symmetric or Qualcomm Ref.
  • No.2400252WO redundant features from multiple bands of a multi-band FFT Discarding this data reduces the overall size of the data stream for further processing and may be referred to a compressing the data stream.
  • a 16-band FFT may include 8 symmetric or redundant bands of after the powers are squared because audio signals are real.
  • signal feature filter 206 may filter out the redundant or symmetric band information and output an audio feature vector 208.
  • output of the signal feature filter may be compressed or otherwise processed prior to output as the audio feature vector 208.
  • the audio feature vector 208 may be provided to a speech synthesis system (e.g., such as the signal synthesis system 120 of FIG.1) for processing by speech synthesis or NHV model.
  • Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and/or other use.
  • FIG.3 is a block diagram illustrating an example of a voice coding system 350 (which can also be referred to as a voice or speech coder or a voice coder- decoder (codec)).
  • a voice encoder 352 of the voice coding system 350 can use a voice coding algorithm to process a speech signal 351.
  • the speech signal 351 can include a digitized speech signal generated from an analog speech signal from a given source.
  • the digitized speech signal can be generated using a filter to eliminate aliasing, a sampler to convert to discrete- time, and an analog-to-digital converter for converting the analog signal to the digital domain.
  • the resulting digitized speech signal (e.g., speech signal 351) is a discrete-time speech signal with sample values (referred to herein as samples) that are also discretized.
  • the voice encoder 352 can generate a compressed signal (including a lower bit-rate stream of data) that represents the speech signal 351 using as few bits as possible, while attempting to maintain a certain quality level for the speech.
  • the voice encoder 352 can use any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., Code-excited linear prediction (CELP), algebraic-CELP (ACELP), or other linear prediction technique) or other voice coding algorithm.
  • the voice encoder 352 can compress the speech signal 351 in an attempt to reduce the bit-rate of the speech signal 351.
  • the bit-rate of a signal is based on the sampling frequency and the number of bits per sample. For instance, the bit-rate of a speech signal can be determined as ⁇ ⁇ ⁇ ⁇ ⁇ , where BR is the bit-rate, S is the sampling frequency, and b is the number of bits per Qualcomm Ref. No.2400252WO sample.
  • the bit-rate (BR) of a signal would be a bit-rate of 128 kilobits per second (kbps).
  • S sampling frequency
  • b 16 bits per sample
  • BR bit-rate
  • the compressed speech signal can then be stored and/or sent to and processed by a voice decoder 354.
  • the voice decoder 354 can communicate with the voice encoder 352, such as to request speech data, send feedback information, and/or provide other communications to the voice encoder 352.
  • the voice encoder 352 or a channel encoder can perform channel coding on the compressed speech signal before the compressed speech signal is sent to the voice decoder 354.
  • channel coding can provide error protection to the bitstream of the compressed speech signal to protect the bitstream from noise and/or interference that can occur during transmission on a communication channel.
  • the voice decoder 354 can decode the data of the compressed speech signal and construct a reconstructed speech signal 355 that approximates the original speech signal 351.
  • the reconstructed speech signal 355 includes a digitized, discrete-time signal that can have the same or similar bit-rate as that of the original speech signal 351.
  • the voice decoder 354 can use an inverse of the voice coding algorithm used by the voice encoder 352, which as noted above can include any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., CELP, ACELP, or other suitable linear prediction technique) or other voice coding algorithm.
  • the reconstructed speech signal 355 can be converted to continuous-time analog signal, such as by performing digital-to-analog conversion and anti-aliasing filtering.
  • Voice coders can exploit the fact that speech signals are highly correlated waveforms.
  • the samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame.
  • each frame can be 10-20 milliseconds (ms) in length.
  • Various voice coding algorithms can be used to encode a speech signal.
  • CELP code-excited linear prediction
  • the CELP model is based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech.
  • the different phonemes e.g., vowels, fricatives, and voice fricatives
  • CELP uses a linear prediction (LP) model to model the vocal tract, and uses entries of a fixed codebook (FCB) as input to the LP model.
  • LP linear prediction
  • FCB fixed codebook
  • long-term linear prediction can be used to model pitch of a speech signal
  • short-term linear prediction can be used to model the spectral shape (phoneme) of the speech signal.
  • Entries in the FCB are based on coding of a residual signal that remains after the long-term and short-term linear prediction modeling is performed.
  • long-term linear prediction and short-term linear prediction models can be used for speech synthesis, and a fixed codebook (FCB) can be searched during encoding to locate the best residual for input to the long-term and short-term linear prediction models.
  • FIG.4 is a block diagram illustrating an example of a CELP-based voice coding system 470, including a voice encoder 472 and a voice decoder 474.
  • the voice encoder 472 can obtain a speech signal 471 and can segment the samples of the speech signal into frames and sub-frames.
  • a frame of N samples can be divided into sub-frames.
  • a frame of 440 samples can be divided into four sub-frames each having 60 samples.
  • the voice encoder 472 chooses the parameters (e.g., gain, filter coefficients or linear prediction (LP) coefficients, etc.) for a synthetic speech signal so as to match as much as possible the synthetic speech signal with the original speech signal.
  • the voice encoder 472 can include a short-term linear prediction (LP) engine 480, a long- term linear prediction (LTP) engine 482, and a fixed codebook (FCB) 484.
  • the short-term LP engine 480 models the spectral shape (phoneme) of the speech signal.
  • the short-term LP engine 480 can perform a short-term LP analysis on each frame to yield linear prediction (LP) coefficients.
  • the input to the short-term LP engine 480 can be the original speech signal or a pre-processed version of the original speech signal.
  • the short-term LP engine 480 can perform linear prediction for each frame by estimating the value of a current speech sample based on a linear combination of past speech samples.
  • a speech signal s(n) can be represented using an autoregressive (AR) model, such as ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , where each sample is represented as a linear combination of the Qualcomm Ref.
  • AR autoregressive
  • No.2400252WO previous m samples plus a prediction error term ⁇ The weighting coefficients a 1 , a 2 , through a m can be referred to as the LP coefficients.
  • the prediction error term ⁇ can be found as follows: ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
  • the short-term LP engine 480 can obtain the LP coefficients.
  • the LP coefficients can be used to form an analysis filter as given in Equation 1, below: ⁇ ⁇ ⁇ ⁇ 1 ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ Eq.
  • the short-term LP engine 480 can solve for ⁇ (which can be referred to as a transfer function) by computing the LP coefficients ( ⁇ ⁇ ) that minimize the error in the above AR model equation ( ⁇ ) or other error metric.
  • the LP coefficients can be determined using a Levinson–Durbin method, a Leroux–Gueguen algorithm, or other suitable technique.
  • the voice encoder 472 can send the LP coefficients to the voice decoder 474.
  • the voice decoder 474 can determine the LP coefficients, in which case the voice encoder 472 may not send the LP coefficients to the voice decoder 474.
  • LTP engine 482 models the pitch of the speech signal.
  • Pitch is a feature that determines the spacing or periodicity of the impulses in a speech signal.
  • speech signals are generated when the airflow from the lungs is periodically interrupted by movements of the vocal cords. The time between successive vocal cord openings corresponds to the pitch period.
  • the LTP engine 482 can be applied to each frame or each sub-frame of a frame after the short- term LP engine 480 is applied to the frame.
  • the LTP engine 482 can predict a current signal sample from a past sample that is one or more pitch periods apart from a current sample (hence the term “long-term”).
  • the current signal sample can be predicted as ⁇ ⁇ ⁇ ⁇ ⁇ , where T denotes the pitch period, ⁇ denotes the pitch gain, and ⁇ ⁇ ⁇ denotes an LP residual for a previous sample one or more pitch periods apart from a current sample.
  • Pitch period can be estimated at every frame. By comparing a frame with past samples, it is possible to identify the period in which the signal repeats itself, resulting in an estimate of the actual pitch period.
  • the LTP engine 482 can be applied separately to each sub-frame. Qualcomm Ref. No.2400252WO [0066]
  • the FCB 484 can include a number (denoted as L) of long-term linear prediction (LTP) residuals.
  • An LTP residual includes the speech signal components that remain after the long-term and short-term linear prediction modeling is performed.
  • the LTP residuals can be, for example, fixed or adaptive and can contain deterministic pulses or random noise (e.g., white noise samples).
  • the voice encoder 472 can pass through the number L of LTP residuals in the FCB 484 a number of times for each segment (e.g., each frame or other group of samples) of the input speech signal, and can calculate an error value (e.g., a mean-squared error value) after each pass.
  • the LTP residuals can be represented using codevectors. The length of each codevector can be equal to the length of each sub-frame, in which case a search of the FCB 484 is performed once every sub- frame.
  • the LTP residual providing the lowest error can be selected by the voice encoder 472.
  • the voice encoder 472 can select an index corresponding to the LTP residual selected from the FCB 484 for a given sub-frame or frame.
  • the voice encoder 472 can send the index to the voice decoder 474 indicating which LTP residual is selected from the FCB 484 for the given sub-frame or frame.
  • a gain associated with the lowest error can also be selected, and send to the voice decoder 474.
  • the voice decoder 474 includes an FCB 494, an LTP engine 492, and a short-term LP engine 490.
  • the FCB 494 has the same LTP residuals (e.g., codevectors) as the FCB 484.
  • the voice decoder 474 can extract an LTP residual from the FCB 494 using the index transmitted to the voice decoder 474 from the voice encoder 472.
  • the extracted LTP residual can be scaled to the appropriate level and filtered by the LTP engine 492 and the short-term LP engine 490 to generate a reconstructed speech signal 475.
  • the LTP engine 492 creates periodicity in the signal associated with the fundamental pitch frequency, and the short-term LP engine 490 generates the spectral envelope of the signal.
  • linear predictive-based coding systems can also be used to code voice signals, including enhanced voice services (EVS), adaptive multi-rate (AMR) voice coding systems, mixed excitation linear prediction (MELP) voice coding systems, linear predictive coding-10 (LPC-10), among others.
  • EVS enhanced voice services
  • AMR adaptive multi-rate
  • MELP mixed excitation linear prediction
  • LPC-10 linear predictive coding-10
  • a voice codec for some applications and/or devices may be needed to deliver higher quality coding of speech signals at low bit-rates, with low complexity, and with low memory requirements.
  • IoT Internet-of-Things
  • existing linear predictive- based codecs cannot meet such requirements.
  • ACELP-based coding systems provide Qualcomm Ref. No.2400252WO high quality, but do not provide low bit-rate or low complexity/low memory.
  • linear- predictive coding systems provide low bit-rate and low complexity/low memory, but do not provide high quality.
  • machine learning systems e.g., using a neural network model
  • a neural network-based voice decoder can generate coefficients for at least one linear filter. The linear filter can then be used to generate a reconstructed signal.
  • a neural network-based voice decoder can be highly complex and resource intensive. For instance, the neural network-based voice decoder will have to perform the operations of a linear predictive filter (LPC), such as the short-term LP engine 480 of FIG.4.
  • LPC linear predictive filter
  • FIG.6 is a diagram illustrating an example of a voice decoding signal synthesis system 500 utilizing a linear time-varying filter 504 with coefficients generated using a neural network (NN) filter estimator 502 and a separate linear predictive coding (LPC) filter 506.
  • the voice decoding signal synthesis system 500 is configured to decode data of the compressed speech signal to generate a reconstructed speech signal ⁇ (also referred to as a synthesized speech sample) for a current time instant n that approximates an original speech signal that was previously compressed by a voice encoder (not shown).
  • the voice encoder can be similar to and can perform some or all of the functions of the voice encoder 252 described above with respect to FIG.2B, or other type of voice encoder.
  • the voice encoder can include a short-term LP engine, an LTP engine, and an FCB.
  • the voice encoder can include a magnitude spectrum generator (e.g., Mel-scale magnitude spectrum or full spectrum magnitude), a short-term linear prediction (LP) engine, and a pitch tracker that detects a fundamental pitch harmonic frequency of the speech and pitch correlation.
  • a magnitude spectrum generator e.g., Mel-scale magnitude spectrum or full spectrum magnitude
  • LP short-term linear prediction
  • pitch tracker that detects a fundamental pitch harmonic frequency of the speech and pitch correlation.
  • the voice encoder can extract (and in some cases quantize) a set of features (referred to as a feature set) from the speech signal, and can send the extracted (and in some cases quantized) feature set to the voice decoding signal synthesis system 500.
  • the features that are computed by the voice encoder can depend on a particular encoder implementation used.
  • Various illustrative examples of feature sets are provided below according to different encoder implementations, which can be extracted by the voice encoder (and in some cases quantized), and sent to the voice decoding signal synthesis system 500.
  • Qualcomm Ref. No.2400252WO feature sets can be extracted by the voice encoder.
  • the voice encoder can extract any set of features, can quantize that feature set, and can send the feature set to the voice decoding signal synthesis system 500.
  • various combinations of features can be extracted as a feature set by the voice encoder.
  • a feature set can include one or any combination of the following features: Linear Prediction (LP) coefficients; Line Spectral Pairs (LSPs); Line Spectral Frequencies (LSFs); pitch lag with integer or fractional accuracy; pitch gain; pitch correlation; Mel-scale frequency cepstral coefficients (also referred to as Mel cepstrum) of the speech signal; Bark-scale frequency cepstral coefficients (also referred to as bark cepstrum) of the speech signal; Mel-scale frequency cepstral coefficients of the LTP residual; Bark-scale frequency cepstral coefficients of the LTP residual; a spectrum (e.g., Discrete Fourier Transform (DFT) or other spectrum) of the speech signal; and/or a spectrum (e.g., DFT or other spectrum) of the LTP residual; voicing level of each frequency band of each speech frame; fundamental frequency of pitch harmonics; pitch correlation of each speech frame; time domain pitch lag of each speech frame.
  • LP Linear Prediction
  • LSPs Line
  • the voice encoder can use any estimation and/or quantization method, such as an engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec) or a neural network-based estimation and/or quantization scheme (e.g., convolutional or fully-connected (dense) or recurrent Autoencoder, or other neural network-based estimation and/or quantization scheme).
  • voice codec e.g. EVS, AMR, or other voice codec
  • a neural network-based estimation and/or quantization scheme e.g., convolutional or fully-connected (dense) or recurrent Autoencoder, or other neural network-based estimation and/or quantization scheme.
  • the voice encoder can also use any frame size, frame overlap, and/or update rate for each feature.
  • the voice encoder can also include extra redundancies in the features to ensure robustness of operation against packet losses.
  • one example of features that can be extracted from a voice signal by the voice encoder includes linear prediction (LP) coefficients and/or line spectral frequencies (LSFs).
  • LP linear prediction
  • LSFs line spectral frequencies
  • the voice encoder can estimate LP coefficients (and/or LSFs) from a speech signal using the Levinson-Durbin algorithm.
  • the LP coefficients and/or LSFs can be Qualcomm Ref. No.2400252WO estimated using an autocovariance method for LP estimation.
  • the LP coefficients can be determined, and an LP to LSF conversion algorithm can be performed to obtain the LSFs. Any other LP and/or LSF estimation engine or algorithm can be used, such as an LP and/or LSF estimation engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
  • Various quantization techniques can be used to quantize the LP coefficients and/or LSFs.
  • the voice encoder can use a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), or other vector quantization technique to quantize the LP coefficients and/or LSFs.
  • SSVQ single stage vector quantization
  • MSVQ multi-stage vector quantization
  • a predictive or adaptive SSVQ or MSVQ can be used to quantize the LP coefficients and/or LSFs.
  • an autoencoder or other neural network based technique can be used by the voice encoder to quantize the LP coefficients and/or LSFs. Any other LP and/or LSF quantization engine or algorithm can be used, such as an LP and/or LSF quantization engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
  • Another example of features that can be extracted from a voice signal by the voice encoder includes pitch lag (integer and/or fractional), pitch gain, and/or pitch correlation.
  • the voice encoder can estimate the pitch lag, pitch gain, and/or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation estimation engine or algorithm (e.g. autocorrelation-based pitch lag estimation).
  • the voice encoder can use a pitch lag, gain, and/or correlation estimation engine (or algorithm) from any suitable voice codec (e.g. EVS, AMR, or other voice codec).
  • voice codec e.g. EVS, AMR, or other voice codec.
  • quantization techniques can be used to quantize the pitch lag, pitch gain, and/or pitch correlation.
  • the voice encoder can quantize the pitch lag, pitch gain, and/or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation quantization engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec).
  • voice codec e.g. EVS, AMR, or other voice codec
  • an autoencoder or other neural network based technique can be used by the voice encoder to quantize the pitch lag, pitch gain, and/or pitch correlation features.
  • Another example of features that can be extracted from a voice signal by the voice encoder includes the Mel cepstrum coefficients and/or Bark cepstrum coefficients of the speech signal, and/or the Mel cepstrum coefficients and/or Bark cepstrum coefficients of the LTP residual. Qualcomm Ref.
  • the voice encoder can use a Mel or Bark frequency cepstrum technique that includes Mel or Bark frequency filter banks computation, filter bank energy computation, logarithm application, and discrete cosine transform (DCT) or truncation of the DCT.
  • DCT discrete cosine transform
  • Various quantization techniques can be used to quantize the Mel cepstrum coefficients and/or Bark cepstrum coefficients. For example, vector quantization (single stage or multistage) or predictive/adaptive vector quantization can be used.
  • an autoencoder or other neural network based technique can be used by the voice encoder to quantize the Mel cepstrum coefficients and/or Bark cepstrum coefficients. Any other suitable cepstrum quantization methods can be used.
  • Another example of features that can be extracted from a voice signal by the voice encoder includes the spectrum of the speech signal and/or the spectrum of the LTP residual.
  • Various estimation techniques can be used to compute the spectrum of the speech signal and/or the LTP residual. For example, a Discrete Fourier transform (DFT), a Fast Fourier Transform (FFT), or other transform of the speech signal can be determined.
  • DFT Discrete Fourier transform
  • FFT Fast Fourier Transform
  • Quantization techniques that can be used to quantize the spectrum of the voice signal can include vector quantization (single stage or multistage) or predictive/adaptive vector quantization.
  • an autoencoder or other neural network based technique can be used by the voice encoder to quantize the spectrum. Any other suitable spectrum quantization methods can be used.
  • any one of the above-described features or any combination of the above-described features can be estimated, quantized, and sent by the voice encoder to the voice decoding signal synthesis system 500 depending on the particular encoder implementation that is used.
  • the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, and pitch correlation.
  • the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal.
  • the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal.
  • the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal.
  • the voice encoder can Qualcomm Ref.
  • No.2400252WO estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal.
  • the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the LTP residual.
  • the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the LTP residual.
  • the voice decoding signal synthesis system 500 includes a neural network filter estimator 502, a linear time-varying filter 504 generated by the neural network, and a linear predictive coding (LPC) filter 506.
  • the LPC filter 506 can include a time-varying LPC filter.
  • the neural network filter estimator 502 is trained to generate filter coefficients for the linear time-varying filter 504.
  • the neural network model of the neural network filter estimator 502 can include any neural network architecture that can be trained to model the filter coefficients for the linear time-varying filter 504.
  • Examples of neural network architectures that can be included in the neural network filter estimator 502 include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and/or other type(s) of neural network architectures or models.
  • the voice decoding signal synthesis system 500 e.g., the neural network model of the neural network filter estimator 502 can be trained using any suitable neural network training technique.
  • the neural network model of the neural network filter estimator 502 can be trained using supervised learning techniques based on backpropagation. For instance, corresponding input and target output pairs can be provided to the neural network filter estimator 502 for training.
  • the input to the neural network filter estimator 502 can include log-Mel-frequency spectrum features or coefficients (e.g., 80 log-Mel features ⁇ , ⁇ 501 shown in FIG. 5).
  • the target output (or label or ground truth) for training the neural network filter estimator 502 can include the target speech sample ⁇ ⁇ ⁇ ⁇ 505 for the current time instant n, as shown in FIG. 5.
  • a loss 507 will be computed based on the reconstructed sample ⁇ ⁇ ⁇ ⁇ and the target output speech sample ⁇ ⁇ ⁇ ⁇ 505 for time instant n.
  • the target output can include a speech signal ⁇ ⁇ ⁇ ⁇ that is generated after passing the target speech through an LPC analysis filter (inverse of LPC filter 506).
  • LPC analysis filter inverse of LPC filter 506
  • the loss will be computed based on the output ⁇ generated by the linear time- varying filter generated by the neural network 504 and the target output ⁇ (e.g., where both are in speech residual domain).
  • Backpropagation can be performed to train the neural network filter estimator 502 using the inputs and the target output. Backpropagation can include a forward pass, a loss function, a backward pass, and a parameter update to update one or more parameters (e.g., weight, bias, or other parameter).
  • training of the neural network filter estimator 502 and/or training of one or more of the machine learning systems or neural networks described herein can be performed using online training (e.g., in some case on-device training), offline training, and/or various combinations of online and offline training.
  • online may refer to time periods during which the input data is processed, for instance for performance of generative voice codec processing implemented by the systems and techniques described herein.
  • offline may refer to idle time periods or time periods during which input data is not being processed. Additionally, offline may be based on one or more time conditions (e.g., after a particular amount of time has expired, such as a day, a week, a month, etc.) and/or may be based on various other conditions such as network and/or server availability, etc., among various others.
  • offline training of a machine learning model e.g., a neural network model
  • a first device e.g., a server device
  • a second device can receive the trained model from the second device.
  • the second device e.g., a mobile device, an XR device, a vehicle or system/component of the vehicle, or other device
  • the forward pass can include passing the input data (e.g., the log-Mel- frequency spectrum features or coefficients, such as the 80 log-Mel features ⁇ , ⁇ 501 shown in FIG.5) through the neural network filter estimator 502.
  • the weights of the neural network model are initially randomized before the neural network filter estimator 502 is trained. For a first training Qualcomm Ref.
  • a loss function can be used to analyze the loss 507 (or error) in the reconstructed or synthesized sample output ⁇ . Any suitable loss function definition can be used.
  • a loss function includes a mean squared error (MSE).
  • MSE is defined as ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , which calculates the sum of one-half times the actual answer the predicted (output) answer squared.
  • the loss can be set to be equal to the value of ⁇ ⁇ .
  • Other loss functions may include a difference of magnitude spectrums between the target and output signals, where the difference may be computed as absolute difference, squared difference, or logarithmic difference between the magnitude spectrum of each speech frame, and then aggregated over all speech frames.
  • the loss (or error) will be high for the first training iterations since the actual values will be much different than the predicted output.
  • the goal of training is to minimize the amount of loss so that the predicted output is the same as the training label.
  • the neural network filter estimator 502 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized.
  • a derivative of the loss with respect to the weights (denoted as dL/dW, where W represents the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network.
  • a weight update can be performed by updating all the weights of the filters.
  • the weights can be updated so that they change in the opposite direction of the gradient.
  • the weight update can be denoted as ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , where w denotes a weight, wi ⁇ denotes the initial weight, and ⁇ denotes a learning rate.
  • rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.
  • a multi-resolution STFT loss ⁇ ⁇ and adversarial losses ⁇ ⁇ and ⁇ can be computed from ⁇ and ⁇ .
  • linear time-varying filters are fully differentiable, gradients can propagate back to the neural network filter estimator 502.
  • Qualcomm Ref. No.2400252WO [0087] Using the filter coefficients generated by the neural network filter estimator 502, the linear time-varying filter 504 can process an excitation signal 503 to generate another signal ⁇ .
  • the signal ⁇ can be used as an excitation signal to excite the LPC filter 506.
  • the linear time- varying filter 504 is a linear filter, which preserves the linearity property between inputs and outputs.
  • a linear filter is associated with a mapping L:RZ ⁇ RZ, ⁇ ⁇ ⁇ ⁇ L ⁇ that has the following property: for any ⁇ , RZ and any ⁇ , ⁇ ⁇ R , L ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ L ⁇ ⁇ ⁇ L ⁇ .
  • time-varying linear filters can be characterized by the set of impulse responses at each time lag h ⁇ ⁇ L ⁇ ⁇ ⁇ for each ⁇ ⁇ Z.
  • the output of a time varying linear filter is L ⁇ ⁇ ⁇ ⁇ ⁇ h ⁇ (where ⁇ is a convolutional operator) or some heuristic combination of filter and impulse responses, e.g. overlap-add on windowed and filtered signal segments, etc.
  • the LPC filter 506 can use the signal ⁇ as input to generate the reconstructed or synthesized speech sample ⁇ for the current time instant n.
  • the LPC filter 506 is a linear filter and in some cases is time varying, as defined above with respect to the linear time-varying filter 504.
  • the LPC filter 506 can be a form of time-varying filter used for processing of speech.
  • the LPC filter 506 includes filter coefficients for each speech frame that can be computed using the autocorrelation of a speech or audio signal.
  • the LPC filter 506 can be used to model the spectral shape (or phenome or envelope) of the speech signal.
  • a signal ⁇ can be filtered by an autoregressive (AR) process synthesizer to obtain an AR signal.
  • AR autoregressive
  • a linear predictor can be used to predict the AR signal (which can be denoted as prediction ⁇ ) as a linear combination of the previous m samples as follows: ⁇ ⁇ ⁇ Qualcomm Ref.
  • ⁇ ⁇ terms ( ⁇ ⁇ , ⁇ ⁇ , ... ⁇ ⁇ ) are estimates of the AR parameters (also referred to as LP coefficients).
  • a residual signal can be the difference between the original AR signal and the predicted AR signal represented as prediction ⁇ (e.g., the difference between the actual sample and the predicted sample).
  • the linear prediction coding can be used to find the best linear prediction coefficients for minimizing a quadratic error function, and thus the error.
  • the linear prediction process removes the short-term correlation from the speech signal.
  • the linear prediction coefficients are an efficient way to represent the short-term spectrum of the speech signal.
  • the LPC filter 506 determines the prediction ⁇ for the current sample n using computed or received coefficients and A transfer function ⁇ ⁇ ⁇ . For instance, in some examples, the LPC filter coefficients are received from the encoder. In other examples, the voice decoding signal synthesis system 500 can derive the LPC filter coefficients, such as using other features (e.g., Mel spectrum features) sent by an encoder to the voice decoding signal synthesis system 500. For instance, the voice decoding signal synthesis system 500 can use Mel spectrum features 501 to derive the LPC filter coefficients for the LPC filter 506.
  • features e.g., Mel spectrum features
  • the LPC filter 506 can determine the final reconstructed (or predicted) sample ⁇ using the output ⁇ from the linear time-varying filter 504 (for the current sample n).
  • further components can be used along with the neural network filter estimator 502 and the linear time-varying filter 504, such as an impulse train generator 514 and a random noise generator 516.
  • the linear time-varying filter 504 can include a harmonic linear time-varying filter 518 and a noise linear time-varying filter 520.
  • a voice encoder can include a pitch tracker 510 and a feature extraction engine 512.
  • the feature extraction engine 512 can be the same as or similar to the feature generator 200 of FIG.2.
  • the original speech signal ⁇ and reconstructed signal ⁇ are divided into non-overlapping frames with frame length L.
  • the term ⁇ can be defined as a frame index
  • the term ⁇ can be defined as a discrete time index
  • the term ⁇ can be defined as a feature index.
  • the total number of frames ⁇ and total number of sampling points ⁇ may follow ⁇ ⁇ ⁇ ⁇ ⁇ .
  • the terms ⁇ , ⁇ , ⁇ , ⁇ , ⁇ , ⁇ are finite duration signals, in which 0 ⁇ ⁇ ⁇ ⁇ 1.
  • Impulse responses h ⁇ , and h ⁇ may be infinitely long, in which ⁇ ⁇ Z.
  • Impulse response h may be causal, in which ⁇ ⁇ Z E Z and ⁇ ⁇ 0.
  • Qualcomm Ref. No.2400252WO [0092]
  • the impulse train generator 514 can generate an impulse train ⁇ from a frame-wise fundamental frequency ⁇ ⁇ ⁇ ⁇ ⁇ output by the pitch tracker 510.
  • the impulse train generator 514 can generate alias-free discrete time impulse trains using additive synthesis.
  • the impulse train generator 514 can use a low-passed sum of sinusoids to generate an impulse train: ⁇ 2 ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ 2 ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ 1 [0093] where ⁇ ⁇ ⁇ is reconstructed from ⁇ ⁇ ⁇ with zero-order hold or linear interpolation, ⁇ ⁇ ⁇ / ⁇ , and ⁇ is the sampling rate. In some cases, the computationally complexity of additive synthesis can be reduced with approximations.
  • the impulse train generator 514 or other component (e.g., a processor) of the voice decoding signal synthesis system 500 can round the fundamental periods to the nearest multiples of the sampling period.
  • the discrete impulse train is sparse.
  • the impulse train generator 514 can then generate the impulse train sequentially (e.g., one pitch mark at a time).
  • the pitch tracker 510 can process the input ⁇ for the time instant n to generate the frame-wise fundamental frequency ⁇ output, which is provided to and processed by the impulse train generator 514 of the voice decoding signal synthesis system 500.
  • the random noise generator 516 of the voice decoding signal synthesis system 500 can sample a noise signal ⁇ from a Gaussian distribution.
  • the neural network filter estimator 502 can estimate impulse responses h ⁇ ⁇ , ⁇ and h ⁇ ⁇ , ⁇ for each frame, given the log-Mel spectrogram ⁇ , ⁇ extracted from the input ⁇ by the feature extraction engine 512 of the encoder.
  • complex cepstrums (h ⁇ ⁇ and h ⁇ ⁇ ) can be used as the internal description of impulse responses (h ⁇ and h ⁇ ) for the neural network filter estimator 502.
  • Complex cepstrums describe the magnitude response and the group delay of filters simultaneously. The group delay of filters affects the timbre of speech.
  • the neural network filter estimator 502 can use mixed-phase filters, with phase characteristics learned from the dataset. Qualcomm Ref. No.2400252WO [0096]
  • the length of a complex cepstrum can be restricted, essentially restricting the levels of detail in the magnitude and phase response. Restricting the length of a complex cepstrum can be used to control the complexity of the filters.
  • the neural network filter estimator 502 can predicts low-frequency coefficients, in which the high-frequency cepstrum coefficients can be set to zero. In one illustrative example, two 10 millisecond (ms) long complex cepstrums are predicted in each frame.
  • the neural network filter estimator 502 can use a discrete Fourier transform (DFT) and an inverse-DFT (IDFT) to generate the impulse responses h ⁇ and h ⁇ .
  • the neural network filter estimator 502 can approximate an infinite impulse response (IIR) (h ⁇ ⁇ , ⁇ and h ⁇ ⁇ , ⁇ ) using finite impulse responses (FIRs).
  • the harmonic LTV filter 518 can filter the impulse train ⁇ from the impulse train generator 514 to generate a harmonic component ⁇ ⁇ ⁇ .
  • the noise LTV filter 520 can filter the noise signal ⁇ to generate a noise component ⁇ ⁇ ⁇ .
  • the voice decoding signal synthesis system 500 can combine (e.g., by summing/adding or otherwise combining) the output of the harmonic LTV filter 518 (the harmonic component ⁇ ⁇ ⁇ ) and the output of the noise LTV filter 520 (the noise component ⁇ ⁇ ⁇ ) can be combined (e.g., summed or otherwise combined) to obtain the excitation signal ⁇ .
  • a computing device may receive audio information via an input device, such as a microphone. In some cases, the received audio information may be sampled at a point in time, or frame.
  • FIG.6 illustrates an audio waveform 600 representing audio information received by a wireless device, in accordance with aspects of the present disclosure.
  • the audio waveform 600 may include regions, such as an active speech region 602, a hangover region 604, and an inactive (e.g., silent) region 606.
  • a speech encoder may include a voice activity detection (VAD) functionality which analyzes the received audio information to determine whether there is speech activity during a particular audio frame.
  • VAD voice activity detection
  • the encoder may also include an indication in the encoded audio data that the encoded audio information includes active speech.
  • an end of the active speech region 602 may correspond to when the speech stops. [0100]
  • the hangover region 604 may follow the active speech region 602.
  • the hangover region 606 may be the region where the speech transitions from active to inactive. For example, when speech stops, the VAD may indicate that speech is no longer being detected. In some cases, an indication from VAD that speech is no longer occurring can be abrupt and result in clipped or abruptly ends of sentences/words, which may sound unnatural or harsh. To avoid this, the encoder may continue to encode the audio information to encoded audio data in the hangover region 604 to allow long and/or lingering sounds (e.g., residual speech which may not be detected as speech by the VAD) to be captured. The encoder may also include an indication in the encoded audio data that the encoded audio information does not include active speech (e.g., is in the hangover region 604).
  • the encoded audio information may include some residual speech, but may primarily include background noise, especially toward the later portions of the hangover region 604.
  • the audio encoder may have a predefined and/or dynamic/adaptive time period after speech has stopped (e.g., no longer detected by the VAD) for the hangover region 604.
  • the hangover region 604 may end.
  • the inactive region 606 may follow the hangover region 604.
  • the encoder may encode captured audio information in a SID.
  • the SID may be transmitted to the other wireless device to be used by the other device to generate comfort noise (e.g., background noise) while in the inactive region 606.
  • a number of SID frames generated/transmitted may be reduced as compared to a number of frames generated/transmitted while in the active speech region 602 and the hangover region 604.
  • certain audio encoders may support adaptive and/or fixed SID intervals where a single SID frame may be transmitted instead of some number of regular audio frames (e.g., when in the active speech region 602 and/or the hangover region 604).
  • EVS enhanced voice services
  • one SID frame may be transmitted during a time period where 8-50 regular audio frames would have been transmitted. While the number of SID frames may be reduced as compared to regular audio frames, Qualcomm Ref. No.2400252WO SID frames are still being regularly transmitted.
  • FIG.7 is a block diagram providing an overview of a technique for decoder silence generation without, or with minimal, coded SIDs 700, in accordance with aspects of the present disclosure.
  • a wireless device may receive an encoded audio data stream and a decoder 701 of the wireless device may decode an audio frame of the encoded audio data stream to obtain decoded active speech from an active region 702 (e.g., of the audio data stream).
  • the decoded active speech from the active region 702 may include speech as well as background noise.
  • the decoded active speech from the active region 702 may be passed to a denoiser 704.
  • the denoiser 704 may filter out the background noise to generate filtered active speech.
  • the filtered active speech may then be subtracted 706 from the decoded active speech from the active region 702 to obtain background noise from the decoded active speech from the active region 702.
  • the background noise may be passed to a spectrum estimator 708.
  • the spectrum estimator may estimate information about the background noise 710 such as a spectral shape (S(t)) (e.g., spectral characteristics) and/or gain (G(t)).
  • S(t) spectral shape
  • G(t) gain
  • the information about the background noise 710 may be passed to an inactive synthesizer 712, which may generate synthesized background noise 714 based on the information about the background noise 710 from the obtained decoded active speech from an active region 702.
  • the encoder may indicate a transition to the inactive region (e.g., marking whether an encoded frame include detected speech or not) without indicating a hangover region.
  • the decoder 701 may stop receiving encoded audio frames (e.g., from another wireless device) when speech is no longer present.
  • the decoder 701 may receive one SID frame indicating that speech is no longer present. In another example, the decoder 701 may receive an indication that speech is no longer present without also receiving encoded audio information (e.g., a description of the silence in the SID frame).
  • the wireless device performing the decoding may receive a radio access network (RAN) message, such as an RRC message, indicating that speech has ended.
  • RAN radio access network
  • the inactive synthesizer 712 may use the information about the background noise 710 from a number of frames (N) prior to the indicated transition to the inactive region to generate the synthesized background noise 714.
  • the decoder 701 may also take in account the hangover region. For example, where the encoder may encode an indication that an audio frame is from a hangover region. The decoder 701 may decode the encoded audio frame to obtain decoded audio frame from the hangover region 716. A classifier 703, based on the indication that the audio frame is from the hangover region, may pass the decoded audio frame from the hangover region 716 to the spectrum estimator 708 to generate information about the background noise 710. Generally, while the hangover region may include some residual speech, the hangover region may primarily include background noise, especially towards an end of the hangover period. [0105] In some cases, then hangover region 716 may be inferred.
  • the encoder may indicate a transition to the inactive region without indicating the hangover region.
  • the decoder 701 and/or classifier 703 may infer that (M) frames prior to the indicated transition to the inactive region comprise the hangover region.
  • the decoded active speech corresponding to those M frames may be input to the spectrum estimator 708 to generate information about the background noise 710.
  • Information about the background noise 710 corresponding to frames from the active region and hangover region may be input to the inactive synthesizer 712.
  • the inactive synthesizer 712 may then generate synthesized background noise 714 based on the information about the background noise 710 in the frames from the active region and/or hangover region.
  • the inactive synthesizer 712 may use the information about the background noise 710 from M frames from the hangover region and N-M frames from the active region.
  • the inactive synthesizer 712 may synthesize the background noise 714 using any technique for synthesizing background noise 714 based on information about the background noise 710.
  • FIG. 8A is a block diagram illustrating an example inactive synthesizer 800, in accordance with aspects of the present disclosure.
  • the inactive synthesizer 800 may correspond with inactive synthesizer 712 of FIG.7.
  • the inactive synthesizer 800 may use a set of gated recurrent units (GRUs) that accept time series information (e.g., information about the background noise in time aligned audio frames) to predict (e.g., generate) the synthesized background noise.
  • GRUs gated recurrent units
  • a first GRU 802 may receive a first input frame 804 and a second Qualcomm Ref. No.2400252WO input frame 806.
  • the first input frame 804 may be an audio frame just prior to the transition (e.g., where the transition occurs at time t, the first input frame 804 may occur at time t-1) to the inactive region (e.g., from the hangover region).
  • the second input frame 806 may be an audio frame prior to the first input frame 804 (e.g., from t-2, which may be from the hangover region or the active region).
  • the first input frame 804 and second input frame 806 may be processed by the first GRU 802 to generate (e.g., predict) a first synthesized noise frame 808 for the inactive region (e.g., for time t).
  • the first synthesized noise frame 808 may be information about the synthesized noise, such as a spectral shape (e.g., spectral characteristics) and/or gain and the actual background noise may be generated based on the information about the synthesized noise, such as by a digital to analog convertor.
  • the first synthesized noise frame 808 may be input to a second GRU 810 along with a third input frame 812.
  • the third input frame 812 may be an audio frame prior to the second input frame 806 (e.g., from time t-3, which may be from the hangover region or the active region).
  • the first synthesized noise frame 808 and the third input frame 812 may be processed by the second GRU 810 to generate a second synthesized noise frame 814 for the inactive region (e.g., for time t+1).
  • the second synthesized noise frame 814 may be input, along with a fourth input frame 816 (e.g., from time t-4, which may be from the hangover region or the active region), may be input to a third GRU 818 to generate a third synthesized noise frame 820 (e.g., for time t+2). This process may be repeated to continue generating synthesized noise frames for the inactive region.
  • a fourth input frame 816 e.g., from time t-4, which may be from the hangover region or the active region
  • a third GRU 818 to generate a third synthesized noise frame 820 (e.g., for time t+2).
  • This process may be repeated to continue generating synthesized noise frames for the inactive region.
  • the SID may be input to the inactive synthesizer and used to generate synthesized noise frames, for example, as the first input frame 804.
  • FIG.8B is a is a block diagram illustrating another example inactive synthesizer 850, in
  • the inactive synthesizer 850 may correspond with inactive synthesizer 712 of FIG.7.
  • the inactive synthesizer 850 may use a set of neural network (NN) based filter estimators to shape random noise.
  • NN neural network
  • a first NN based filter estimator 852 may receive a first input frame 854.
  • the first input frame 854 may be an audio frame just prior to the transition (e.g., where the transition occurs at time t, the first input frame 854 may occur at time t-1) to the inactive region (e.g., from the hangover region).
  • the first NN based filter estimator 852 may estimate a filter based on the first input frame 854.
  • Random noise 856 may be passed through a noise filter 858, which may color (e.g., introduce a spectral shape to) the random noise, resulting in filtered random noise.
  • This filtered random noise may be further shaped by the estimated filter from the first input frame to generate a first synthesized noise frame 860 for the inactive region (e.g., for time t).
  • the first synthesized noise frame 860 may be input to a second NN based filter estimator 862 along with the estimated filter of the first NN based filter estimator 852, and a second input frame 864 to estimate a filter.
  • the second input frame 864 may be an audio frame prior to the first input frame 854 (e.g., from t-2, which may be from the hangover region or the active region).
  • the estimated filter of the second NN based filter estimator 862 may be used to shape random noise 865 passed through a noise filter 866 to generate a second synthesized noise frame 868 for the inactive region (e.g., for time t+1).
  • the second synthesized noise frame 868, a third input frame 870 (e.g., from t-3), and the estimated filter of the second NN based filter estimator 862 may be input to a third NN filter estimator 872 and used to generate a third synthesized noise frame 874 in a manner substantially similar to that used to generate the second synthesized noise frame 868. This process may be repeated to continue generating synthesized noise frames for the inactive region.
  • the SID may be input to the inactive synthesizer and used to generate synthesized noise frames, for example, as the first input frame 854.
  • FIG.9A is a block diagram illustrating a denoiser 900, in accordance with aspects of the present disclosure.
  • the denoiser 900 may correspond with denoiser 704 of FIG. 7.
  • the denoiser 900 may be implemented based on non-negative matrix factorization (NMF).
  • the denoiser 900 may include a pretrained speech dictionary 902.
  • a speech dictionary 902 and a noise dictionary may be trained as a part of training.
  • the pretrained speech dictionary 902 may include speech-like characteristics (e.g., spectral characteristics that represent parts of speech), while a noise dictionary may include spectral characteristics that represent noise.
  • a product of the pretrained speech dictionary 902 and an activation function 904 may represent the speech portion of the input audio data (e.g., decoded active speech from the active region 702 of FIG. 7) as represented by spectrogram 906.
  • FIG. 9B is a block diagram illustrating NMF denoising 950, in accordance with aspects of the present disclosure.
  • an input y(t) 952 may include a speech component s(t) and a noise component n(t).
  • a short-time-Fourier-transform 954 may be applied to the input 952 and a magnitude 956 determined.
  • Activations 958 may be estimated based on a noise dictionary 960 and a speech dictionary 962 to determine portions of the input 952 which may be represented by the noise dictionary 960 and other portions that may be represented by the speech dictionary 962.
  • the noise portion may then be estimated as a product of the noise dictionary 960 and the noise activation 958.
  • speech portion may be estimated as a product of the speech dictionary 962 and the speech activation 958.
  • the speech portion may be returned by the denoiser 900 and subtracted from the original decoded active speech (e.g., subtracted 706 of FIG.7).
  • FIG.10 is a block diagram illustrating a technique for location aware background noise synthetization 1000, in accordance with aspects of the present disclosure.
  • location information 1002 for another wireless device sending the encoded audio data stream may be included as a part of, or in addition to, the encoded audio data stream.
  • the location information 1002 may indicate where the other wireless device is located, such as at a train station, on a busy street, at the beach, etc.
  • the location information 1002 may be input to an ambient noise prediction engine 1004.
  • the ambient noise prediction engine 1004 may predict a type of ambient noise (e.g., background noise) 1006 that may be present based on the location information 1002.
  • the predicted type of ambient noise 1006 may be input to an inactive synthesizer 1008.
  • the inactive synthesizer 1008 may correspond to inactive synthesizer 712 of FIG.7.
  • the inactive synthesizer 712 may take into account the predicted type of ambient noise 1006 when generating the synthesized background noise.
  • FIG. 11 is a block diagram illustrating a technique for random background noise synthetization 1100, in accordance with aspects of the present disclosure. In FIG.
  • a random noise generator 1102 may be used to generate random noise based on, for example, a random number generator.
  • the random noise from the random noise generator 1102 may be input to an Qualcomm Ref. No.2400252WO inactive synthesizer 1104.
  • the inactive synthesizer 1104 may correspond to inactive synthesizer 712 of FIG.7.
  • inactive synthesizer 1104 may use the random noise from the random noise generator 1102 as the synthesized background noise.
  • the inactive synthesizer 1104 may mix the random noise from the random noise generator 1102 with synthesized background noise based on the information about the background noise from an active region and/or hangover region.
  • a spectrum estimator such as spectrum estimator 708 of FIG.7, may estimate information about the background noise such as the spectral shape and/or gain.
  • the spectrum estimator may be configured to perform additional spectral characteristic measurements.
  • FIG. 12 illustrates a spectrum estimator 1200 for measuring additional spectral characteristics, in accordance with aspects of the present disclosure.
  • a spectrum estimator 1200 may include a spectrum magnitude engine 1204, a dominant spectral sample calculator 1206, and a spectral descriptor calculator.
  • Noise 1202 e.g., from a denoiser or from decoded audio from the hangover region
  • the spectrum magnitude engine 1204 may generate a Magnitude spectrum of the noise.
  • the Magnitude spectrum of the noise may describe the noise using frequency and amplitude.
  • the Magnitude spectrum may be input to the dominant spectral sample calculator 1206 and the spectral descriptor calculator 1208.
  • the dominant spectral sample calculator 1206 apply a probability density function to the Magnitude spectrum and spectral samples higher than a m th percentile may be selected.
  • the dominant spectral sample calculator 1206 may determine k-disjoint contiguous maximum sum segments of spectrums to indicate k-contiguous dominant spectral locations.
  • the dominant spectral locations may be passed to the spectral descriptor calculator 1208.
  • spectral characteristics for noise may be, at least in part, determined on dominant spectrums for increased accuracy.
  • the spectral descriptor calculator 1208 may generate additional descriptions 1210 of the structural features of the noise for the dominant spectrums. For example, the spectral descriptor calculator 1208 may determine first-n moments for the dominant spectrum where the first moment may indicate an energy concentration frequency, the second moment may indicate bandwidth around a centroid, skewness (indicating asymmetry), and kurtosis (indicating peakiness) for output Qualcomm Ref. No.2400252WO as a part of the additional descriptions 1210.
  • the spectral descriptor calculator 1208 may also determine Mel-band/critical band energies of dominant spectral samples, perform auto-regressive (AR)/AR integrated moving average (ARIMA) modeling of the Magnitude spectrums, and determine a spectral slope using spectral entropy of the Magnitude spectrums for output as a part of the additional descriptions 1210.
  • FIG.13 is a block diagram illustrating another technique for decoder silence generation without, or with minimal, coded SIDs 1300, in accordance with aspects of the present disclosure.
  • FIG.13 is a block diagram illustrating another technique for decoder silence generation without, or with minimal, coded SIDs 1300, in accordance with aspects of the present disclosure.
  • decoded active speech from an active region 1302 may be passed into a noise extractor or ML model 1304.
  • a noise extractor may operate in a manner similar to that described with respect to FIGs.9A and 9B.
  • a noise dictionary may be used to extract a noise portion from the decoded active speech from the active region 1302. This noise portion may be passed to the spectrum estimator 1308 to estimate information about the noise portion.
  • the decoded active speech from the active region 1302 may be passed to a ML model 1304.
  • the ML model 1304 may be trained to subtract speech or isolate background noise from audio frames.
  • ML model 1304 may be a neural network, deep learning network, convolutional neural network, or any other type of machine learning network.
  • a single ML model 1304 may be used, or multiple ML models 1304 may be used.
  • the ML model 1304 may output background noise from the decoded active speech from the active region 1302 to the spectrum estimator 1308 to estimate information about the noise portion.
  • FIG.14 is signal diagram 1400 illustrating signals for decoder silence generation without coded silence descriptions, in accordance with aspects of the present disclosure.
  • a wireless node 1404 may configure a wireless device, such as a sending UE 1402 and receiving UE 1406 with timing information, wireless resources, etc. to help optimize scheduling.
  • the UE may indicate to the wireless node that no SIDs will be sent.
  • a voice bearer may be established 1408 as between a sending UE 1402 (e.g., UE encoding Qualcomm Ref. No.2400252WO audio data for transmission) and a receiving UE 1406 (e.g., UE decoding received audio data) via the wireless node 1404.
  • a sending UE 1402 may send UE assistance information (UAI) 1410 to the wireless node 1404 indicating that the sending UE 1402 is not sending SIDs. Sending an indication to the wireless node 1404 may be useful to help with scheduling, ensuring compatibility, etc.
  • the wireless node 1404 may transmit a message 1412 (e.g., RRC message or other control plane message) to the receiving UE 1406 indicating that the sending UE 1402 is not sending SIDs.
  • the sending UE 1402 enter a talk spurt and may transmit a set of voice packets 1414 to the receiving UE 1406.
  • signaling indicating that SIDs are not being used and/or signaling that SIDs should be used may be transmitted via any RAN protocol and this transmission may utilize a new message/packet data unit (PDU) or utilize an existing message/PDU.
  • PDU packet/packet data unit
  • the sending UE 1402 may transmit a UAI message indicating an end of the talk spurt 1418 to the wireless node 1404.
  • the end of the talk spurt may be after a hangover period. In other cases, the end of the talk spurt may be after speech stops being detected by the sending UE 1402 (e.g., without a hangover period).
  • the wireless node 1404 may transmit a message 1420 (e.g., RRC message) to the receiving UE 1406 indicating that the talk spurt has ended. In some cases, the message 1420 may indicate to the receiving UE 1406, that an inactive region has begun.
  • sending a RAN message such as an RRC message indicating that a talk spurt has ended (e.g., an inactive region has been reached) may be a smaller sized message without a description of the background noise sent on a control plane as compared to a SID, which may be sent as a user plane message.
  • messages sent on the user plane may be intended for the operation of user applications.
  • Messages sent on the control plane may carry RAN messages for establishing, managing, and or controlling the different components of the wireless network.
  • the encoder may include an indication in packets from the hangover period indicating an end of the Qualcomm Ref.
  • No.2400252WO talk spurt/entry to the hangover period.
  • the decoder upon receiving an indication of the hangover period, can assume that if the packets stop, it's because the talk spurt has ended (e.g., in the hangover period). Otherwise, if the decoder does not receive the indication of the hangover period and the packets stop, then the decoder may treat this stoppage as a signal loss.
  • activation of SIDs may be configured, activated, and/or deactivated by the wireless node 1404. For example, certain UEs may not support decoder silence generation and the wireless node 1404 or other UE may indicate to the sender UE 1402 to send SIDs.
  • FIG. 15 is a block diagram illustrating an example of an audio codec system 1500 that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer.
  • the audio codec system 1500 can also be referred to as a voice coding system (e.g., vocoder), a signal synthesis system, a voice coding signal synthesis system, etc.
  • the audio codec system 1500 can be a generative audio codec (e.g., a generative voice codec) including one or more feedback recurrent autoencoders (FRAEs) that may be used to generate encoded audio data, and one or more neural synthesizers that may be used to generate reconstructed audio data based on the encoded audio data.
  • the audio codec system 1500 e.g., a generative voice codec
  • the encoder 1505 can be used to generate encoded features that are transmitted to the decoder 1510 over a channel 1540.
  • the encoder 1505 can receive audio 1515 and generate encoded features corresponding to the audio 1515.
  • the encoded features can be generated using a feedback recurrent autoencoder (FRAE) 1525 included in the encoder 1505.
  • FRAE feedback recurrent autoencoder
  • the encoder 1505 can include one or more FRAEs, where each FRAE of the one or more FRAEs is configured to generate a corresponding one or more encoded features associated with the audio 1515.
  • the audio 1515 may include and/or comprise a speech signal, a voice signal, etc.
  • the decoder 1510 can receive encoded features (e.g., from the encoder 1505, over the channel 1540) and generate reconstructed audio 1555 based at least in part on the encoded features received from the encoder 1505.
  • the reconstructed audio 1555 may include and/or comprise a Qualcomm Ref. No.2400252WO reconstructed speech signal, a reconstructed voice signal, etc.
  • the reconstructed audio 1555 can be generated using a neural network-based speech synthesizer 1550 (e.g., also referred to as a neural speech synthesizer and/or a neural synthesizer) included in the decoder 1510.
  • the reconstructed audio 1555 generated using the neural speech synthesizer 1550 may also be referred to as synthesized speech.
  • the decoder 1510 can include a FRAE decoder 1545, which can be used to decode the encoded features received from the encoder 1505.
  • the FRAE decoder 1545 can be the same as or similar to a decoder implemented by the FRAE 1525 included in the encoder 1505.
  • the decoder 1510 can include one or more FRAE decoders 1545, which can correspond to one or more FRAE autoencoders 1525 included in the encoder 1505.
  • the number of FRAE decoders 1545 included in the decoder 1510 can be equal to the number of FRAE autoencoders 1525 included in the encoder 1505.
  • the audio 1515 includes speech.
  • the audio 1515 may be an example of the speech signal 101 of FIG.1, and/or the speech signal 201 of FIG.2, etc.
  • the encoder 1505 of the audio codec system 1500 can be configured to generate one or more encoded representations of the input audio 1515.
  • the one or more encoded representations can include spectral envelope information associated with the audio 1515, and/or can include pitch information associated with the audio 1515, etc.
  • the encoder 1505 can extract spectral envelope features from the audio 1515 using spectral envelope feature extraction engine 1520, and can encode the extracted spectral envelope features using the FRAE 1525.
  • the encoded spectral features z can be transmitted from the FRAE 1525 to the decoder 1510, using the channel 1540.
  • the encoder 1505 can extract pitch information from the audio 1515 using pitch extraction engine 1530, and can perform quantization 1535 to generate quantized pitch information associated with the audio 1515.
  • the quantized pitch information can be transmitted from the pitch quantizer 1535 to the decoder 1510, using the channel 1540.
  • the encoder 1505 can be configured to extract spectral features (e.g., spectral envelope features) from the audio 1515 using the spectral envelope feature extraction engine 1520.
  • the spectral envelope feature extraction engine 1520 can perform a cepstrum computation to determine cepstrum information and/or one or more cepstral coefficients corresponding to the audio 1515.
  • the spectral features extracted from the audio 1515 using the spectral envelope feature extraction engine 1520 can include mel- frequency cepstral coefficients (MFCC), such as MFCC-24 features (e.g., 24-dimensional MFCC Qualcomm Ref. No.2400252WO features).
  • MFCC mel- frequency cepstral coefficients
  • the spectral envelope feature extraction engine 1520 applies one or more mel-scaled filter(s) (e.g., a mel-scaled filterbank), a logarithmic compression, and/or a discrete cosine transform (DCT) to the audio 1515 (e.g., and/or to a magnitude spectrum associated with the audio 1515).
  • the encoder 1505 processes the extracted spectral features using a feedback recurrent autoencoder (FRAE) 1525 to generate encoded features z.
  • the FRAE 1525 includes a decoder and an encoder.
  • the FRAE 1525 can implement feedback of state information h between the decoder and the encoder included in the FRAE 1525.
  • the state information h can be determined or obtained at the decoder of the FRAE 1525, and feedback of the state information h can be performed for the decoder of the FRAE 1525 (e.g., the state information h is fed back to the decoder of the FRAE 1525) and for the encoder of the FRAE 1525 (e.g., the state information h is fed back from the decoder of the FRAE 1525 to the encoder of the FRAE 1525).
  • the spectral envelope feature extraction engine 420 can be configured to generate and/or determine (e.g., extract) one or more types of spectral envelope features.
  • the extracted spectral envelope features may include cepstrum information and/or cepstral coefficients.
  • cepstral liftering can be performed based on DCT truncation to exclude or remove pitch information and only capture spectral envelope information in the extracted cepstrum or cepstral coefficients.
  • the extracted spectral envelope features can include companded (e.g., log) filterbank energies associated with the input audio 415.
  • companded filterbank energies can be determined based on applying an inverse DCT (e.g., IDCT) to a liftered cepstrum.
  • the extracted spectral envelope features can include filterbank energies determined based on uncompanding (e.g., exp) the companded filterbank energies.
  • the extracted spectral envelope features can include a full resolution spectrum (e.g., DFT domain with companded (e.g., log, linear amplitude, etc.) information, etc.). The full resolution spectrum can be smoothed to the envelope of the spectrum, for example based on interpolating the companded or uncompanded filterbank energies.
  • the extracted spectral envelope features can include one or more linear prediction (LP) coefficients, for example determined based on processing one or more speech frames of the input audio 415 using the Levinson-Durbin algorithm (e.g., autocorrelation technique), and/or using a covariance technique.
  • the extracted spectral envelope features can include one or more of line spectral frequency (LSF) information and/or line spectral pair (LSP) information.
  • LSF line spectral frequency
  • LSP line spectral pair
  • the encoded features or encoded feature information transmitted from the generative voice codec system encoder 405 can comprise a latent representation z between the encoder of the FRAE 425 and the decoder of the FRAE 425.
  • the encoder 405 can be configured to pass (e.g., transmit) the encoded features z through a channel 440 to the decoder 410.
  • the encoder 405 also processes the audio 415 using a pitch extraction engine 430 to extract pitch information from the audio 415.
  • the audio 415 can be processed in parallel by the spectral envelope feature extraction engine 420 and the pitch extraction engine 430.
  • the encoder 405 can include a denoiser 418 that processes the input audio 415 and provides a de-noised audio to the spectral envelope feature extraction engine 420 and/or to the pitch extraction engine 430.
  • the pitch extraction engine 430 can generate one or more types of pitch information.
  • the pitch extraction engine 430 can output pitch information in the frequency-domain (e.g., f 0 pitch information of the audio 415, in units of Hertz (Hz)) and/or can output pitch information in the time-domain (e.g., pitch lag in samples, with or without a fractional component, and/or pitch lag in milliseconds).
  • the pitch extraction engine 430 can be used to generate pitch estimation indicative of a pitch lag (or pitch delay) from the audio 415, and/or to identify a pitch correlation from the audio 415.
  • the pitch information generated by the pitch extraction engine 430 can include a pitch lag and a pitch correlation.
  • the encoder 405 can be configured to processes the pitch information of the audio 415 (e.g., pitch, pitch lag, and/or pitch correlation, determined using the pitch extraction engine 430) using a quantizer 435 Q() to generate a quantized pitch signal.
  • the encoder 405 passes the quantized pitch signal through the channel 440 to the decoder 410.
  • the quantizer 435 Q() may also be referred to as a pitch quantizer.
  • the quantizer 435 Q() can perform vector quantization (VQ) to generate the quantized pitch signal for transmission to the decoder 410 over the channel 440.
  • the quantizer 435 Q() can perform single- stage VQ and/or can perform multi-stage VQ (e.g., MSVQ). In some cases, the quantizer 435 Q() can be implemented using one or more FRAEs. In some aspects, the quantizer 435 Q() can be configured to implement forward error correction (FEC) for the quantized pitch signal that is transmitted to the decoder 410 over the channel 440.
  • FEC forward error correction
  • the quantizer 435 Q() can Qualcomm Ref. No.2400252WO implement multiple description coding (MDC) for the transmission of the quantized pitch signal over the channel 440, can implement full-redundancy FEC for the transmission of the quantized pitch signal over the channel 440, etc.
  • MDC multiple description coding
  • the generative voice codec system decoder 410 can be configured to receive the encoded features z from the FRAE 425 included in the generative voice codec system encoder 405. For example, the decoder 410 can receive the encoded features z via the channel 440. The decoder 410 decodes the encoded features z using a FRAE 445 to generate decoded features. In some examples, the FRAE 445 of the decoder 410 includes only a decoder, without an encoder. In some examples, the FRAE 445 of the decoder 410 can include an encoder.
  • the decoded features generated by the FRAE 445 include mel-frequency cepstral coefficients (MFCC), such as MFCC- 24 features (e.g., 24-dimensional MFCC features).
  • MFCC mel-frequency cepstral coefficients
  • the decoded features determined using the FRAE decoder 445 can be provided to a neural speech synthesizer 450, which can be configured to generate a reconstructed audio 455 (e.g., synthesized speech) based at least in part on the decoded spectral envelope features from the FRAE decoder 445.
  • the neural speech synthesizer 450 receives as input the decoded spectral envelope features (e.g., from the FRAE decoder 445) and the received pitch encoding information (e.g., the quantized pitch signal received by the decoder 410 over the channel 440 and from the pitch quantizer 435 of the encoder 405).
  • the neural speech synthesizer 450 can generate the reconstructed audio 455 (e.g., synthesized speech) based on processing the decoded spectral envelope features along with a pitch signal (e.g., the quantized pitch signal or a reconstructed variant thereof).
  • the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and/or the pitch correlation of the audio 415) from the encoder 405 via the channel 440.
  • the decoder 410 passes the quantized pitch signal to the neural speech synthesizer 450, and the neural speech synthesizer 450 processes the decoded spectral envelope features along with the quantized pitch signal to generate reconstructed audio 455 (e.g., synthesized speech).
  • the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and/or the pitch correlation of the audio 415) from the encoder 405 via the channel 440, and can perform dequantization of the quantized pitch signal to obtain a reconstructed Qualcomm Ref. No.2400252WO pitch signal.
  • the reconstructed pitch signal can be passed to the neural speech synthesizer, and used in combination with the decoded spectral envelope features from the FRAE decoder 445 to generate the reconstructed audio 455 (e.g., synthesized speech).
  • the decoder 410 can include a reconstruction engine that is separate from the neural speech synthesizer 450.
  • the reconstruction engine processes the quantized pitch signal to reconstruct the pitch signal (e.g., including the pitch, the pitch lag, and/or the pitch correlation) before the neural speech synthesizer 450 using the reconstructed pitch signal (e.g., the pitch, the pitch lag, and/or the pitch correlation) to generate the reconstructed audio 455.
  • the output of the neural speech synthesizer 450 can be provided to one or more linear predictive coding (LPC) layers 452 included in the decoder 410.
  • LPC 452 can be used to perform linear prediction analysis and/or linear prediction synthesis, based on the output of the neural speech synthesizer 450.
  • the LPC 452 can perform linear prediction based on the output of the neural speech synthesizer 450, to generate the reconstructed audio 455 (e.g., synthesized speech).
  • the decoder 410 does not include the LPC 452, and the neural speech synthesizer 450 can be trained and/or configured to generate the reconstructed audio 455 directly (e.g., the output of the neural speech synthesizer 450 can be the reconstructed audio 455).
  • the LPC layers 452 can be used to implement a linear prediction (LP) synthesis filter, based on: ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ [0139]
  • represents (e.g., the input signal to the LPC layers 452)
  • represents the output signal (e.g., the output of the LPC layers 452, which can be the reconstructed audio 455)
  • ⁇ ⁇ represents the linear prediction coefficients associated with implementing the LP synthesis filter
  • is a value corresponding to the LP filter order.
  • the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.
  • the LP synthesis filter associated with the LPC layers 452 can be implemented in the time domain or in the frequency domain.
  • the Qualcomm Ref. No.2400252WO LP synthesis filter can be implemented based on a difference equation.
  • the LP synthesis filter can be implemented in the time domain by convolving the input signal x[n] with the LP filter impulse response or an approximation of the LP filter impulse response.
  • the LP filter can be associated with an infinite impulse response (IIR), which can be approximated using a finite segment of the IIR (e.g., such as the first N samples of the IIR for a configured integer value of N, etc.).
  • IIR infinite impulse response
  • the LP synthesis filter associated with the LPC layers 452 can be implemented in the frequency domain based on multiplying the FFT of the input signal x[n] with the frequency response of the LP filter, and then determining the IFFT of the result to convert the output of the frequency domain LP synthesis filter into a time domain signal corresponding to the reconstructed audio 455.
  • the LP coefficients ⁇ ⁇ can be estimated from the decoded spectral envelope features obtained using the FRAE decoder 445.
  • the spectral envelope features can be LP coefficients (e.g., determined by the spectral envelope feature extraction engine 420 based on using the Levinson-Durbin algorithm or other autocorrelation technique, or using a covariance technique, to process the input speech frames of the audio 415), and the decoded LP coefficients on the decoder side can be used for the LP filter implemented by the LPC layers 452.
  • the extracted spectral envelope features may be LSFs or LSPs, and the decoded LSFs or LSPs obtained using the FRAE decoder 445 can be converted to the corresponding LP coefficients ⁇ ⁇ .
  • the LP coefficients ⁇ ⁇ of the LPC layers 452 can be estimated from the spectral envelope features using a neural network-based and/or a DSP-based technique.
  • the extracted spectral envelope features may comprise a cepstrum, and a DSP-based technique can be used to convert the decoded cepstrum into filterbank energies, and subsequently interpolate the filterbank energies to estimate an FFT square magnitude at all FFT bins (e.g., including bins beyond the centers of filterbank filters).
  • the DSP-based technique can include applying an IFFT to obtain an estimate of the autocorrelation sequence, and utilizing the Levinson- Durbin algorithm on the estimated autocorrelation sequence to obtain the LP coefficients ⁇ ⁇ .
  • the spectral envelope feature extraction 420 can be performed based at least in part on spectral feature learning.
  • spectral feature learning can be performed by one or more machine learning models configured and/or trained to generate as output the one Qualcomm Ref. No.2400252WO or more extracted spectral envelope features.
  • a cepstrum or cepstrum information may be an optional input to a machine learning spectral feature learning model used to implement the spectral envelope feature extraction 420.
  • the spectral envelope feature extraction 420 can be implemented using one or more machine learning spectral feature learning models or engines, which can be jointly trained with the neural speech synthesizer 450, an NHV used to implement the neural speech synthesizer (e.g., such as the NHV-based neural speech synthesizer 650 of FIG.6), an LPC network used to implement the neural speech synthesizer (e.g., such as the LPC network-based neural speech synthesizer 750 of FIG.7), etc.
  • joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on a short-time Fourier transform (STFT) loss and/or STFT loss function.
  • STFT short-time Fourier transform
  • joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on an adversarial and/or generative adversarial network (GAN) loss.
  • GAN generative adversarial network
  • joint training can be performed using a multi-resolution STFT loss, based on calculating STFT amplitude spectrograms from ground truth speech information and the synthesized speech information 455.
  • the multi-resolution STFT loss can then be determined as a sum of mean absolute error and mean absolute error in the log domain.
  • the STFT amplitude spectrograms can be calculated at different window lengths to capture representations of the error and/or loss at different time and/or frequency resolutions.
  • joint training based on adversarial or GAN loss can be used to learn temporal fine structures in speech signals.
  • An adversarial or GAN loss can be used to match the distribution of real speech and the synthesized speech 455.
  • the neural speech synthesizer 450 (and/or an NHV-based neural speech synthesizer) can be used as the generator network.
  • a separate discriminator network can be used to train based on the adversarial loss, with both the NHV and the discriminator jointly trained using adversarial training techniques.
  • the discriminator network can be implemented as a classifier configured to classify whether an input audio represents real speech or synthesized speech.
  • the generator network can attempt to fool the discriminator network to classify synthesized speech as real speech. Over the course of training, the generator network improves and begins producing (e.g., generating or synthesizing) speech that is very similar to the real speech.
  • the discriminator network can be implemented using WaveNet. Qualcomm Ref. No.2400252WO [0145]
  • the neural speech synthesizer 450 can be implemented using one or more trained neural networks. For example, the one or more trained neural networks can be trained to perform speech synthesis and/or signal synthesis.
  • the neural speech synthesizer 450 can be implemented using or based on an LPCNet machine learning architecture, a WaveNet machine learning architecture, a WaveRNN machine learning architecture, etc.
  • the neural speech synthesizer 450 can be provided as a neural homomorphic vocoder (NHV).
  • FIG. 16 illustrates an example of an audio codec system 1600 (e.g., generative voice codec) that includes a decoder 1610 configured to implement an NHV-based neural speech synthesizer with a technique for random background noise synthesis, in accordance with aspects of the present disclosure.
  • an NHV is a type of neural vocoder that can synthesize speech with source-filter models controlled by one or more neural networks.
  • An NHV-based neural vocoder may include one or more neural networks in a source-filter model that can synthesize speech based on filtering impulse trains and noise with linear time-varying (LTV) filters, with the one or more neural networks used to control the LTV filters by estimating complex cepstrums of time-varying impulse responses given acoustic features.
  • Traditional or non-neural vocoders may operate based on decomposing speech into various parameters such as pitch, timbre, rhythm, etc., which can subsequently be manipulated and resynthesized to generate a desired output audio or voice signal.
  • Neural vocoders can apply transformations and manipulations to speech signals directly within a learned feature space of the neural network.
  • the learned feature space used by neural vocoders may capture more complex relationships and characteristics of speech than non-neural network-based signal processing and/or vocoder techniques.
  • neural vocoders can be trained to learn a mapping between raw speech waveforms and the spectral or cepstral representations of the speech.
  • An NHV system can apply one or more homomorphic processing techniques within the learned space corresponding to the mapping.
  • the decoder 1610 may also be referred to as an NHV synthesizer decoder.
  • the audio codec system 1600 of FIG.16 may be similar to the audio codec system 1500 of FIG.15.
  • the channel 1640 of FIG.16 can be the same as or similar to the channel 1540 of FIG.15, and may be associated with an encoder that is the same as or similar to the encoder 1505 of FIG.15, etc.
  • Qualcomm Ref. No.2400252WO [0148]
  • the decoder 1610 can include a neural speech synthesizer 1650 that is the same as or similar to the neural speech synthesizer 1550 included in the decoder 1510 of FIG. 15.
  • the neural speech synthesizer 1650 can receive a first input comprising decoded spectral envelope features (e.g., determined by a FRAE decoder 1645 that is the same as or similar to the FRAE decoder 1545 of FIG.15).
  • the neural speech synthesizer 1650 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 1550 of FIG. 15.
  • the NHV synthesizer decoder 1610 can include the FRAE decoder 1645, which may be the same as or similar to the FRAE decoder 1545 of FIG.
  • the pitch dequantization engine 1638 can be used to process a received pitch encoding obtained by the decoder 1610 over the channel 1640.
  • the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 1610 over the channel 1640 by a corresponding encoder (e.g., an encoder associated with the decoder 1610, which may be the same as or similar to the encoder 1505 of FIG.15, etc.).
  • the pitch dequantization engine 1638 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 1640 (e.g., to dequantize the quantized pitch encoding received by the decoder 1610 over the channel 1640).
  • the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1638 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and/or pitch correlation information.
  • the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1638 can include information indicative of a voiced or unvoiced (V/UV) classification.
  • the neural speech synthesizer 1650 can be implemented as a neural homomorphic vocoder (NHV) speech synthesizer and/or an NHV-based neural speech synthesizer.
  • the neural speech synthesizer 1650 can include a neural filter estimator 1652 configured to generate respective filter specification or filter configuration information to parameterize one or more linear time-varying (LTV) filters of the NHV speech synthesizer 1650.
  • the neural filter estimator 1652 can be trained to generate respective filter coefficients for a noise linear time-varying filter (LTVF) 1675 and to generate respective filter coefficients for a harmonic LTVF 1665.
  • the neural network model of the neural filter estimator 1652 can include any neural network architecture that can be trained to model the filter coefficients for the noise LTVF 1675 and the harmonic LTVF 1665 (e.g., and/or that can be trained to model the filter coefficients for one or more additional filters implemented by the NHV speech synthesizer 1650, in either the time-domain, the frequency-domain, or combinations thereof).
  • Examples of neural network architectures that may be included in the neural network filter estimator 1652 can include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and/or other type(s) of neural networks, etc.
  • the neural filter estimator 1652 can be configured to receive a stream of decoded spectral envelope features from the FRAE decoder 1645. Based on the decoded spectral envelope features, the neural filter estimator 1652 can generate corresponding filter characterization parameters for each filter of one or more filters included in the NHV speech synthesizer 1650.
  • the respective filter characterization parameters can be output from the neural filter estimator 1652 and used to parameterize and/or configure corresponding learned filters (e.g., learned linear filters, etc.) for each respective set of filter characterization parameters.
  • the neural filter estimator 1652 can generate a first set of filters corresponding to the noise LTVF 1675, can generate a second set of filter characterization parameters corresponding to the harmonic LTVF 1665, etc.
  • the one or more filters included in the NHV speech synthesizer 1650 e.g., the noise LTVF 1675, the harmonic LTVF 1665, etc.
  • the noise LTVF 1675, the harmonic LTVF 1665, and/or various other filters that may be included in the NHV speech synthesizer 1650 can be specified or characterized (e.g., using the filter characterization parameters determined by the neural filter estimator 1652) as cepstrums, as frequency responses, as time-domain impulse responses, as difference equation coefficients, etc.
  • the neural filter estimator 1652 can be configured to receive as input the decoded spectral envelope features from the FRAE decoder 1645, and may additionally receive as input at least a portion of the dequantized pitch information generated by the pitch dequantization engine 1638.
  • the neural filter estimator 1652 may receive as input the decoded Qualcomm Ref.
  • No.2400252WO spectral envelope features and pitch dequantization information e.g., such as f0 pitch information in the frequency domain, pitch lag or pitch delay in the time domain (e.g., in units of samples or milliseconds, etc.), etc.
  • the pitch dequantization information provided as input to the neural filter estimator 1652 may include voiced/unvoiced (V/UV) classification information indicating whether the underlying audio represented in the encoded information received by the decoder 1610 over the channel 1640 corresponds to voiced or unvoiced sounds, speech, etc.
  • V/UV voiced/unvoiced
  • the neural filter estimator 1652 can generate the corresponding filter characterization parameters for each NHV filter (e.g., noise LTVF 1675, harmonic LTVF 1665, etc.) based on the decoded spectral envelope features and the dequantized pitch information.
  • the noise LTVF 1675 and the harmonic LTVF 1665 can be implemented as time domain filters or frequency domain filters. In some examples, the noise LTVF 1675 and the harmonic LTVF 1665 can be implemented in the time domain or the frequency domain, independent of whether the neural filter estimator 1652 is configured to generate the corresponding filter characterization parameters in the time domain or the frequency domain.
  • the neural filter estimator 1652 can output impulse response-based filter characterization parameters, and the operation of noise LTVF 1675 and/or harmonic LTVF 1665 can be implemented in the time domain as convolutions with the impulse response.
  • the operation of noise LTVF 1675 and/or harmonic LTVF 1665 may be implemented in the frequency domain based on converting the impulse response to frequency response (e.g., using an FFT transform) and multiplying with the FFT of the input signal, and subsequently converting back to the time domain using an IFFT transform.
  • the NHV speech synthesizer 1650 can include a pulse train generator 1660.
  • the dequantized pitch information (e.g., generated using the pitch dequantization engine 1638) can be provided to the pulse train generator 1660.
  • the pulse train generator 1660 can generate a pulse train based at least in part on the pitch frequency (e.g., f0) and/or pitch lag or pitch delay information obtained from the pitch dequantization engine 1638 of the decoder 1610.
  • the pulse train generator 1660 can be implemented as a cosine Qualcomm Ref. No.2400252WO sum pulse generator, which can be configured to process the dequantized pitch information (e.g., obtained from the pitch dequantization engine 1638) to generate a pulse train p[n].
  • the pulse train may also be referred to as an impulse train.
  • the pulse train generator 1660 can be a differentiable cosine sum pulse generator, and/or can be a non-differentiable cosine sum pulse generator (e.g., among various other pulse generators).
  • the pulse train generated by the pulse train generator 1660 can be processed by the harmonic LTVF 1665, using the corresponding harmonic filter characterization parameters determined by the neural filter estimator 1652 for the harmonic LTVF 1665.
  • the harmonic LTVF 1665 can generate a harmonic output based on processing the pulse train from the pulse train generator 1660.
  • the pulse train generator 1660 may receive an additional input from the pitch dequantization engine 1638, indicative of a voice or unvoiced (e.g., V/UV) classification.
  • the pulse train generator 1660 can be configured to generate a 0 output or a noise output that is provided to the harmonic LTVF 1665 instead of the pulse train p[n]ep[n] provided from the pulse train generator 1660 to the harmonic LTVF 1665 in response to a voiced (V) indication or classification from the pitch dequantization engine 1638).
  • the NHV speech synthesizer 1650 can include a noise generator 1670 that is configured to generate a noise signal (e.g., white noise, etc.) for processing by the noise LTVF 1675.
  • the noise LTVF 1675 can be parameterized based on the respective noise filter characterization parameters generated by the neural filter estimator 1652 for the noise LTVF 1675, and can subsequently be used to process the noise signal generated by the noise generator 1670. Based on processing the noise signal from the noise generator 1670, the noise LTVF 1675 can generate a noise-filtered output.
  • a hangover detector 1681 may detect that an audio frame is from the hangover region.
  • the hangover detector 1681 may be a portion of a decoder (e.g., decoder 701 of FIG.7) or classifier (e.g., classifier 703 of FIG.7) that may detect (and/or infer) an indication that an audio frame is from a hangover region. Based on the detection that the audio frame is from the hangover region, the hangover detector 1681 may indicate to a filter coefficients copier 1682 to copy the coefficients of the noise LTVF 1675 to noise filter 1684.
  • Qualcomm Ref. No.2400252WO noise filter 1684 may operate as an inactive synthesizer (e.g., inactive synthesizer 712 of FIG.7, inactive synthesizer 1104 of FIG.
  • the random noise generator 1686 may correspond to random noise generator 1102 of FIG. 11.
  • a hangover region may include inactive frames following speech and thus filter coefficients estimated in the hangover region may be used for noise generation and the noise filter 1684 may mix the random noise from the random noise generator 1686 with synthesized background noise based on the copied filter coefficients.
  • the random noise generator 1686 may be the same as noise generator 1670, or random noise generator 1686 may be a separate noise generator from noise generator 1670.
  • the NHV speech synthesizer 1650 can include a combination function 1680 to combine the harmonic output (e.g., from the harmonic LTVF 1665) and the noise-filtered output (e.g., from the noise LTVF 1675) to generate a predicted sample (S' t ) that may be input to a LPC 1654 during active speech.
  • the combination function 1680 may combine the harmonic output (e.g., from the harmonic LTVF 1665) and the synthesized noise output from the noise filter 1684 to generate a predicted sample (S' t ) that may be input to a LPC 1654.
  • the combination function 1680 can include an adder, a multiplier, a divider, weighted sum, a weighted product, a weighted ratio, an average, a weighted average, a weighted mean, a weighted median, a weighted mode, or a combination thereof.
  • the reconstructed audio 1655 may be an example of the reconstructed speech signal 105 of FIG.1, the reconstructed speech signal 205 of FIG.2, etc.
  • FIG.17 is a block diagram illustrating an example of a decoder 1710 of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear prediction coding (LPC) network with a technique for random background noise synthesis, in accordance with aspects of the present disclosure, in accordance with aspects of the present disclosure.
  • the decoder 1710 can be included in an audio codec system (e.g., generative voice codec system) 1700 that can be used to generate reconstructed audio (e.g., synthesized speech) 1755.
  • an audio codec system e.g., generative voice codec system
  • the generative voice codec system 1700 of FIG.17 can be the same as or similar to the generative voice codec system 1500 of FIG. 15, 1600 of FIG. 16, etc.
  • the channel 1740 of FIG.17 can be the same as or similar to the channel 1540 of FIG.15, the channel 1640 of FIG.16, etc.
  • the decoder 1710 of FIG.17 can be the same as or similar to the decoder 1510 of FIG. 15, the decoder 1610 of FIG.16, etc.
  • the decoder 1710 can include an FRAE decoder 1745 the same as or similar to the FRAE decoder 1545 of FIG. 15, 1645 of FIG. 16, etc.
  • the decoder 1710 can include a neural speech synthesizer 1750 that can be the same as or similar to the neural speech synthesizer 1550 of FIG.15, 1650 of FIG.16, etc.
  • the decoder 1710 can generate reconstructed audio 1755 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 1555 of FIG.15, 1655 of FIG.16, etc.
  • the neural speech synthesizer 1750 can receive a first input comprising decoded spectral envelope features (e.g., determined by the FRAE decoder 1745).
  • the neural speech synthesizer 1750 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 1550 of FIG.15, etc.
  • the decoder 1710 may include a pitch de-quantization engine 1738 (e.g., de Q()).
  • the pitch de-quantization engine 1738 can be used to process a received pitch encoding obtained by the decoder 1710 over the channel 1740.
  • the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 1710 over the channel 1740 by a corresponding encoder (e.g., an encoder associated with the decoder 1710, which may be the same as or similar to the encoder 1505 of FIG.15, etc.).
  • the pitch dequantization engine 1738 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 1740 (e.g., to dequantize the quantized pitch encoding received by the decoder 1710 over the channel 1740).
  • the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1738 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and/or pitch correlation information.
  • the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1738 can include information indicative of a voiced or unvoiced (V/UV) classification.
  • Qualcomm Ref. No.2400252WO the neural speech synthesizer 1750 can be implemented as a linear predictive coding (LPC) network.
  • the LPC network-based neural speech synthesizer 1750 can be used to implement a linear prediction (LP) synthesis filter, based on ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
  • represents the input signal to the LP filter
  • represents the output signal
  • ⁇ ⁇ represents the linear prediction coefficients associated with implementing the LP synthesis filter
  • is a value corresponding to the LP filter order.
  • the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.
  • the LPC-based neural speech synthesizer 1750 can include a frame rate network 1752 configured to process inputs comprising the decoded features obtained using the FRAE decoder 1745 and the decoded pitch information obtained using the pitch dequantization engine 1738.
  • An LPC estimation engine 1762 can process the decoded features from the FRAE decoder 1745 to determine one or more estimated LP coefficients.
  • the LPC estimation engine 1762 can generate estimated LP coefficients ⁇ ⁇ , based on an input comprising the decoded features determined using the FRAE decoder 1745.
  • the LP coefficients estimated using the LPC estimation engine 1762 can be provided to an LP prediction engine 1764, configured to generate as output a prediction p[n], where ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
  • the LP prediction engine 1764 can generate the LP prediction p[n] based on a first input comprising the LP coefficients ⁇ ⁇ estimated by the LPC estimation engine 1762, and a feedback input s[n-1].
  • the estimated or predicted LP coefficients p[n] can be provided as input to a sample rate network 1772 and a downstream combination or summation operation 1780.
  • the sample rate network 1772 can receive additional inputs comprising the output of the frame rate network 1752, the feedback s[n-1] from the previous step n-1, and an intermediate feedback value e[n-1] from the same previous step n-1.
  • the output of the sample rate network 1772 can be the probability distribution P(e[n]), which is a probability distribution for e[n].
  • a sampling engine 1774 can perform sampling from the probability distribution P(e[n]) to obtain a realization or representation of e[n].
  • the representation of e[n] determined by the sampling engine 1774 can be combined with the LP prediction p[n] (e.g., determined by the LP prediction engine 1764), using the combination or summation operation 1780 to thereby generate as output the signal s[n].
  • the output signal s[n] can be the same as the reconstructed audio 1755 of the LPC network- based neural speech synthesizer 1750 and/or decoder 1710.
  • the representation of e[n] determined by the sampling engine 1774 can additionally be provided to a first feedback calculation 1778, which generates as output e[n-1] provided as an additional input to the sample rate network 1772.
  • the output signal s[n] of the combination or summation operation 1780 can be output as the reconstructed audio 1755 and may additionally be provided to a second feedback calculation 1779, which generates the representation s[n-1] based on the input s[n].
  • the representation s[n-1] can be provided as a feedback input to the LP prediction engine 1764 and to the sample rate network 1772.
  • a hangover detector 1781 may detect that an audio frame is from the hangover region.
  • the hangover detector 1781 may be a portion of a decoder that may detect (and/or infer) an indication that an audio frame is from a hangover region in a manner similar to that discussed above with respect to decoder 701 of FIG.7. Based on the detection that the audio frame is from the hangover region, the hangover detector 1781 may indicate to a noise filter 1784 that the frame is from the hangover region.
  • noise filter 1784 may operate as an inactive synthesizer (e.g., inactive synthesizer 712 of FIG.7, inactive synthesizer 1104 of FIG.11) to use the random noise from the random noise generator 1786 and output the random noise to the combination or summation operation 1780 to mix in as the synthesized background noise.
  • the random noise generator 1786 may correspond to random noise generator 1102 of FIG. 11.
  • a hangover region may include inactive frames following speech and thus filter coefficients estimated in the hangover region may be used for noise generation.
  • FIG.18 is a flow diagram illustrating an example of a process 1800 for audio playback, in accordance with aspects of the present disclosure.
  • the process 1800 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, processor 484 of FIG.4, DSP 482 of FIG.4, processor 1910 of FIG.19, etc.) of the computing device (e.g., UE 104 or UE 190 of FIGs.1-3, wireless device 407 of FIG.7, receiving UE 1406 of FIG.14, computing system Qualcomm Ref. No.2400252WO 1900 of FIG. 19, etc.).
  • a computing device or apparatus
  • a component e.g., a chipset, codec, processor 484 of FIG.4, DSP 482 of FIG.4, processor 1910 of FIG.19, etc.
  • the computing device e.g., UE 104 or UE 190 of FIGs.1-3, wireless device 407 of FIG.7, receiving UE 1406 of FIG.14, computing system Qualcomm Ref. No.2400252WO 1900 of FIG. 19, etc.
  • the computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device.
  • the computing device may be or may include UE device, such as the UE 104 or UE 190 of FIGs.1-3.
  • the operations of the process 1800 may be implemented as software components that are executed and run on one or more processors.
  • the computing device (or component thereof) may receive a first decoded audio frame.
  • the first decoded audio frame includes active speech.
  • the computing device may decode an audio frame of the encoded audio data stream to obtain decoded active speech.
  • the computing device may receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device.
  • SID silence descriptor
  • the computing device may generate information associated with background noise (e.g., information about the background noise 710 of FIG.7) in the first decoded audio frame.
  • the computing device may generate the information associated with the background noise in the first decoded audio frame by: denoising (e.g., by denoiser 704 of FIG 7) the first decoded audio frame to generate a denoised audio frame; subtracting (e.g., subtracted 706 of FIG.7) the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generating the information associated with the background noise based on the background noise frame.
  • the information about the background noise in the first decoded audio frame is generated by a machine learning model.
  • the information associated with background noise is determined based on at least one of a noise dictionary (e.g., noise dictionary 960 of FIG.9B) or a speech dictionary (e.g., speech dictionary 962 of FIG. 9B).
  • the information associated with background noise is generated by a neural synthesizer (e.g., neural speech synthesizer 1550 of FIG.15, neural speech synthesizer 1650 of FIG.16, etc.) of a decoder (e.g., decoder 1510 of FIG.15, decoder 1610 of FIG.16, etc.).
  • the information associated with the background noise is received from a noise linear time-varying filter (e.g., noise LTVF 1675 of FIG.16) of the neural speech synthesizer.
  • the synthesized background noise is combined with an output of a harmonic linear time-varying filter (e.g., harmonic LTVF Qualcomm Ref. No.2400252WO 1665 of FIG.16) of the neural speech synthesizer to generate a predicted sample.
  • the predicted sample is input to a linear predictive coding (LPC) network (e.g., LPC 1552 of FIG.
  • LPC linear predictive coding
  • the computing device may detect a transition to an inactive region after the first decoded audio frame.
  • the encoder may indicate a transition to the inactive region
  • the decoder 701 of FIG.7 and/or classifier 703 of FIG.7 may determine that there is a transition to the inactive region.
  • the computing device may to detect the transition to the inactive region after the first decoded audio frame by receiving a radio access network message (e.g., message 1420 of FIG.14) indicating the transition to the inactive region.
  • a radio access network message e.g., message 1420 of FIG.14
  • the computing device may synthesize background noise (e.g., synthesized background noise 714 of FIG.7) based on the information associated with the background noise in response to the detected transition to the inactive region.
  • the computing device may receive a third decoded audio frame (e.g., third input frame 870 of FIG.8B), the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region.
  • the background noise is synthesized by a machine learning model (e.g., GRU 802, 810, 818 of FIG.8A, inactive synthesizer 850 of FIG.8B, ambient noise prediction engine 1004 of FIG.10, ML model 1304 of FIG.13, etc.).
  • the computing device may receive an encoded audio frame from a second device; receive location information (e.g., location information 1002 of FIG.10) for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information.
  • the ambient noise prediction engine 1004 of FIG. 10 may predict a type of ambient noise (e.g., background noise) 1006 FIG. 10 that may be present based on the location information 1002 FIG. 10.
  • the background noise is synthesized based on random noise. For example, a random Qualcomm Ref.
  • No.2400252WO noise generator 1102 of FIG.11 may be used to generate random noise based on, for example, a random number generator.
  • the techniques or processes described herein may be performed by a computing device, an apparatus, and/or any other computing device.
  • the computing device or apparatus may include a processor, microprocessor, microcomputer, or other component of a device that is configured to carry out the steps of processes described herein.
  • the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames.
  • the computing device may include a camera device, which may or may not include a video codec.
  • the computing device may include a mobile device with a camera (e.g., a camera device such as a digital camera, an IP camera or the like, a mobile phone or tablet including a camera, or other type of device with a camera).
  • the computing device may include a display for displaying images.
  • a camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives the captured video data.
  • the computing device may further include a network interface, transceiver, and/or transmitter configured to communicate the video data.
  • the network interface, transceiver, and/or transmitter may be configured to communicate Internet Protocol (IP) based data or other network data.
  • IP Internet Protocol
  • the processes described herein can be implemented in hardware, computer instructions, or a combination thereof.
  • the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations.
  • computer- executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types.
  • the order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes.
  • the devices or apparatuses configured to perform the operations of the process 1800 and/or other processes described herein may include a processor, microprocessor, micro-computer, or other component of a device that is configured to carry out the steps of the process 1800 and/or other process.
  • such devices or apparatuses may include one or more sensors configured to capture image data and/or other sensor measurements.
  • such computing device or apparatus may include one or more sensors and/or a camera configured to capture one or more images or videos.
  • such device or apparatus may include a display for displaying images.
  • the one or more sensors and/or camera are separate from the device or apparatus, in which case the device or apparatus receives the sensed data.
  • Such device or apparatus may further include a network interface configured to communicate data.
  • the components of the device or apparatus configured to carry out one or more operations of the process 1800 and/or other processes described herein can be implemented in circuitry.
  • the components can include and/or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and/or other suitable electronic circuits), and/or can include and/or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
  • programmable electronic circuits e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and/or other suitable electronic circuits
  • the computing device may further include a display (as an example of the output device or in addition to the output device), a network interface configured to communicate and/or receive the data, any combination thereof, and/or other component(s).
  • the network interface may be configured to communicate and/or receive Internet Protocol (IP) based data or other type of data.
  • IP Internet Protocol
  • computer- executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types.
  • the order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes.
  • the processes described herein e.g., process 1800 and/or other processes
  • the code may be stored on a computer- readable or machine-readable storage medium, for example, in the form of a computer program including a plurality of instructions executable by one or more processors.
  • the computer-readable or machine-readable storage medium may be non-transitory.
  • the processes described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof.
  • the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors.
  • FIG.19 is a diagram illustrating an example of a system for implementing certain aspects of the present technology.
  • computing system 1900 may be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 1905.
  • Connection 1905 may be a physical connection using a bus, or a direct connection into processor 1910, such as in a chipset architecture.
  • Connection 1905 may also be a virtual connection, networked connection, or logical connection.
  • computing system 1900 is a distributed system in which the functions described in this disclosure may be distributed within a datacenter, multiple data centers, a peer network, etc.
  • one or more of the described system components represents many such components each performing some or all of the function for which the component is described.
  • the components may be physical or virtual devices.
  • Example system 1900 includes at least one processing unit (CPU or processor) 1910 and connection 1905 that communicatively couples various system components including system memory 1915, such as read-only memory (ROM) 1920 and random access memory (RAM) 1925 to processor 1910.
  • Computing system 1900 may include a cache 1912 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1910. Qualcomm Ref.
  • Processor 1910 may include any general purpose processor and a hardware service or software service, such as services 1932, 1934, and 1936 stored in storage device 1930, configured to control processor 1910 as well as a special-purpose processor where software instructions are incorporated into the actual processor design.
  • Processor 1910 may essentially be a completely self- contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc.
  • a multi-core processor may be symmetric or asymmetric.
  • computing system 1900 includes an input device 1945, which may represent any number of input mechanisms, such as a microphone for speech, a touch- sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc.
  • Computing system 1900 may also include output device 1935, which may be one or more of a number of output mechanisms. In some instances, multimodal systems may enable a user to provide multiple types of input/output to communicate with computing system 1900. [0193] Computing system 1900 may include communications interface 1940, which may generally govern and manage the user input and system output.
  • the communication interface may perform or facilitate receipt and/or transmission wired or wireless communications using wired and/or wireless transceivers, including those making use of an audio jack/plug, a microphone jack/plug, a universal serial bus (USB) port/plug, an AppleTM LightningTM port/plug, an Ethernet port/plug, a fiber optic port/plug, a proprietary wired port/plug, 3G, 4G, 5G and/or other cellular data network wireless signal transfer, a BluetoothTM wireless signal transfer, a BluetoothTM low energy (BLE) wireless signal transfer, an IBEACONTM wireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (
  • the communications interface 1940 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to Qualcomm Ref. No.2400252WO determine a location of the computing system 1900 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems.
  • GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS.
  • GPS Global Positioning System
  • GLONASS Russia-based Global Navigation Satellite System
  • BDS BeiDou Navigation Satellite System
  • Galileo GNSS Europe-based Galileo GNSS
  • Storage device 1930 may be a non-volatile and/or non-transitory and/or computer- readable memory device and may be a hard disk or other types of computer readable media which may store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip/stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini
  • the storage device 1930 may include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1910, it causes the system to perform a function.
  • a hardware service that performs a particular function may include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1910, connection 1905, output device 1935, Qualcomm Ref. No.2400252WO etc., to carry out the function.
  • the term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and/or data.
  • a code segment may be coupled to another code segment or a hardware circuit by passing and/or receiving information, data, arguments, parameters, or memory contents.
  • Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
  • circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail.
  • well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.
  • those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.
  • Processes and methods according to the above-described examples may be implemented using computer-executable instructions that are stored or otherwise available from computer- readable media. Such instructions may include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used may be accessible over a network.
  • the computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of Qualcomm Ref.
  • No.2400252WO computer-readable media that may be used to store instructions, information used, and/or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
  • the computer-readable storage devices, mediums, and memories may include a cable or wireless signal containing a bitstream and the like.
  • non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
  • the program code or code segments to perform the necessary tasks may be stored in a computer-readable or machine-readable medium.
  • a processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also may be embodied in peripherals or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
  • the instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
  • Qualcomm Ref. No.2400252WO The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices.
  • the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and/or operations described above.
  • the computer-readable data storage medium may form part of a computer program product, which may include packaging materials.
  • the computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non- volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like.
  • RAM random access memory
  • SDRAM synchronous dynamic random access memory
  • ROM read-only memory
  • NVRAM non- volatile random access memory
  • EEPROM electrically erasable programmable read-only memory
  • FLASH memory magnetic or optical data storage media, and the like.
  • the techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that may be accessed, read, and/or executed by a computer, such as propagated signals or waves.
  • the program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry.
  • DSPs digital signal processors
  • ASICs application specific integrated circuits
  • FPGAs field programmable logic arrays
  • Such a processor may be configured to perform any of the techniques described in this disclosure.
  • a general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine.
  • a processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein. Qualcomm Ref.
  • Coupled to or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and/or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and/or other suitable communication interface) either directly or indirectly.
  • Claim language or other language reciting “at least one of” a set and/or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B.
  • claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C.
  • the language “at least one of” a set and/or “one or more” of a set does not limit the set to the items listed in the set.
  • claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of Qualcomm Ref. No.2400252WO operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z.
  • claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
  • one element may perform all functions, or more than one element may collectively perform the functions.
  • each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and/or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function).
  • one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
  • an entity e.g., any entity or device described herein
  • the entity may be configured to cause one or more elements (individually or collectively) to perform the functions.
  • the one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and/or any combination thereof.
  • the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions.
  • a device for audio playback comprising: one or more memories; and one or more processors coupled to the one or more memories and configured to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information Qualcomm Ref.
  • No.2400252WO associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • Aspect 2 The device of Aspect 1, wherein, to generate the information associated with the background noise in the first decoded audio frame, the one or more processors are configured to: denoise the first decoded audio frame to generate a denoised audio frame; subtract the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generate the information associated with the background noise based on the background noise frame.
  • Aspect 4 The device of any of Aspects 1-3, wherein the background noise is synthesized by a machine learning model. [0219] Aspect 5.
  • Aspect 6 The device of any of Aspects 1-4, wherein the one or more processors are configured to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information.
  • Aspect 6 The device of any of Aspects 1-5, wherein the background noise is synthesized based on random noise.
  • Aspect 7. The device of any of Aspects 1-6, wherein the one or more processors are configured to receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device.
  • SID silence descriptor
  • Aspect 9 The device of any of Aspects 1-7, wherein the background noise is synthesized based on random noise.
  • Aspect 9. The device of any of Aspects 1-8, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model.
  • Aspect 10. The device of any of Aspects 1-9, wherein, to detect the transition to the inactive region after the first decoded audio frame, the one or more processors are configured to receive a radio access network message indicating the transition to the inactive region.
  • Aspect 11 The device of any of Aspects 1-10, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary.
  • a method for audio playback comprising: receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • Aspect 13 The method of Aspect 12, wherein generating the information associated with the background noise in the first decoded audio frame comprises: denoising the first decoded audio frame to generate a denoised audio frame; subtracting the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generating the information associated with the background noise based on the background noise frame.
  • Aspect 14 The method of any of Aspects 12-13, comprising: receiving a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generating information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region.
  • Aspect 15 The method of any of Aspects 12-14, wherein the background noise is synthesized by a machine learning model.
  • Aspect 16 The method of any of Aspects 12-15, comprising: receiving an encoded audio frame from a second device; receiving location information for the second device; and decoding Qualcomm Ref.
  • Aspect 17 The method of any of Aspects 12-16, wherein the background noise is synthesized based on random noise.
  • Aspect 18 The method of any of Aspects 12-17, comprising receiving, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device.
  • SID silence descriptor
  • Aspect 19 The method of any of Aspects 12-18, wherein the background noise is synthesized based on random noise.
  • Aspect 20 The method of any of Aspects 12-18, wherein the background noise is synthesized based on random noise.
  • Aspect 21 The method of any of Aspects 12-20, wherein detecting the transition to the inactive region after the first decoded audio frame comprises receiving a radio access network message indicating the transition to the inactive region.
  • Aspect 22 The method of any of Aspects 12-21, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary.
  • a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
  • Aspect 26 The non-transitory computer-readable medium of any of Aspects 23-24, wherein instructions case the one or more processors to: receive a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region.
  • Aspect 26 The non-transitory computer-readable medium of any of Aspects 23-25, wherein the background noise is synthesized by a machine learning model.
  • Aspect 28 The non-transitory computer-readable medium of any of Aspects 23-26, wherein instructions case the one or more processors to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information.
  • Aspect 28 The non-transitory computer-readable medium of any of Aspects 23-27, wherein the background noise is synthesized based on random noise.
  • Aspect 29 Aspect 29.
  • Aspect 30 The non-transitory computer-readable medium of any of Aspects 23-28, wherein the instructions case the one or more processors to receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device.
  • Aspect 30 The non-transitory computer-readable medium of any of Aspects 23-29, wherein the background noise is synthesized based on random noise.
  • Aspect 31 The non-transitory computer-readable medium of any of Aspects 23-30, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model.
  • Aspect 32 The non-transitory computer-readable medium of any of Aspects 23-30, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model.
  • Aspect 33 The non-transitory computer-readable medium of any of Aspects 23-31, wherein, to detect the transition to the inactive region after the first decoded audio frame, the instructions case the one or more processors to receive a radio access network message indicating the transition to the inactive region.
  • Qualcomm Ref. No.2400252WO [0247]
  • Aspect 33 The non-transitory computer-readable medium of any of Aspects 23-32, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary.
  • Aspect 34 The device of any of Aspects 1-11, wherein the information associated with background noise is generated by a neural synthesizer of a decoder.
  • Aspect 35 The device of any of Aspects 1-11, wherein the information associated with background noise is generated by a neural synthesizer of a decoder.
  • Aspect 35 wherein the information associated with the background noise is received from a noise linear time-varying filter of the neural speech synthesizer.
  • Aspect 36 The device of Aspect 35, wherein the synthesized background noise is combined with an output of a harmonic linear time-varying filter of the neural speech synthesizer to generate a predicted sample.
  • Aspect 37 The device of Aspect 36, wherein the predicted sample is input to a linear predictive coding (LPC) network to generate reconstructed audio.
  • LPC linear predictive coding
  • Aspect 34 An apparatus comprising means for performing a method according to any of Aspects 12 to 22 and Aspects 34-37.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Quality & Reliability (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

Disclosed are systems and techniques for audio playback. For instance, a process can include receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region.

Description

Qualcomm Ref. No.2400252WO DECODER SILENCE GENERATION WITHOUT CODED SILENCE DESCRIPTION FIELD [0001] The present disclosure generally relates to audio coding (e.g., audio encoding and/or decoding). For example, aspects of the present disclosure relate to systems and techniques for decoder silence generation without coded silence description. BACKGROUND [0002] Audio coding (also referred to as voice coding and/or speech coding) is a technique used to represent a digitized audio signal using as few bits as possible (thus compressing the speech data), while attempting to maintain a certain level of audio quality. An audio or voice encoder is used to encode (or compress) the digitized audio (e.g., speech, music, etc.) signal to a lower bit- rate stream of data. The lower bit-rate stream of data can be input to an audio or voice decoder, which decodes the stream of data and constructs an approximation or reconstruction of the original signal. The audio or voice encoder-decoder structure can be referred to as an audio coder (or voice coder or speech coder) or an audio/voice/speech coder-decoder (codec). [0003] Audio coders exploit the fact that speech signals are highly correlated waveforms. Some speech coding techniques are based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech. The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter). SUMMARY [0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below. Qualcomm Ref. No.2400252WO [0005] Disclosed are systems, methods, apparatuses, and computer-readable media for audio coding (e.g., encoding and/or decoding audio data). In one illustrative example, a device for audio playback is provided. The first device includes: one or more memories comprising instructions; and one or more processors coupled to the one or more memories and configured to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0006] As another example, a method for audio playback is provided. The method includes receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0007] In another example, a non-transitory computer-readable medium having stored thereon instructions is provided. The instructions, when executed by one or more processors, cause the one or more processors to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0008] As another example, an apparatus for audio playback is provided. The apparatus includes means for receiving a first decoded audio frame, the first decoded audio frame including active speech; means for generating information associated with background noise in the first decoded audio frame; means for detecting a transition to an inactive region after the first decoded audio frame; and means for synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0009] Aspects generally include a method, apparatus, system, computer program product, non- transitory computer-readable medium, user equipment, base station, wireless communication Qualcomm Ref. No.2400252WO device, and/or processing system as substantially described herein with reference to and as illustrated by the drawings and specification. [0010] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims. [0011] While aspects are described in the present disclosure by illustration to some examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios. Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and/or packaging arrangements. For example, some aspects may be implemented via integrated chip implementations or other non-module- component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail/purchasing devices, medical devices, and/or artificial intelligence devices). Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and/or system-level components. Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects. For example, transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and/or summers). It is intended that aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and/or end-user devices of varying size, shape, and constitution. Qualcomm Ref. No.2400252WO [0012] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim. [0013] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS [0014] Examples of various implementations are described in detail below with reference to the following figures: [0015] FIG. 1 is a block diagram illustrating an example speech processing system, in accordance with some examples; [0016] FIG.2 is a block diagram illustrating an example feature generator, in accordance with some examples; [0017] FIG.3 is a block diagram illustrating an example of a voice coding system, in accordance with some examples; [0018] FIG. 4 is a block diagram illustrating an example of a code-excited linear prediction (CELP)-based voice coding system utilizing a fixed codebook (FCB), in accordance with some examples; [0019] FIG. 5 is a block diagram illustrating an example of a voice coding signal synthesis system utilizing a linear time-varying filter generated using a neural network model and a separate linear predictive coding (LPC) filter, in accordance with some examples; [0020] FIG. 6 illustrates an audio waveform representing audio information received by a wireless device, in accordance with aspects of the present disclosure; [0021] FIG. 7 is a block diagram providing an overview of a technique for decoder silence generation without, or with minimal, coded SIDs, in accordance with aspects of the present disclosure; Qualcomm Ref. No.2400252WO [0022] FIG. 8A is a block diagram illustrating an example inactive synthesizer, in accordance with aspects of the present disclosure; [0023] FIG. 8B is a is a block diagram illustrating another example inactive synthesizer, in accordance with aspects of the present disclosure; [0024] FIG. 9A is a block diagram illustrating a denoiser, in accordance with aspects of the present disclosure; [0025] FIG. 9B is a block diagram illustrating non-negative matrix factorization (NMF) denoising, in accordance with aspects of the present disclosure; [0026] FIG.10 is a block diagram illustrating a technique for location aware background noise synthetization, in accordance with aspects of the present disclosure; [0027] FIG. 11 is a block diagram illustrating a technique for random background noise synthetization, in accordance with aspects of the present disclosure; [0028] FIG.12 illustrates a spectrum estimator for measuring additional spectral characteristics, in accordance with aspects of the present disclosure; [0029] FIG.13 is a block diagram illustrating another technique for decoder silence generation without, or with minimal, coded SIDs, in accordance with aspects of the present disclosure; [0030] FIG. 14 is signal diagram illustrating signals for decoder silence generation without coded silence descriptions, in accordance with aspects of the present disclosure; [0031] FIG.15 is a block diagram illustrating an example of an audio codec system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer; [0032] FIG. 16 illustrates an example of an audio codec system (e.g., generative voice codec) that includes a decoder configured to implement an NHV-based neural speech synthesizer with a technique for random background noise synthesis, in accordance with aspects of the present disclosure; [0033] FIG.17 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using Qualcomm Ref. No.2400252WO a neural speech synthesizer comprising a linear prediction coding (LPC) network with a technique for random background noise synthesis, in accordance with aspects of the present disclosure; [0034] Fig. 18 is a flow diagram illustrating an example of a process for audio playback, in accordance with aspects of the present disclosure; and [0035] FIG.19 is a diagram illustrating an example of a computing system, according to aspects of the present disclosure. DETAILED DESCRIPTION [0036] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive. [0037] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example aspects will provide those skilled in the art with an enabling description for implementing an example aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims. [0038] In human speech, there may be occasional breaks (e.g., pauses, stops, etc.). During these breaks, while no speech may be occurring, there may still be background noise occurring. During a voice call, this background noise during breaks may go unnoticed, but if the background noise is replaced by actual silence, the absence of this background noise can be uncomfortable to users. In some cases, SID frames may be used to describe the background noise during breaks. The SID may include information about background noise, and this information may be used to generate comfort noise by a decoding device. As used herein, silence may refer to audio frames without speech, but with background noise, while absolute silence may refer to an absence of speech and background noise. In some cases, radio resource and/or encoder device power may be limited and techniques to reduce encoding and transmitting SID frames may be useful. Qualcomm Ref. No.2400252WO [0039] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for decoder silence generation without coded silence description. In some cases, rather than transmitting SID frames, SID frames may be minimized or eliminated by predicting and synthesizing comfort noise at the decoder (e.g., decoding device). In some cases, speech may include regions, such as an active speech region, a hangover region, and an inactive (e.g., silent) region. In an active speech region, speech is present along with background noise. In the hangover region, less speech may be present and background noise may be present. In some cases, it may be useful to generate information associated with the background noise while in the active speech region. This information associated with the background noise may be used to synthesize background noise when a transition to the inactive region is detected. In some cases, the information associated with the background noise may be obtained by applying a denoise filter to an audio frame from the active speech region and subtracting the denoised version of the audio frame from the audio frame to obtain a background noise frame. Information associated with the background noise may then be generated based on the background noise frame. In some cases, a machine learning (ML) model may be used to predict the background noise from the audio frame. In some cases, non-negative matrix factorization (NMF) may be used to extract the background noise. In some cases, an audio frame from the handover region may be received. As indicated above, audio frames from the handover region may include primarily background noise. The information about the background noise may be generated using audio frames from the active speech region and audio frames from the handover region. In some cases, the synthesized background noise may be synthesized based on location information from the sending device. In some cases, the synthesized background noise may be synthesized in part by a ML model. In some cases, the synthesized background noise may be synthesized based on random noise. In some cases, a transition to an inactive region may be signaled by a radio access network message, such as an RRC message. [0040] Autoencoders have gained popularity in recent years as they are able to learn efficient representations of input data without the need of labels (e.g., based on performing unsupervised learning). Various types of autoencoders exist and are well explained in “Autoencoder and its various variants” 2018 IEEE International Conference on Systems, Man, and Cybernetics, Zhang et. al. In speech coding, autoencoders conditioned on log Mel Cepstrum inputs and/or spectrogram inputs have been used for speech compression in "Enhancing into the Codec: Noise Robust Speech Qualcomm Ref. No.2400252WO Coding with Vector-Quantized Autoencoders" arXiv:2102.06610v1 [eess.AS] 12 Feb 2021, Casebeer et. al. [0041] Speech coding is a lossy compression process. Autoencoders perform dimensionality reduction (e.g., an N-dimensional vector input vector is passed into the encoder of the autoencoder, and lossy compression is performed to represent the important aspects of the input vector in an M- dimensional vector, where M is smaller than N, e.g. by an order of magnitude). A drawback of using an autoencoder for speech coding is that it cannot necessarily exploit the temporal relationship between sets of input data. U.S. Patent No.11,526,734 (“the ‘734 patent”, assigned to Qualcomm, Inc.), "Method and Apparatus for recurrent auto encoding", Yang et. al. improves upon an autoencoder to exploit temporal redundancies and correlations and describes a feedback recurrent autoencoder ("FRAE"). The FRAE in the '734 patent can be used for training and application of compression of sequential data with temporal correlation. The recurrent structure of the FRAE can be used to efficiently extract the redundancy embedded along the time-dimension of sequential data and enables compact discrete representation of the data at the bottleneck in a sequential fashion. For example, Table 1 of the '734 patent illustrates the MSE (Mean Squared Error) used as Mel-scale mean-square-error configured as a reconstruction loss for training of the FRAE, where the MSE of each frequency bin is scaled according to its weight at Mel-frequency both for latent feedback and output feedback. [0042] The FRAE described in the ‘734 patent has two advantageous features not present in autoencoders: (1) recurrent layers, e.g. LSTM or GRU layers, have memory of the past; and (2) feedback from the decoder of the autoencoder to the encoder of the autoencoder. The feedback connection 150 in Fig. 2, Fig. 3 and Fig. 7 of the ‘734 patent provides additional historical information from a state (ht) in the decoder to the encoder indicative of how reconstruction of a prior input has fared. This feedback loop is present during training and inference, which allows the encoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder. Thus, the feedback loop is analogous to a mode switch input in that it is not encoded by the encoder, but it is used as an input that influences how the encoder operates on the next input vector. [0043] In addition, the ‘734 patent describes a second feedback connection (152) from the state (ht) of the decoder for a first iteration of series of inputs Xt, to the next iteration of series of inputs Qualcomm Ref. No.2400252WO Xt+1. This second feedback connection (152) enables the decoder to learn from its previous reconstruction attempts, providing additional historical context about how reconstruction of a prior input has fared. This second feedback loop is also present during both training and inference, which allows the decoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder. [0044] The ‘734 patent also describes other optional feedback connections. For example, an embedding vector (z) may be fed back via a third feedback connection (356 in Fig.3 of the ‘734 patent) to the encoder. As another optional example, the '734 patent describes that the FRAE can be a variational autoencoder. In this example, an output of the decoder (754 in Fig.7 of the ‘734 patent) is sample and parameterized to a generate an autoregressive prior which can be used as a fourth feedback connection to condition a prior model for the next latent space embedded vector (zt+1). Similar to the previously described feedback connection, each of these optional connections is present both training and inference, and thus are trained as part of the training process thereby improving functionality of the FRAE. [0045] Additional aspects of the present disclosure are described in more detail below. [0046] FIG. 1 illustrates an example implementation of a system-on-a-chip (SoC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computational device (e.g., neural network with weights), delays, frequency bin information, task information, among other information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with a CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and/or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118. [0047] The SoC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and the like, and a multimedia processor 112 that may, Qualcomm Ref. No.2400252WO for example, detect and recognize gestures, speech, and/or other interactive user action(s) or input(s). In one implementation, the NPU 108 is implemented in the CPU 102, DSP 106, and/or GPU 104. The SoC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and/or a signal synthesis system 120. For example, the signal synthesis system 120 can be implemented or configured as a speech synthesis system, including a neural speech decoder and/or a neural homomorphic vocoder (NHV) system, which can be used to generate speech (e.g., perform speech synthesis), can be implemented in a text-to-speech (TTS) system, etc. In some examples, the sensor processor 114 can be associated with or connected to one or more sensors for providing sensor input(s) to sensor processor 114. For example, the one or more sensors and the sensor processor 114 can be provided in, coupled to, or otherwise associated with a same computing device. [0048] In some examples, the one or more sensors can include one or more microphones for receiving sound (e.g., an audio input), including sound or audio inputs that can be used to perform various speech synthesis tasks and/or to generate a reconstructed speech signal from the audio input, etc. In some cases, the sound or audio input received by the one or more microphones (and/or other sensors) may be digitized into data packets for analysis and/or transmission. The audio input may include ambient sounds in the vicinity of a computing device associated with the SoC 100 and/or may include speech from a user of the computing device associated with the SoC 100. In some cases, a computing device associated with the SoC 100 can additionally, or alternatively, be communicatively coupled to one or more peripheral devices (not shown) and/or configured to communicate with one or more remote computing devices or external resources, for example using a wireless transceiver and a communication network, such as a cellular communication network. [0049] SoC 100, DSP 106, NPU 108 and/or signal synthesis (e.g., neural speech decoder, NHV, etc.) system 120 may be configured to perform audio signal processing. For example, the signal synthesis system 120 may be configured to perform steps for speech synthesis and/or neural homomorphic vocoding, etc. As another example, one or more portions of the steps, such as feature generation, for speech synthesis and/or NHV may be performed by the signal synthesis system 120 while the DSP 106/NPU 108 performs other steps, such as steps using one or more machine learning networks and/or machine learning techniques according to aspects of the present disclosure and as described herein. Qualcomm Ref. No.2400252WO [0050] FIG.2 depicts an example of a feature generator 200, in accordance with aspects of the present disclosure. It should be understood that many techniques may be used to generate feature vectors for an audio input and that feature generator 200 is just a single example of a technique that may be used to generate feature vectors. [0051] Feature generator 200 receives an audio signal at signal pre-processor 202. As above, the audio signal may be from an audio source of an electronic device, such a microphone. Signal pre- processor 202 may perform various pre-processing steps on the received audio signal. For example, signal pre-processor 202 may split the audio signal into parallel audio signals and delay one of the signals by a predetermined amount of time to prepare the audio signals for input into a Fast-Fourier Transform (FFT) circuit. As another example, signal pre-processor 202 may perform a windowing function, such as a Hamming, Hann, Blackman-Harris, Kaiser-Bessel window function, or other sine-based window function, which may improve the performance of further processing stages, such as signal domain transformer 204. Generally, a windowing (or window) function in may be used to reduce the amplitude of discontinuities at the boundaries of each finite sequence of received audio signal data to improve further processing. As another example, signal pre-processor 202 may convert the audio signal data from parallel to serial, or vice versa, for further processing. The pre-processed audio signal from the signal pre-processor 202 may be provided to signal domain transformer 204, which may transform the pre-processed audio signal from a first domain into a second domain, such as from a time domain into a frequency domain. [0052] In some aspects, signal domain transformer 204 implements a Fourier transform, such as a Fast-Fourier transform (FFT). For example, in some cases, the Fast Fourier transform may be a 16-band (or bin, channel, or point) FFT, which generates a compact feature set that may be efficiency processed by a model. In some cases, a Fourier transform provides fine spectral domain information about the incoming audio signal as compared to conventional single channel processing, such as conventional hardware SNR threshold detection. The result of signal domain transformer 204 is a set of audio features, such as a set of voltages, powers, or energies per frequency band in the transformed data. [0053] The set of audio features may then be provided to signal feature filter 206, which may reduce the size of or compress the feature set in the audio feature data. In some aspects, signal feature filter 206 may discard certain features from the audio feature set, such as symmetric or Qualcomm Ref. No.2400252WO redundant features from multiple bands of a multi-band FFT. Discarding this data reduces the overall size of the data stream for further processing and may be referred to a compressing the data stream. For example, in some cases, a 16-band FFT may include 8 symmetric or redundant bands of after the powers are squared because audio signals are real. Thus, signal feature filter 206 may filter out the redundant or symmetric band information and output an audio feature vector 208. In some cases, output of the signal feature filter may be compressed or otherwise processed prior to output as the audio feature vector 208. The audio feature vector 208 may be provided to a speech synthesis system (e.g., such as the signal synthesis system 120 of FIG.1) for processing by speech synthesis or NHV model. [0054] Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and/or other use. FIG.3 is a block diagram illustrating an example of a voice coding system 350 (which can also be referred to as a voice or speech coder or a voice coder- decoder (codec)). A voice encoder 352 of the voice coding system 350 can use a voice coding algorithm to process a speech signal 351. The speech signal 351 can include a digitized speech signal generated from an analog speech signal from a given source. For instance, the digitized speech signal can be generated using a filter to eliminate aliasing, a sampler to convert to discrete- time, and an analog-to-digital converter for converting the analog signal to the digital domain. The resulting digitized speech signal (e.g., speech signal 351) is a discrete-time speech signal with sample values (referred to herein as samples) that are also discretized. [0055] Using the voice coding algorithm, the voice encoder 352 can generate a compressed signal (including a lower bit-rate stream of data) that represents the speech signal 351 using as few bits as possible, while attempting to maintain a certain quality level for the speech. The voice encoder 352 can use any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., Code-excited linear prediction (CELP), algebraic-CELP (ACELP), or other linear prediction technique) or other voice coding algorithm. [0056] The voice encoder 352 can compress the speech signal 351 in an attempt to reduce the bit-rate of the speech signal 351. The bit-rate of a signal is based on the sampling frequency and the number of bits per sample. For instance, the bit-rate of a speech signal can be determined as ^^^^ ൌ ^^ ∗ ^^, where BR is the bit-rate, S is the sampling frequency, and b is the number of bits per Qualcomm Ref. No.2400252WO sample. In one illustrative example, at a sampling frequency (S) of 8 kilohertz (kHz) and at 16 bits per sample (b), the bit-rate (BR) of a signal would be a bit-rate of 128 kilobits per second (kbps). [0057] The compressed speech signal can then be stored and/or sent to and processed by a voice decoder 354. In some examples, the voice decoder 354 can communicate with the voice encoder 352, such as to request speech data, send feedback information, and/or provide other communications to the voice encoder 352. In some examples, the voice encoder 352 or a channel encoder can perform channel coding on the compressed speech signal before the compressed speech signal is sent to the voice decoder 354. For instance, channel coding can provide error protection to the bitstream of the compressed speech signal to protect the bitstream from noise and/or interference that can occur during transmission on a communication channel. [0058] The voice decoder 354 can decode the data of the compressed speech signal and construct a reconstructed speech signal 355 that approximates the original speech signal 351. The reconstructed speech signal 355 includes a digitized, discrete-time signal that can have the same or similar bit-rate as that of the original speech signal 351. The voice decoder 354 can use an inverse of the voice coding algorithm used by the voice encoder 352, which as noted above can include any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., CELP, ACELP, or other suitable linear prediction technique) or other voice coding algorithm. In some cases, the reconstructed speech signal 355 can be converted to continuous-time analog signal, such as by performing digital-to-analog conversion and anti-aliasing filtering. [0059] Voice coders can exploit the fact that speech signals are highly correlated waveforms. The samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame. In one illustrative example, each frame can be 10-20 milliseconds (ms) in length. [0060] Various voice coding algorithms can be used to encode a speech signal. For instance, code-excited linear prediction (CELP) is one example of a voice coding algorithm. The CELP model is based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech. The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter). Qualcomm Ref. No.2400252WO [0061] In general, CELP uses a linear prediction (LP) model to model the vocal tract, and uses entries of a fixed codebook (FCB) as input to the LP model. For instance, long-term linear prediction can be used to model pitch of a speech signal, and short-term linear prediction can be used to model the spectral shape (phoneme) of the speech signal. Entries in the FCB are based on coding of a residual signal that remains after the long-term and short-term linear prediction modeling is performed. For example, long-term linear prediction and short-term linear prediction models can be used for speech synthesis, and a fixed codebook (FCB) can be searched during encoding to locate the best residual for input to the long-term and short-term linear prediction models. The FCB provides the residual speech components not captured by the short-term and long-term linear prediction models. A residual, and a corresponding index, can be selected at the encoder based on an analysis-by-synthesis process that is performed to choose the best parameters so as to match the original speech signal as closely as possible. The index can be sent to the decoder, which can extract the corresponding LTP residual from the FCB based on the index. [0062] FIG.4 is a block diagram illustrating an example of a CELP-based voice coding system 470, including a voice encoder 472 and a voice decoder 474. The voice encoder 472 can obtain a speech signal 471 and can segment the samples of the speech signal into frames and sub-frames. For instance, a frame of N samples can be divided into sub-frames. In one illustrative example, a frame of 440 samples can be divided into four sub-frames each having 60 samples. For each frame, sub-frame, or sample, the voice encoder 472 chooses the parameters (e.g., gain, filter coefficients or linear prediction (LP) coefficients, etc.) for a synthetic speech signal so as to match as much as possible the synthetic speech signal with the original speech signal. [0063] The voice encoder 472 can include a short-term linear prediction (LP) engine 480, a long- term linear prediction (LTP) engine 482, and a fixed codebook (FCB) 484. The short-term LP engine 480 models the spectral shape (phoneme) of the speech signal. For example, the short-term LP engine 480 can perform a short-term LP analysis on each frame to yield linear prediction (LP) coefficients. In some examples, the input to the short-term LP engine 480 can be the original speech signal or a pre-processed version of the original speech signal. In some implementations, the short-term LP engine 480 can perform linear prediction for each frame by estimating the value of a current speech sample based on a linear combination of past speech samples. For example, a speech signal s(n) can be represented using an autoregressive (AR) model, such as ^^^^^^ ൌ ^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^ , where each sample is represented as a linear combination of the Qualcomm Ref. No.2400252WO previous m samples plus a prediction error term ^^^^^^. The weighting coefficients a1, a2, through am can be referred to as the LP coefficients. The prediction error term ^^^^^^ can be found as follows: ^^^^^^ ൌ ^^^^^^ െ ∑ ^ ^ୀ^ ^^^^^^^^ െ ^^^ . By minimizing the mean square prediction error with respect to the filter coefficients, the short-term LP engine 480 can obtain the LP coefficients. The LP coefficients can be used to form an analysis filter as given in Equation 1, below: ^ ^^^^^^ ൌ 1 െ ^ ^^ ^ ି^ ^ ^ Eq. (1) [0064] The short-term LP engine 480 can solve for ^^^^^^ (which can be referred to as a transfer function) by computing the LP coefficients (^^^) that minimize the error in the above AR model equation (^^^^^^) or other error metric. In some implementations, the LP coefficients can be determined using a Levinson–Durbin method, a Leroux–Gueguen algorithm, or other suitable technique. In some examples, the voice encoder 472 can send the LP coefficients to the voice decoder 474. In some examples, the voice decoder 474 can determine the LP coefficients, in which case the voice encoder 472 may not send the LP coefficients to the voice decoder 474. In some examples, Line Spectral Frequencies (LSFs) can be computed instead of or in addition to LP coefficients. [0065] The LTP engine 482 models the pitch of the speech signal. Pitch is a feature that determines the spacing or periodicity of the impulses in a speech signal. For example, speech signals are generated when the airflow from the lungs is periodically interrupted by movements of the vocal cords. The time between successive vocal cord openings corresponds to the pitch period. The LTP engine 482 can be applied to each frame or each sub-frame of a frame after the short- term LP engine 480 is applied to the frame. The LTP engine 482 can predict a current signal sample from a past sample that is one or more pitch periods apart from a current sample (hence the term “long-term”). For instance, the current signal sample can be predicted as ^^^^^^^ ൌ ^^^^^^^^ െ ^^^, where T denotes the pitch period, ^^^ denotes the pitch gain, and ^^^^^ െ ^^^ denotes an LP residual for a previous sample one or more pitch periods apart from a current sample. Pitch period can be estimated at every frame. By comparing a frame with past samples, it is possible to identify the period in which the signal repeats itself, resulting in an estimate of the actual pitch period. The LTP engine 482 can be applied separately to each sub-frame. Qualcomm Ref. No.2400252WO [0066] The FCB 484 can include a number (denoted as L) of long-term linear prediction (LTP) residuals. An LTP residual includes the speech signal components that remain after the long-term and short-term linear prediction modeling is performed. The LTP residuals can be, for example, fixed or adaptive and can contain deterministic pulses or random noise (e.g., white noise samples). The voice encoder 472 can pass through the number L of LTP residuals in the FCB 484 a number of times for each segment (e.g., each frame or other group of samples) of the input speech signal, and can calculate an error value (e.g., a mean-squared error value) after each pass. The LTP residuals can be represented using codevectors. The length of each codevector can be equal to the length of each sub-frame, in which case a search of the FCB 484 is performed once every sub- frame. The LTP residual providing the lowest error can be selected by the voice encoder 472. The voice encoder 472 can select an index corresponding to the LTP residual selected from the FCB 484 for a given sub-frame or frame. The voice encoder 472 can send the index to the voice decoder 474 indicating which LTP residual is selected from the FCB 484 for the given sub-frame or frame. A gain associated with the lowest error can also be selected, and send to the voice decoder 474. [0067] The voice decoder 474 includes an FCB 494, an LTP engine 492, and a short-term LP engine 490. The FCB 494 has the same LTP residuals (e.g., codevectors) as the FCB 484. The voice decoder 474 can extract an LTP residual from the FCB 494 using the index transmitted to the voice decoder 474 from the voice encoder 472. The extracted LTP residual can be scaled to the appropriate level and filtered by the LTP engine 492 and the short-term LP engine 490 to generate a reconstructed speech signal 475. The LTP engine 492 creates periodicity in the signal associated with the fundamental pitch frequency, and the short-term LP engine 490 generates the spectral envelope of the signal. [0068] Other linear predictive-based coding systems can also be used to code voice signals, including enhanced voice services (EVS), adaptive multi-rate (AMR) voice coding systems, mixed excitation linear prediction (MELP) voice coding systems, linear predictive coding-10 (LPC-10), among others. [0069] A voice codec for some applications and/or devices (e.g., Internet-of-Things (IoT) applications and devices) may be needed to deliver higher quality coding of speech signals at low bit-rates, with low complexity, and with low memory requirements. Existing linear predictive- based codecs cannot meet such requirements. For example, ACELP-based coding systems provide Qualcomm Ref. No.2400252WO high quality, but do not provide low bit-rate or low complexity/low memory. Other linear- predictive coding systems provide low bit-rate and low complexity/low memory, but do not provide high quality. In some cases, machine learning systems (e.g., using a neural network model) can be used to generate reconstructed voice or audio signals. For example, using features extracted from a frame of audio data, a neural network-based voice decoder can generate coefficients for at least one linear filter. The linear filter can then be used to generate a reconstructed signal. However, such a neural network-based voice decoder can be highly complex and resource intensive. For instance, the neural network-based voice decoder will have to perform the operations of a linear predictive filter (LPC), such as the short-term LP engine 480 of FIG.4. Such LPC operations can include complex operations that require the use of a large amount of computing resources by the neural network-based voice decoder. [0070] FIG.6 is a diagram illustrating an example of a voice decoding signal synthesis system 500 utilizing a linear time-varying filter 504 with coefficients generated using a neural network (NN) filter estimator 502 and a separate linear predictive coding (LPC) filter 506. The voice decoding signal synthesis system 500 is configured to decode data of the compressed speech signal to generate a reconstructed speech signal ^̂^^^^^ (also referred to as a synthesized speech sample) for a current time instant n that approximates an original speech signal that was previously compressed by a voice encoder (not shown). The voice encoder can be similar to and can perform some or all of the functions of the voice encoder 252 described above with respect to FIG.2B, or other type of voice encoder. For example, the voice encoder can include a short-term LP engine, an LTP engine, and an FCB. In another example, the voice encoder can include a magnitude spectrum generator (e.g., Mel-scale magnitude spectrum or full spectrum magnitude), a short-term linear prediction (LP) engine, and a pitch tracker that detects a fundamental pitch harmonic frequency of the speech and pitch correlation. [0071] The voice encoder can extract (and in some cases quantize) a set of features (referred to as a feature set) from the speech signal, and can send the extracted (and in some cases quantized) feature set to the voice decoding signal synthesis system 500. The features that are computed by the voice encoder can depend on a particular encoder implementation used. Various illustrative examples of feature sets are provided below according to different encoder implementations, which can be extracted by the voice encoder (and in some cases quantized), and sent to the voice decoding signal synthesis system 500. However, one of ordinary skill will appreciate that other Qualcomm Ref. No.2400252WO feature sets can be extracted by the voice encoder. For example, the voice encoder can extract any set of features, can quantize that feature set, and can send the feature set to the voice decoding signal synthesis system 500. [0072] As noted above, various combinations of features can be extracted as a feature set by the voice encoder. For example, a feature set can include one or any combination of the following features: Linear Prediction (LP) coefficients; Line Spectral Pairs (LSPs); Line Spectral Frequencies (LSFs); pitch lag with integer or fractional accuracy; pitch gain; pitch correlation; Mel-scale frequency cepstral coefficients (also referred to as Mel cepstrum) of the speech signal; Bark-scale frequency cepstral coefficients (also referred to as bark cepstrum) of the speech signal; Mel-scale frequency cepstral coefficients of the LTP residual; Bark-scale frequency cepstral coefficients of the LTP residual; a spectrum (e.g., Discrete Fourier Transform (DFT) or other spectrum) of the speech signal; and/or a spectrum (e.g., DFT or other spectrum) of the LTP residual; voicing level of each frequency band of each speech frame; fundamental frequency of pitch harmonics; pitch correlation of each speech frame; time domain pitch lag of each speech frame. [0073] For any one or more of the other features listed above, the voice encoder can use any estimation and/or quantization method, such as an engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec) or a neural network-based estimation and/or quantization scheme (e.g., convolutional or fully-connected (dense) or recurrent Autoencoder, or other neural network-based estimation and/or quantization scheme). The voice encoder can also use any frame size, frame overlap, and/or update rate for each feature. The voice encoder can also include extra redundancies in the features to ensure robustness of operation against packet losses. Examples of estimation and quantization methods for each example feature are provided below for illustrative purposes, where other examples of estimation and quantization methods can be used by the voice encoder. [0074] As noted above, one example of features that can be extracted from a voice signal by the voice encoder includes linear prediction (LP) coefficients and/or line spectral frequencies (LSFs). Various estimation techniques can be used to compute the LP coefficients and/or LSFs. For example, the voice encoder can estimate LP coefficients (and/or LSFs) from a speech signal using the Levinson-Durbin algorithm. In some examples, the LP coefficients and/or LSFs can be Qualcomm Ref. No.2400252WO estimated using an autocovariance method for LP estimation. In some cases, the LP coefficients can be determined, and an LP to LSF conversion algorithm can be performed to obtain the LSFs. Any other LP and/or LSF estimation engine or algorithm can be used, such as an LP and/or LSF estimation engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec). [0075] Various quantization techniques can be used to quantize the LP coefficients and/or LSFs. For example, the voice encoder can use a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), or other vector quantization technique to quantize the LP coefficients and/or LSFs. In some cases, a predictive or adaptive SSVQ or MSVQ (or other vector quantization technique) can be used to quantize the LP coefficients and/or LSFs. In another example, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the LP coefficients and/or LSFs. Any other LP and/or LSF quantization engine or algorithm can be used, such as an LP and/or LSF quantization engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec). [0076] Another example of features that can be extracted from a voice signal by the voice encoder includes pitch lag (integer and/or fractional), pitch gain, and/or pitch correlation. Various estimation techniques can be used to compute the pitch lag, pitch gain, and/or pitch correlation. For example, the voice encoder can estimate the pitch lag, pitch gain, and/or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation estimation engine or algorithm (e.g. autocorrelation-based pitch lag estimation). For example, the voice encoder can use a pitch lag, gain, and/or correlation estimation engine (or algorithm) from any suitable voice codec (e.g. EVS, AMR, or other voice codec). Various quantization techniques can be used to quantize the pitch lag, pitch gain, and/or pitch correlation. For example, the voice encoder can quantize the pitch lag, pitch gain, and/or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation quantization engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec). In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the pitch lag, pitch gain, and/or pitch correlation features. [0077] Another example of features that can be extracted from a voice signal by the voice encoder includes the Mel cepstrum coefficients and/or Bark cepstrum coefficients of the speech signal, and/or the Mel cepstrum coefficients and/or Bark cepstrum coefficients of the LTP residual. Qualcomm Ref. No.2400252WO Various estimation techniques can be used to compute the Mel cepstrum coefficients and/or Bark cepstrum coefficients. For example, the voice encoder can use a Mel or Bark frequency cepstrum technique that includes Mel or Bark frequency filter banks computation, filter bank energy computation, logarithm application, and discrete cosine transform (DCT) or truncation of the DCT. Various quantization techniques can be used to quantize the Mel cepstrum coefficients and/or Bark cepstrum coefficients. For example, vector quantization (single stage or multistage) or predictive/adaptive vector quantization can be used. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the Mel cepstrum coefficients and/or Bark cepstrum coefficients. Any other suitable cepstrum quantization methods can be used. [0078] Another example of features that can be extracted from a voice signal by the voice encoder includes the spectrum of the speech signal and/or the spectrum of the LTP residual. Various estimation techniques can be used to compute the spectrum of the speech signal and/or the LTP residual. For example, a Discrete Fourier transform (DFT), a Fast Fourier Transform (FFT), or other transform of the speech signal can be determined. Quantization techniques that can be used to quantize the spectrum of the voice signal can include vector quantization (single stage or multistage) or predictive/adaptive vector quantization. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the spectrum. Any other suitable spectrum quantization methods can be used. [0079] As noted above, any one of the above-described features or any combination of the above-described features can be estimated, quantized, and sent by the voice encoder to the voice decoding signal synthesis system 500 depending on the particular encoder implementation that is used. In one illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, and pitch correlation. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal. In another illustrative example, the voice encoder can Qualcomm Ref. No.2400252WO estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the LTP residual. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the LTP residual. [0080] The voice decoding signal synthesis system 500 includes a neural network filter estimator 502, a linear time-varying filter 504 generated by the neural network, and a linear predictive coding (LPC) filter 506. The LPC filter 506 can include a time-varying LPC filter. The neural network filter estimator 502 is trained to generate filter coefficients for the linear time-varying filter 504. The neural network model of the neural network filter estimator 502 can include any neural network architecture that can be trained to model the filter coefficients for the linear time-varying filter 504. Examples of neural network architectures that can be included in the neural network filter estimator 502 include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and/or other type(s) of neural network architectures or models. [0081] The voice decoding signal synthesis system 500 (e.g., the neural network model of the neural network filter estimator 502) can be trained using any suitable neural network training technique. In some examples, the neural network model of the neural network filter estimator 502 can be trained using supervised learning techniques based on backpropagation. For instance, corresponding input and target output pairs can be provided to the neural network filter estimator 502 for training. In one example, for each time instant n, the input to the neural network filter estimator 502 can include log-Mel-frequency spectrum features or coefficients (e.g., 80 log-Mel features ^^^^^, ^^^ 501 shown in FIG. 5). In some examples, the target output (or label or ground truth) for training the neural network filter estimator 502 can include the target speech sample ^^^^^^ 505 for the current time instant n, as shown in FIG. 5. In such examples, a loss 507 will be computed based on the reconstructed sample ^̂^ ^ ^^ ^ and the target output speech sample ^^ ^ ^^ ^ 505 for time instant n. In some examples, the target output can include a speech signal ^^^^^^ that is generated after passing the target speech through an LPC analysis filter (inverse of LPC filter 506). Qualcomm Ref. No.2400252WO In this case, the loss will be computed based on the output ^̂^^^^^ generated by the linear time- varying filter generated by the neural network 504 and the target output ^^^^^^ (e.g., where both are in speech residual domain). [0082] Backpropagation can be performed to train the neural network filter estimator 502 using the inputs and the target output. Backpropagation can include a forward pass, a loss function, a backward pass, and a parameter update to update one or more parameters (e.g., weight, bias, or other parameter). The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of inputs until the neural network filter estimator 502 is trained well enough so that the weights (and/or other parameters) of the various layers are accurately tuned. [0083] In some aspects, training of the neural network filter estimator 502 and/or training of one or more of the machine learning systems or neural networks described herein can be performed using online training (e.g., in some case on-device training), offline training, and/or various combinations of online and offline training. [0084] In some cases, online may refer to time periods during which the input data is processed, for instance for performance of generative voice codec processing implemented by the systems and techniques described herein. In some examples, offline may refer to idle time periods or time periods during which input data is not being processed. Additionally, offline may be based on one or more time conditions (e.g., after a particular amount of time has expired, such as a day, a week, a month, etc.) and/or may be based on various other conditions such as network and/or server availability, etc., among various others. In some aspects, offline training of a machine learning model (e.g., a neural network model) can be performed by a first device (e.g., a server device) to generate a pre-trained model, and a second device can receive the trained model from the second device. In some cases, the second device (e.g., a mobile device, an XR device, a vehicle or system/component of the vehicle, or other device) can perform online (or on-device) training of the pre-trained model to further adapt or tune the parameters of the model. [0085] In some cases, the forward pass can include passing the input data (e.g., the log-Mel- frequency spectrum features or coefficients, such as the 80 log-Mel features ^^^^^, ^^^ 501 shown in FIG.5) through the neural network filter estimator 502. The weights of the neural network model are initially randomized before the neural network filter estimator 502 is trained. For a first training Qualcomm Ref. No.2400252WO iteration for the neural network filter estimator 502, the output will likely include values that do not give preference to any particular output due to the weights being randomly selected at initialization. With the initial weights, the neural network filter estimator 502 is unable to determine low level features and thus cannot make an accurate estimation of the filter coefficients for the linear time-varying filter 504. A loss function can be used to analyze the loss 507 (or error) in the reconstructed or synthesized sample output ^̂^^^^^. Any suitable loss function definition can be used. One example of a loss function includes a mean squared error (MSE). The MSE is defined as ^^௧^௧^^ ൌ ∑ ^ ଶ ^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^ଶ , which calculates the sum of one-half times the actual answer the predicted (output) answer squared. The loss can be set to be equal to the value of ^^௧^௧^^. Other loss functions may include a difference of magnitude spectrums between the target and output signals, where the difference may be computed as absolute difference, squared difference, or logarithmic difference between the magnitude spectrum of each speech frame, and then aggregated over all speech frames. [0086] The loss (or error) will be high for the first training iterations since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. The neural network filter estimator 502 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL/dW, where W represents the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as ^^ ൌ ^^^ െ ^^ ௗ^, where w denotes a weight, wi ௗ^ denotes the initial weight, and η denotes a learning rate. rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates. In some examples, to train the neural network filter estimator 502, a multi-resolution STFT loss ^^ and adversarial losses ^^ and ^^^ can be computed from ^^^^^^ and ^^^^^^ . Because linear time-varying filters are fully differentiable, gradients can propagate back to the neural network filter estimator 502. Qualcomm Ref. No.2400252WO [0087] Using the filter coefficients generated by the neural network filter estimator 502, the linear time-varying filter 504 can process an excitation signal 503 to generate another signal ^̂^^^^^. The signal ^̂^^^^^ can be used as an excitation signal to excite the LPC filter 506. The linear time- varying filter 504 is a linear filter, which preserves the linearity property between inputs and outputs. For instance, a linear filter is associated with a mapping ℒ:ℝℤ → ℝℤ, ^^^^^^ → ^^^^^^ ൌ ℒ^^^^^^^^ that has the following property: for any ^^^^^^^, ℝℤ and any ^^, ^^ ∈ ℝ , ℒ^^^ ^^^^^^^ ^ ^^ ^^ଶ^^^^^ ൌ ^^ ℒ^^^^^^^^^ ^ ^^ ℒ^^^ଶ^^^^^. In one if input1 produces output1 and input2 produces output2, then a combined input of (input1+input2) will produce output = output1+output2. The time-varying nature of the linear time-varying filter 504 indicates that the filter response depends on the time of excitation of the linear time-varying filter 504 (e.g., a new set of coefficients used to filter each frame (block of time) of input at the time of excitation). In some cases, time-varying linear filters can be characterized by the set of impulse responses at each time lag ℎ^^^^^ ൌ ℒ^^^^^^ െ ^^^^ for each ^^ ∈ ℤ. In some examples, the output of a time varying linear filter is ℒ^^^^^^^^ ൌ ∑^ ^^^^^^ ⋆ ℎ^^^^^ (where ⋆ is a convolutional operator) or some heuristic combination of filter and impulse responses, e.g. overlap-add on windowed and filtered signal segments, etc. [0088] The LPC filter 506 can use the signal ^̂^^^^^ as input to generate the reconstructed or synthesized speech sample ^̂^^^^^ for the current time instant n. The LPC filter 506 is a linear filter and in some cases is time varying, as defined above with respect to the linear time-varying filter 504. The LPC filter 506 can be a form of time-varying filter used for processing of speech. The LPC filter 506 includes filter coefficients for each speech frame that can be computed using the autocorrelation of a speech or audio signal. [0089] In some examples, the LPC filter 506 can be used to model the spectral shape (or phenome or envelope) of the speech signal. For example, at the voice encoder, a signal ^^^^^^ can be filtered by an autoregressive (AR) process synthesizer to obtain an AR signal. As described above, a linear predictor can be used to predict the AR signal (which can be denoted as prediction ^̂^^^^^) as a linear combination of the previous m samples as follows: ^ ^^^^^^^ Qualcomm Ref. No.2400252WO [0090] where the ^^^^ terms (^^^^, ^^^, … ^^^^) are estimates of the AR parameters (also referred to as LP coefficients). A residual signal can be the difference between the original AR signal and the predicted AR signal represented as prediction ^̂^^^^^ (e.g., the difference between the actual sample and the predicted sample). At the voice encoder, the linear prediction coding can be used to find the best linear prediction coefficients for minimizing a quadratic error function, and thus the error. The linear prediction process removes the short-term correlation from the speech signal. The linear prediction coefficients are an efficient way to represent the short-term spectrum of the speech signal. At the voice decoding signal synthesis system 500, the LPC filter 506 determines the prediction ^̂^^^^^ for the current sample n using computed or received coefficients and A transfer function ^^^^^^^. For instance, in some examples, the LPC filter coefficients are received from the encoder. In other examples, the voice decoding signal synthesis system 500 can derive the LPC filter coefficients, such as using other features (e.g., Mel spectrum features) sent by an encoder to the voice decoding signal synthesis system 500. For instance, the voice decoding signal synthesis system 500 can use Mel spectrum features 501 to derive the LPC filter coefficients for the LPC filter 506. The LPC filter 506 can determine the final reconstructed (or predicted) sample ^̂^^^^^ using the output ^̂^^^^^ from the linear time-varying filter 504 (for the current sample n). [0091] In some examples, further components can be used along with the neural network filter estimator 502 and the linear time-varying filter 504, such as an impulse train generator 514 and a random noise generator 516. In some examples, the linear time-varying filter 504 can include a harmonic linear time-varying filter 518 and a noise linear time-varying filter 520. A voice encoder can include a pitch tracker 510 and a feature extraction engine 512. In some examples, the feature extraction engine 512 can be the same as or similar to the feature generator 200 of FIG.2. In some cases, the original speech signal ^^ and reconstructed signal ^^ are divided into non-overlapping frames with frame length L. The term ^^ can be defined as a frame index, the term ^^ can be defined as a discrete time index, and the term ^^ can be defined as a feature index. The total number of frames ^^ and total number of sampling points ^^ may follow ^^ ൌ ^^ ൈ ^^. In ^^^, ^^, ℎ^, ℎ^, 0 ^ ^^ െ 1. The terms ^^, ^^, ^^, ^^, ^^^, ^^^ are finite duration signals, in which 0 ^ ^^ ^ ^^ െ 1. Impulse responses ℎ^, and ℎ^ may be infinitely long, in which ^^ ∈ ℤ. Impulse response h may be causal, in which ^^ ∈ ℤ E Z and ^^ ^ 0. Qualcomm Ref. No.2400252WO [0092] To perform the speech synthesis process, the impulse train generator 514 can generate an impulse train ^^^^^^ from a frame-wise fundamental frequency ^^^ ^^^^ output by the pitch tracker 510. In one illustrative example, the impulse train generator 514 can generate alias-free discrete time impulse trains using additive synthesis. For instance, as illustrated in equation (1) below, the impulse train generator 514 can use a low-passed sum of sinusoids to generate an impulse train: ^ 2^^^^^^^^^ ^ ^^^ ^^^^^^ ^^ ௧ 2^^^^ ^^ ^^ ^ ^ ^^ ^ ^^ ൌ ^ ^ ^ ^^^^^ , ^ ^ ^ 1 [0093] where ^^^^^^^ is reconstructed from ^^^^^^^ with zero-order hold or linear interpolation, ^^^^^^ ൌ ^^^^^/^^^^, and ^^^ is the sampling rate. In some cases, the computationally complexity of additive synthesis can be reduced with approximations. For example, the impulse train generator 514 or other component (e.g., a processor) of the voice decoding signal synthesis system 500 can round the fundamental periods to the nearest multiples of the sampling period. In such an example, the discrete impulse train is sparse. The impulse train generator 514 can then generate the impulse train sequentially (e.g., one pitch mark at a time). [0094] The pitch tracker 510 can process the input ^^^^^^ for the time instant n to generate the frame-wise fundamental frequency ^^^^^^^ output, which is provided to and processed by the impulse train generator 514 of the voice decoding signal synthesis system 500. The random noise generator 516 of the voice decoding signal synthesis system 500 can sample a noise signal ^^^^^^ from a Gaussian distribution. [0095] The neural network filter estimator 502 can estimate impulse responses ℎ^ ^^^,^^^ and ℎ^ ^^^,^^^ for each frame, given the log-Mel spectrogram ^^^^^, ^^^ extracted from the input ^^^^^^ by the feature extraction engine 512 of the encoder. In some aspects, complex cepstrums (ℎ ^ ^ and ℎ ^ ^) can be used as the internal description of impulse responses (ℎ^ and ℎ^) for the neural network filter estimator 502. Complex cepstrums describe the magnitude response and the group delay of filters simultaneously. The group delay of filters affects the timbre of speech. In some cases, instead of using linear-phase or minimum-phase filters, the neural network filter estimator 502 can use mixed-phase filters, with phase characteristics learned from the dataset. Qualcomm Ref. No.2400252WO [0096] In some examples, the length of a complex cepstrum can be restricted, essentially restricting the levels of detail in the magnitude and phase response. Restricting the length of a complex cepstrum can be used to control the complexity of the filters. In some cases, the neural network filter estimator 502 can predicts low-frequency coefficients, in which the high-frequency cepstrum coefficients can be set to zero. In one illustrative example, two 10 millisecond (ms) long complex cepstrums are predicted in each frame. In some cases, the neural network filter estimator 502 can use a discrete Fourier transform (DFT) and an inverse-DFT (IDFT) to generate the impulse responses ℎ^ and ℎ^. In some cases, the neural network filter estimator 502 can approximate an infinite impulse response (IIR) (ℎ^^^^,^^^ and ℎ^^^^,^^^) using finite impulse responses (FIRs). The DFT size can be set to at least a threshold size (e.g., N=1024) to avoid aliasing. [0097] Using the impulse response ℎ^ ^^^,^^^, the harmonic LTV filter 518 can filter the impulse train ^^^^^^ from the impulse train generator 514 to generate a harmonic component ^^^^^^^. Using the impulse response ℎ^ ^^^,^^^, the noise LTV filter 520 can filter the noise signal ^^^^^^ to generate a noise component ^^^^^^^. The voice decoding signal synthesis system 500 can combine (e.g., by summing/adding or otherwise combining) the output of the harmonic LTV filter 518 (the harmonic component ^^^^^^^) and the output of the noise LTV filter 520 (the noise component ^^^^^^^) can be combined (e.g., summed or otherwise combined) to obtain the excitation signal ^^^^^^. [0098] As discussed above, a computing device may receive audio information via an input device, such as a microphone. In some cases, the received audio information may be sampled at a point in time, or frame. The audio frames may be encoded as encoded audio data that may be transmitted to another device, for example, during a phone call with the other device. The encoded audio data may be decoded and the audio frames may then be played back in sequence to reproduce the audio information. [0099] FIG.6 illustrates an audio waveform 600 representing audio information received by a wireless device, in accordance with aspects of the present disclosure. As shown, the audio waveform 600 may include regions, such as an active speech region 602, a hangover region 604, and an inactive (e.g., silent) region 606. In some cases, a speech encoder (e.g., a coding portion of a codec for encoding audio information) may include a voice activity detection (VAD) functionality which analyzes the received audio information to determine whether there is speech activity during a particular audio frame. For frames in the active speech region 602, VAD may Qualcomm Ref. No.2400252WO indicate that there is speech being detected in these frames and the encoder may encode the audio information to encoded audio data. The encoder may also include an indication in the encoded audio data that the encoded audio information includes active speech. In some cases, an end of the active speech region 602 may correspond to when the speech stops. [0100] The hangover region 604 may follow the active speech region 602. The hangover region 606 may be the region where the speech transitions from active to inactive. For example, when speech stops, the VAD may indicate that speech is no longer being detected. In some cases, an indication from VAD that speech is no longer occurring can be abrupt and result in clipped or abruptly ends of sentences/words, which may sound unnatural or harsh. To avoid this, the encoder may continue to encode the audio information to encoded audio data in the hangover region 604 to allow long and/or lingering sounds (e.g., residual speech which may not be detected as speech by the VAD) to be captured. The encoder may also include an indication in the encoded audio data that the encoded audio information does not include active speech (e.g., is in the hangover region 604). In some cases, the encoded audio information may include some residual speech, but may primarily include background noise, especially toward the later portions of the hangover region 604. In some cases, the audio encoder may have a predefined and/or dynamic/adaptive time period after speech has stopped (e.g., no longer detected by the VAD) for the hangover region 604. In some cases, after a predefined/adaptive/dynamic time period, the hangover region 604 may end. [0101] The inactive region 606 may follow the hangover region 604. In the inactive region 606, the encoder may encode captured audio information in a SID. The SID may be transmitted to the other wireless device to be used by the other device to generate comfort noise (e.g., background noise) while in the inactive region 606. In some cases, a number of SID frames generated/transmitted may be reduced as compared to a number of frames generated/transmitted while in the active speech region 602 and the hangover region 604. For example, certain audio encoders may support adaptive and/or fixed SID intervals where a single SID frame may be transmitted instead of some number of regular audio frames (e.g., when in the active speech region 602 and/or the hangover region 604). For example, in enhanced voice services (EVS), one SID frame may be transmitted during a time period where 8-50 regular audio frames would have been transmitted. While the number of SID frames may be reduced as compared to regular audio frames, Qualcomm Ref. No.2400252WO SID frames are still being regularly transmitted. In some cases, it may be useful to further minimize or eliminate SID frames. [0102] In some cases, SID frames may be minimized or eliminated by predicting and synthesizing comfort noise at the decoder (e.g., decoding device) with minimal number (or no) SID frames. FIG.7 is a block diagram providing an overview of a technique for decoder silence generation without, or with minimal, coded SIDs 700, in accordance with aspects of the present disclosure. In FIG.7, a wireless device may receive an encoded audio data stream and a decoder 701 of the wireless device may decode an audio frame of the encoded audio data stream to obtain decoded active speech from an active region 702 (e.g., of the audio data stream). In some cases, the decoded active speech from the active region 702 may include speech as well as background noise. The decoded active speech from the active region 702 may be passed to a denoiser 704. The denoiser 704 may filter out the background noise to generate filtered active speech. The filtered active speech may then be subtracted 706 from the decoded active speech from the active region 702 to obtain background noise from the decoded active speech from the active region 702. The background noise may be passed to a spectrum estimator 708. The spectrum estimator may estimate information about the background noise 710 such as a spectral shape (S(t)) (e.g., spectral characteristics) and/or gain (G(t)). The information about the background noise 710 may be passed to an inactive synthesizer 712, which may generate synthesized background noise 714 based on the information about the background noise 710 from the obtained decoded active speech from an active region 702. [0103] As an example, in cases where an active region transitions directly to an inactive region (e.g., inactive region 606 of FIG. 6), the encoder may indicate a transition to the inactive region (e.g., marking whether an encoded frame include detected speech or not) without indicating a hangover region. For example, the decoder 701 may stop receiving encoded audio frames (e.g., from another wireless device) when speech is no longer present. As another example, the decoder 701 may receive one SID frame indicating that speech is no longer present. In another example, the decoder 701 may receive an indication that speech is no longer present without also receiving encoded audio information (e.g., a description of the silence in the SID frame). For example, the wireless device performing the decoding may receive a radio access network (RAN) message, such as an RRC message, indicating that speech has ended. Based on the indicated transition to the Qualcomm Ref. No.2400252WO inactive region, the inactive synthesizer 712 may use the information about the background noise 710 from a number of frames (N) prior to the indicated transition to the inactive region to generate the synthesized background noise 714. [0104] In some cases, the decoder 701 may also take in account the hangover region. For example, where the encoder may encode an indication that an audio frame is from a hangover region. The decoder 701 may decode the encoded audio frame to obtain decoded audio frame from the hangover region 716. A classifier 703, based on the indication that the audio frame is from the hangover region, may pass the decoded audio frame from the hangover region 716 to the spectrum estimator 708 to generate information about the background noise 710. Generally, while the hangover region may include some residual speech, the hangover region may primarily include background noise, especially towards an end of the hangover period. [0105] In some cases, then hangover region 716 may be inferred. For example, the encoder may indicate a transition to the inactive region without indicating the hangover region. The decoder 701 and/or classifier 703 may infer that (M) frames prior to the indicated transition to the inactive region comprise the hangover region. The decoded active speech corresponding to those M frames may be input to the spectrum estimator 708 to generate information about the background noise 710. Information about the background noise 710 corresponding to frames from the active region and hangover region may be input to the inactive synthesizer 712. The inactive synthesizer 712 may then generate synthesized background noise 714 based on the information about the background noise 710 in the frames from the active region and/or hangover region. For example, the inactive synthesizer 712 may use the information about the background noise 710 from M frames from the hangover region and N-M frames from the active region. The inactive synthesizer 712 may synthesize the background noise 714 using any technique for synthesizing background noise 714 based on information about the background noise 710. [0106] FIG. 8A is a block diagram illustrating an example inactive synthesizer 800, in accordance with aspects of the present disclosure. In some cases, the inactive synthesizer 800 may correspond with inactive synthesizer 712 of FIG.7. As an example, the inactive synthesizer 800 may use a set of gated recurrent units (GRUs) that accept time series information (e.g., information about the background noise in time aligned audio frames) to predict (e.g., generate) the synthesized background noise. For example, a first GRU 802 may receive a first input frame 804 and a second Qualcomm Ref. No.2400252WO input frame 806. The first input frame 804 may be an audio frame just prior to the transition (e.g., where the transition occurs at time t, the first input frame 804 may occur at time t-1) to the inactive region (e.g., from the hangover region). The second input frame 806 may be an audio frame prior to the first input frame 804 (e.g., from t-2, which may be from the hangover region or the active region). The first input frame 804 and second input frame 806 may be processed by the first GRU 802 to generate (e.g., predict) a first synthesized noise frame 808 for the inactive region (e.g., for time t). In some cases, the first synthesized noise frame 808 may be information about the synthesized noise, such as a spectral shape (e.g., spectral characteristics) and/or gain and the actual background noise may be generated based on the information about the synthesized noise, such as by a digital to analog convertor. [0107] The first synthesized noise frame 808 may be input to a second GRU 810 along with a third input frame 812. The third input frame 812 may be an audio frame prior to the second input frame 806 (e.g., from time t-3, which may be from the hangover region or the active region). The first synthesized noise frame 808 and the third input frame 812 may be processed by the second GRU 810 to generate a second synthesized noise frame 814 for the inactive region (e.g., for time t+1). The second synthesized noise frame 814 may be input, along with a fourth input frame 816 (e.g., from time t-4, which may be from the hangover region or the active region), may be input to a third GRU 818 to generate a third synthesized noise frame 820 (e.g., for time t+2). This process may be repeated to continue generating synthesized noise frames for the inactive region. Of note, if an SID is received, the SID may be input to the inactive synthesizer and used to generate synthesized noise frames, for example, as the first input frame 804. [0108] FIG.8B is a is a block diagram illustrating another example inactive synthesizer 850, in accordance with aspects of the present disclosure. In some cases, the inactive synthesizer 850 may correspond with inactive synthesizer 712 of FIG.7. As an example, the inactive synthesizer 850 may use a set of neural network (NN) based filter estimators to shape random noise. For example, a first NN based filter estimator 852 may receive a first input frame 854. The first input frame 854 may be an audio frame just prior to the transition (e.g., where the transition occurs at time t, the first input frame 854 may occur at time t-1) to the inactive region (e.g., from the hangover region). The first NN based filter estimator 852 may estimate a filter based on the first input frame 854. For example, a digital filter may include filter coefficients that may determine a frequency Qualcomm Ref. No.2400252WO response of the digital filter. Estimating a filter may refer to estimating a filter coefficient. Random noise 856 may be passed through a noise filter 858, which may color (e.g., introduce a spectral shape to) the random noise, resulting in filtered random noise. This filtered random noise may be further shaped by the estimated filter from the first input frame to generate a first synthesized noise frame 860 for the inactive region (e.g., for time t). The first synthesized noise frame 860 may be input to a second NN based filter estimator 862 along with the estimated filter of the first NN based filter estimator 852, and a second input frame 864 to estimate a filter. The second input frame 864 may be an audio frame prior to the first input frame 854 (e.g., from t-2, which may be from the hangover region or the active region). The estimated filter of the second NN based filter estimator 862 may be used to shape random noise 865 passed through a noise filter 866 to generate a second synthesized noise frame 868 for the inactive region (e.g., for time t+1). The second synthesized noise frame 868, a third input frame 870 (e.g., from t-3), and the estimated filter of the second NN based filter estimator 862 may be input to a third NN filter estimator 872 and used to generate a third synthesized noise frame 874 in a manner substantially similar to that used to generate the second synthesized noise frame 868. This process may be repeated to continue generating synthesized noise frames for the inactive region. Of note, if an SID is received, the SID may be input to the inactive synthesizer and used to generate synthesized noise frames, for example, as the first input frame 854. [0109] FIG.9A is a block diagram illustrating a denoiser 900, in accordance with aspects of the present disclosure. The denoiser 900 may correspond with denoiser 704 of FIG. 7. The denoiser 900 may be implemented based on non-negative matrix factorization (NMF). The denoiser 900 may include a pretrained speech dictionary 902. For example, as a part of training, a speech dictionary 902 and a noise dictionary may be trained. The pretrained speech dictionary 902 may include speech-like characteristics (e.g., spectral characteristics that represent parts of speech), while a noise dictionary may include spectral characteristics that represent noise. A product of the pretrained speech dictionary 902 and an activation function 904 may represent the speech portion of the input audio data (e.g., decoded active speech from the active region 702 of FIG. 7) as represented by spectrogram 906. The activation function 904 may be estimated from the input audio data to determine how to combine the different speech-like characteristics to generate the speech portion of the input audio data. Qualcomm Ref. No.2400252WO [0110] As an example, FIG. 9B is a block diagram illustrating NMF denoising 950, in accordance with aspects of the present disclosure. As shown, an input y(t) 952 may include a speech component s(t) and a noise component n(t). A short-time-Fourier-transform 954 may be applied to the input 952 and a magnitude 956 determined. Activations 958 may be estimated based on a noise dictionary 960 and a speech dictionary 962 to determine portions of the input 952 which may be represented by the noise dictionary 960 and other portions that may be represented by the speech dictionary 962. The noise portion may then be estimated as a product of the noise dictionary 960 and the noise activation 958. Similarly, speech portion may be estimated as a product of the speech dictionary 962 and the speech activation 958. In some cases, the speech portion may be returned by the denoiser 900 and subtracted from the original decoded active speech (e.g., subtracted 706 of FIG.7). In other cases, the noise portion may be returned by the denoiser 900 and input directly to a spectrum estimator, such as spectrum estimator 708 of FIG.7 (e.g., without the subtraction 706 of FIG.7). [0111] FIG.10 is a block diagram illustrating a technique for location aware background noise synthetization 1000, in accordance with aspects of the present disclosure. In FIG. 10, location information 1002 for another wireless device sending the encoded audio data stream (e.g., not the wireless device decoding the encoded audio data stream) may be included as a part of, or in addition to, the encoded audio data stream. The location information 1002 may indicate where the other wireless device is located, such as at a train station, on a busy street, at the beach, etc. In some cases, the location information 1002 may be input to an ambient noise prediction engine 1004. The ambient noise prediction engine 1004 may predict a type of ambient noise (e.g., background noise) 1006 that may be present based on the location information 1002. The predicted type of ambient noise 1006 may be input to an inactive synthesizer 1008. The inactive synthesizer 1008 may correspond to inactive synthesizer 712 of FIG.7. In some cases, the inactive synthesizer 712 may take into account the predicted type of ambient noise 1006 when generating the synthesized background noise. [0112] FIG. 11 is a block diagram illustrating a technique for random background noise synthetization 1100, in accordance with aspects of the present disclosure. In FIG. 10, a random noise generator 1102 may be used to generate random noise based on, for example, a random number generator. The random noise from the random noise generator 1102 may be input to an Qualcomm Ref. No.2400252WO inactive synthesizer 1104. The inactive synthesizer 1104 may correspond to inactive synthesizer 712 of FIG.7. In some cases, inactive synthesizer 1104 may use the random noise from the random noise generator 1102 as the synthesized background noise. In other cases, the inactive synthesizer 1104 may mix the random noise from the random noise generator 1102 with synthesized background noise based on the information about the background noise from an active region and/or hangover region. [0113] As indicated above, a spectrum estimator, such as spectrum estimator 708 of FIG.7, may estimate information about the background noise such as the spectral shape and/or gain. In some cases, the spectrum estimator may be configured to perform additional spectral characteristic measurements. FIG. 12 illustrates a spectrum estimator 1200 for measuring additional spectral characteristics, in accordance with aspects of the present disclosure. In some cases, a spectrum estimator 1200 may include a spectrum magnitude engine 1204, a dominant spectral sample calculator 1206, and a spectral descriptor calculator. Noise 1202 (e.g., from a denoiser or from decoded audio from the hangover region) or may be input to the spectrum magnitude engine 1204. The spectrum magnitude engine 1204 may generate a Magnitude spectrum of the noise. The Magnitude spectrum of the noise may describe the noise using frequency and amplitude. The Magnitude spectrum may be input to the dominant spectral sample calculator 1206 and the spectral descriptor calculator 1208. [0114] The dominant spectral sample calculator 1206 apply a probability density function to the Magnitude spectrum and spectral samples higher than a mth percentile may be selected. The dominant spectral sample calculator 1206 may determine k-disjoint contiguous maximum sum segments of spectrums to indicate k-contiguous dominant spectral locations. The dominant spectral locations may be passed to the spectral descriptor calculator 1208. In some cases, spectral characteristics for noise may be, at least in part, determined on dominant spectrums for increased accuracy. [0115] The spectral descriptor calculator 1208 may generate additional descriptions 1210 of the structural features of the noise for the dominant spectrums. For example, the spectral descriptor calculator 1208 may determine first-n moments for the dominant spectrum where the first moment may indicate an energy concentration frequency, the second moment may indicate bandwidth around a centroid, skewness (indicating asymmetry), and kurtosis (indicating peakiness) for output Qualcomm Ref. No.2400252WO as a part of the additional descriptions 1210. The spectral descriptor calculator 1208 may also determine Mel-band/critical band energies of dominant spectral samples, perform auto-regressive (AR)/AR integrated moving average (ARIMA) modeling of the Magnitude spectrums, and determine a spectral slope using spectral entropy of the Magnitude spectrums for output as a part of the additional descriptions 1210. [0116] FIG.13 is a block diagram illustrating another technique for decoder silence generation without, or with minimal, coded SIDs 1300, in accordance with aspects of the present disclosure. FIG. 13 includes a decoder 1301, classifier 1303, spectrum estimator 1308, and inactive synthesizer 1312 which may be substantially similar to the decoder 701, classifier 703, spectrum estimator 708, and inactive synthesizer 712, respectively, of FIG. 7. In FIG. 13, decoded active speech from an active region 1302 may be passed into a noise extractor or ML model 1304. [0117] As an example, a noise extractor may operate in a manner similar to that described with respect to FIGs.9A and 9B. For example, as discussed above, a noise dictionary may be used to extract a noise portion from the decoded active speech from the active region 1302. This noise portion may be passed to the spectrum estimator 1308 to estimate information about the noise portion. [0118] As another example, the decoded active speech from the active region 1302 may be passed to a ML model 1304. The ML model 1304 may be trained to subtract speech or isolate background noise from audio frames. In some cases, ML model 1304 may be a neural network, deep learning network, convolutional neural network, or any other type of machine learning network. In some cases, a single ML model 1304 may be used, or multiple ML models 1304 may be used. The ML model 1304 may output background noise from the decoded active speech from the active region 1302 to the spectrum estimator 1308 to estimate information about the noise portion. [0119] FIG.14 is signal diagram 1400 illustrating signals for decoder silence generation without coded silence descriptions, in accordance with aspects of the present disclosure. Generally, a wireless node 1404, such as a gNB, may configure a wireless device, such as a sending UE 1402 and receiving UE 1406 with timing information, wireless resources, etc. to help optimize scheduling. In some cases, the UE may indicate to the wireless node that no SIDs will be sent. For example, a voice bearer may be established 1408 as between a sending UE 1402 (e.g., UE encoding Qualcomm Ref. No.2400252WO audio data for transmission) and a receiving UE 1406 (e.g., UE decoding received audio data) via the wireless node 1404. While shown as a single, common, wireless node 1404, it should be understood that the sending UE 1402 and receiving UE 1406 may be connected to different wireless nodes. In some cases, a sending UE 1402 may send UE assistance information (UAI) 1410 to the wireless node 1404 indicating that the sending UE 1402 is not sending SIDs. Sending an indication to the wireless node 1404 may be useful to help with scheduling, ensuring compatibility, etc. The wireless node 1404 may transmit a message 1412 (e.g., RRC message or other control plane message) to the receiving UE 1406 indicating that the sending UE 1402 is not sending SIDs. In this example, the sending UE 1402 enter a talk spurt and may transmit a set of voice packets 1414 to the receiving UE 1406. Of note, while described in the context of a UAI message and RRC messages, it should be understood that signaling indicating that SIDs are not being used and/or signaling that SIDs should be used may be transmitted via any RAN protocol and this transmission may utilize a new message/packet data unit (PDU) or utilize an existing message/PDU. [0120] After a last voice packet 1416 of a talk spurt is sent to the receiving UE 1406, the sending UE 1402 may transmit a UAI message indicating an end of the talk spurt 1418 to the wireless node 1404. In some cases, the end of the talk spurt may be after a hangover period. In other cases, the end of the talk spurt may be after speech stops being detected by the sending UE 1402 (e.g., without a hangover period). Based on the UAI message indicating the end of the talk spurt 1418, the wireless node 1404 may transmit a message 1420 (e.g., RRC message) to the receiving UE 1406 indicating that the talk spurt has ended. In some cases, the message 1420 may indicate to the receiving UE 1406, that an inactive region has begun. In some cases, sending a RAN message, such as an RRC message indicating that a talk spurt has ended (e.g., an inactive region has been reached) may be a smaller sized message without a description of the background noise sent on a control plane as compared to a SID, which may be sent as a user plane message. In some cases, messages sent on the user plane may be intended for the operation of user applications. Messages sent on the control plane may carry RAN messages for establishing, managing, and or controlling the different components of the wireless network. [0121] In some cases, to differentiate between a loss of signal and an end of a talk spurt, the encoder may include an indication in packets from the hangover period indicating an end of the Qualcomm Ref. No.2400252WO talk spurt/entry to the hangover period. This way, the decoder, upon receiving an indication of the hangover period, can assume that if the packets stop, it's because the talk spurt has ended (e.g., in the hangover period). Otherwise, if the decoder does not receive the indication of the hangover period and the packets stop, then the decoder may treat this stoppage as a signal loss. [0122] In some cases, activation of SIDs may be configured, activated, and/or deactivated by the wireless node 1404. For example, certain UEs may not support decoder silence generation and the wireless node 1404 or other UE may indicate to the sender UE 1402 to send SIDs. In some cases, the wireless node 1404 may transmit a message 1422 (e.g., RRC message) to the sending UE 1402 indicating that the sending UE should activate sending SIDs. In response to the message 1422, the sending UE 1402 may send a SID 1424 to the receiving UE 1406. [0123] FIG. 15 is a block diagram illustrating an example of an audio codec system 1500 that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer. In some aspects, the audio codec system 1500 can also be referred to as a voice coding system (e.g., vocoder), a signal synthesis system, a voice coding signal synthesis system, etc. In one illustrative example, the audio codec system 1500 can be a generative audio codec (e.g., a generative voice codec) including one or more feedback recurrent autoencoders (FRAEs) that may be used to generate encoded audio data, and one or more neural synthesizers that may be used to generate reconstructed audio data based on the encoded audio data. [0124] For example, the audio codec system 1500 (e.g., a generative voice codec) can include an encoder 1505 (e.g., a transmitter) and a decoder 1510 (e.g., a receiver). The encoder 1505 can be used to generate encoded features that are transmitted to the decoder 1510 over a channel 1540. For example, the encoder 1505 can receive audio 1515 and generate encoded features corresponding to the audio 1515. The encoded features can be generated using a feedback recurrent autoencoder (FRAE) 1525 included in the encoder 1505. In some examples, the encoder 1505 can include one or more FRAEs, where each FRAE of the one or more FRAEs is configured to generate a corresponding one or more encoded features associated with the audio 1515. In some aspects, the audio 1515 may include and/or comprise a speech signal, a voice signal, etc. [0125] The decoder 1510 can receive encoded features (e.g., from the encoder 1505, over the channel 1540) and generate reconstructed audio 1555 based at least in part on the encoded features received from the encoder 1505. The reconstructed audio 1555 may include and/or comprise a Qualcomm Ref. No.2400252WO reconstructed speech signal, a reconstructed voice signal, etc. The reconstructed audio 1555 can be generated using a neural network-based speech synthesizer 1550 (e.g., also referred to as a neural speech synthesizer and/or a neural synthesizer) included in the decoder 1510. The reconstructed audio 1555 generated using the neural speech synthesizer 1550 may also be referred to as synthesized speech. In some examples, the decoder 1510 can include a FRAE decoder 1545, which can be used to decode the encoded features received from the encoder 1505. In some aspects, the FRAE decoder 1545 can be the same as or similar to a decoder implemented by the FRAE 1525 included in the encoder 1505. In some examples, the decoder 1510 can include one or more FRAE decoders 1545, which can correspond to one or more FRAE autoencoders 1525 included in the encoder 1505. For example, the number of FRAE decoders 1545 included in the decoder 1510 can be equal to the number of FRAE autoencoders 1525 included in the encoder 1505. [0126] In some examples, the audio 1515 includes speech. The audio 1515 may be an example of the speech signal 101 of FIG.1, and/or the speech signal 201 of FIG.2, etc. The encoder 1505 of the audio codec system 1500 can be configured to generate one or more encoded representations of the input audio 1515. For example, the one or more encoded representations can include spectral envelope information associated with the audio 1515, and/or can include pitch information associated with the audio 1515, etc. In some aspects, the encoder 1505 can extract spectral envelope features from the audio 1515 using spectral envelope feature extraction engine 1520, and can encode the extracted spectral envelope features using the FRAE 1525. The encoded spectral features z can be transmitted from the FRAE 1525 to the decoder 1510, using the channel 1540. In some examples, the encoder 1505 can extract pitch information from the audio 1515 using pitch extraction engine 1530, and can perform quantization 1535 to generate quantized pitch information associated with the audio 1515. The quantized pitch information can be transmitted from the pitch quantizer 1535 to the decoder 1510, using the channel 1540. [0127] In some aspects, the encoder 1505 can be configured to extract spectral features (e.g., spectral envelope features) from the audio 1515 using the spectral envelope feature extraction engine 1520. In one illustrative example, the spectral envelope feature extraction engine 1520 can perform a cepstrum computation to determine cepstrum information and/or one or more cepstral coefficients corresponding to the audio 1515. In some cases, the spectral features extracted from the audio 1515 using the spectral envelope feature extraction engine 1520 can include mel- frequency cepstral coefficients (MFCC), such as MFCC-24 features (e.g., 24-dimensional MFCC Qualcomm Ref. No.2400252WO features). In some examples, the spectral envelope feature extraction engine 1520 applies one or more mel-scaled filter(s) (e.g., a mel-scaled filterbank), a logarithmic compression, and/or a discrete cosine transform (DCT) to the audio 1515 (e.g., and/or to a magnitude spectrum associated with the audio 1515). The encoder 1505 processes the extracted spectral features using a feedback recurrent autoencoder (FRAE) 1525 to generate encoded features z. The FRAE 1525 includes a decoder and an encoder. The FRAE 1525 can implement feedback of state information h between the decoder and the encoder included in the FRAE 1525. For example, the state information h can be determined or obtained at the decoder of the FRAE 1525, and feedback of the state information h can be performed for the decoder of the FRAE 1525 (e.g., the state information h is fed back to the decoder of the FRAE 1525) and for the encoder of the FRAE 1525 (e.g., the state information h is fed back from the decoder of the FRAE 1525 to the encoder of the FRAE 1525). [0128] In some examples, the spectral envelope feature extraction engine 420 can be configured to generate and/or determine (e.g., extract) one or more types of spectral envelope features. For example, the extracted spectral envelope features may include cepstrum information and/or cepstral coefficients. In some cases, cepstral liftering can be performed based on DCT truncation to exclude or remove pitch information and only capture spectral envelope information in the extracted cepstrum or cepstral coefficients. In some examples, the extracted spectral envelope features can include companded (e.g., log) filterbank energies associated with the input audio 415. For example, companded filterbank energies can be determined based on applying an inverse DCT (e.g., IDCT) to a liftered cepstrum. In some cases, the extracted spectral envelope features can include filterbank energies determined based on uncompanding (e.g., exp) the companded filterbank energies. In some examples, the extracted spectral envelope features can include a full resolution spectrum (e.g., DFT domain with companded (e.g., log, linear amplitude, etc.) information, etc.). The full resolution spectrum can be smoothed to the envelope of the spectrum, for example based on interpolating the companded or uncompanded filterbank energies. In some aspects, the extracted spectral envelope features can include one or more linear prediction (LP) coefficients, for example determined based on processing one or more speech frames of the input audio 415 using the Levinson-Durbin algorithm (e.g., autocorrelation technique), and/or using a covariance technique. In some examples, the extracted spectral envelope features can include one or more of line spectral frequency (LSF) information and/or line spectral pair (LSP) information. Qualcomm Ref. No.2400252WO [0129] In some aspects, the encoded features or encoded feature information transmitted from the generative voice codec system encoder 405 can comprise a latent representation z between the encoder of the FRAE 425 and the decoder of the FRAE 425. The encoder 405 can be configured to pass (e.g., transmit) the encoded features z through a channel 440 to the decoder 410. [0130] The encoder 405 also processes the audio 415 using a pitch extraction engine 430 to extract pitch information from the audio 415. In some aspects, the audio 415 can be processed in parallel by the spectral envelope feature extraction engine 420 and the pitch extraction engine 430. In some cases, the encoder 405 can include a denoiser 418 that processes the input audio 415 and provides a de-noised audio to the spectral envelope feature extraction engine 420 and/or to the pitch extraction engine 430. [0131] Based on the audio 415, the pitch extraction engine 430 can generate one or more types of pitch information. For example, the pitch extraction engine 430 can output pitch information in the frequency-domain (e.g., f0 pitch information of the audio 415, in units of Hertz (Hz)) and/or can output pitch information in the time-domain (e.g., pitch lag in samples, with or without a fractional component, and/or pitch lag in milliseconds). In some cases, the pitch extraction engine 430 can be used to generate pitch estimation indicative of a pitch lag (or pitch delay) from the audio 415, and/or to identify a pitch correlation from the audio 415. In some aspects, the pitch information generated by the pitch extraction engine 430 can include a pitch lag and a pitch correlation. [0132] The encoder 405 can be configured to processes the pitch information of the audio 415 (e.g., pitch, pitch lag, and/or pitch correlation, determined using the pitch extraction engine 430) using a quantizer 435 Q() to generate a quantized pitch signal. The encoder 405 passes the quantized pitch signal through the channel 440 to the decoder 410. In some examples, the quantizer 435 Q() may also be referred to as a pitch quantizer. In some aspects, the quantizer 435 Q() can perform vector quantization (VQ) to generate the quantized pitch signal for transmission to the decoder 410 over the channel 440. In some examples, the quantizer 435 Q() can perform single- stage VQ and/or can perform multi-stage VQ (e.g., MSVQ). In some cases, the quantizer 435 Q() can be implemented using one or more FRAEs. In some aspects, the quantizer 435 Q() can be configured to implement forward error correction (FEC) for the quantized pitch signal that is transmitted to the decoder 410 over the channel 440. For example, the quantizer 435 Q() can Qualcomm Ref. No.2400252WO implement multiple description coding (MDC) for the transmission of the quantized pitch signal over the channel 440, can implement full-redundancy FEC for the transmission of the quantized pitch signal over the channel 440, etc. [0133] The generative voice codec system decoder 410 can be configured to receive the encoded features z from the FRAE 425 included in the generative voice codec system encoder 405. For example, the decoder 410 can receive the encoded features z via the channel 440. The decoder 410 decodes the encoded features z using a FRAE 445 to generate decoded features. In some examples, the FRAE 445 of the decoder 410 includes only a decoder, without an encoder. In some examples, the FRAE 445 of the decoder 410 can include an encoder. In some examples, the decoded features generated by the FRAE 445 include mel-frequency cepstral coefficients (MFCC), such as MFCC- 24 features (e.g., 24-dimensional MFCC features). [0134] The decoded features determined using the FRAE decoder 445 can be provided to a neural speech synthesizer 450, which can be configured to generate a reconstructed audio 455 (e.g., synthesized speech) based at least in part on the decoded spectral envelope features from the FRAE decoder 445. In one illustrative example, the neural speech synthesizer 450 receives as input the decoded spectral envelope features (e.g., from the FRAE decoder 445) and the received pitch encoding information (e.g., the quantized pitch signal received by the decoder 410 over the channel 440 and from the pitch quantizer 435 of the encoder 405). In some aspects, the neural speech synthesizer 450 can generate the reconstructed audio 455 (e.g., synthesized speech) based on processing the decoded spectral envelope features along with a pitch signal (e.g., the quantized pitch signal or a reconstructed variant thereof). [0135] For example, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and/or the pitch correlation of the audio 415) from the encoder 405 via the channel 440. In some examples, the decoder 410 passes the quantized pitch signal to the neural speech synthesizer 450, and the neural speech synthesizer 450 processes the decoded spectral envelope features along with the quantized pitch signal to generate reconstructed audio 455 (e.g., synthesized speech). [0136] In some cases, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and/or the pitch correlation of the audio 415) from the encoder 405 via the channel 440, and can perform dequantization of the quantized pitch signal to obtain a reconstructed Qualcomm Ref. No.2400252WO pitch signal. The reconstructed pitch signal can be passed to the neural speech synthesizer, and used in combination with the decoded spectral envelope features from the FRAE decoder 445 to generate the reconstructed audio 455 (e.g., synthesized speech). In some examples, the decoder 410 can include a reconstruction engine that is separate from the neural speech synthesizer 450. The reconstruction engine processes the quantized pitch signal to reconstruct the pitch signal (e.g., including the pitch, the pitch lag, and/or the pitch correlation) before the neural speech synthesizer 450 using the reconstructed pitch signal (e.g., the pitch, the pitch lag, and/or the pitch correlation) to generate the reconstructed audio 455. [0137] In some examples, the output of the neural speech synthesizer 450 can be provided to one or more linear predictive coding (LPC) layers 452 included in the decoder 410. The LPC 452 can be used to perform linear prediction analysis and/or linear prediction synthesis, based on the output of the neural speech synthesizer 450. For example, the LPC 452 can perform linear prediction based on the output of the neural speech synthesizer 450, to generate the reconstructed audio 455 (e.g., synthesized speech). In some aspects, the decoder 410 does not include the LPC 452, and the neural speech synthesizer 450 can be trained and/or configured to generate the reconstructed audio 455 directly (e.g., the output of the neural speech synthesizer 450 can be the reconstructed audio 455). [0138] In some aspects, the LPC layers 452 can be used to implement a linear prediction (LP) synthesis filter, based on: ^ ^^^^^^ ^^^ ^ ^^^^^^ [0139] Here, ^^^^^^ represents (e.g., the input signal to the LPC layers 452), ^^^^^^ represents the output signal (e.g., the output of the LPC layers 452, which can be the reconstructed audio 455), ^^^ represents the linear prediction coefficients associated with implementing the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values. [0140] In some aspects, the LP synthesis filter associated with the LPC layers 452 can be implemented in the time domain or in the frequency domain. For example, in the time domain, the Qualcomm Ref. No.2400252WO LP synthesis filter can be implemented based on a difference equation. In some cases, the LP synthesis filter can be implemented in the time domain by convolving the input signal x[n] with the LP filter impulse response or an approximation of the LP filter impulse response. For example, the LP filter can be associated with an infinite impulse response (IIR), which can be approximated using a finite segment of the IIR (e.g., such as the first N samples of the IIR for a configured integer value of N, etc.). The LP synthesis filter associated with the LPC layers 452 can be implemented in the frequency domain based on multiplying the FFT of the input signal x[n] with the frequency response of the LP filter, and then determining the IFFT of the result to convert the output of the frequency domain LP synthesis filter into a time domain signal corresponding to the reconstructed audio 455. [0141] In some examples, the LP coefficients ^^^ can be estimated from the decoded spectral envelope features obtained using the FRAE decoder 445. In some cases, the spectral envelope features can be LP coefficients (e.g., determined by the spectral envelope feature extraction engine 420 based on using the Levinson-Durbin algorithm or other autocorrelation technique, or using a covariance technique, to process the input speech frames of the audio 415), and the decoded LP coefficients on the decoder side can be used for the LP filter implemented by the LPC layers 452. [0142] In some examples, the extracted spectral envelope features may be LSFs or LSPs, and the decoded LSFs or LSPs obtained using the FRAE decoder 445 can be converted to the corresponding LP coefficients ^^^. In some cases, the LP coefficients ^^^ of the LPC layers 452 can be estimated from the spectral envelope features using a neural network-based and/or a DSP-based technique. For example, the extracted spectral envelope features may comprise a cepstrum, and a DSP-based technique can be used to convert the decoded cepstrum into filterbank energies, and subsequently interpolate the filterbank energies to estimate an FFT square magnitude at all FFT bins (e.g., including bins beyond the centers of filterbank filters). After estimating the FFT square magnitude based on the interpolated filterbank energies, the DSP-based technique can include applying an IFFT to obtain an estimate of the autocorrelation sequence, and utilizing the Levinson- Durbin algorithm on the estimated autocorrelation sequence to obtain the LP coefficients ^^^. [0143] In some examples, the spectral envelope feature extraction 420 can be performed based at least in part on spectral feature learning. For example, spectral feature learning can be performed by one or more machine learning models configured and/or trained to generate as output the one Qualcomm Ref. No.2400252WO or more extracted spectral envelope features. In some cases, a cepstrum or cepstrum information may be an optional input to a machine learning spectral feature learning model used to implement the spectral envelope feature extraction 420. In some cases, the spectral envelope feature extraction 420 can be implemented using one or more machine learning spectral feature learning models or engines, which can be jointly trained with the neural speech synthesizer 450, an NHV used to implement the neural speech synthesizer (e.g., such as the NHV-based neural speech synthesizer 650 of FIG.6), an LPC network used to implement the neural speech synthesizer (e.g., such as the LPC network-based neural speech synthesizer 750 of FIG.7), etc. [0144] In some cases, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on a short-time Fourier transform (STFT) loss and/or STFT loss function. In some aspects, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on an adversarial and/or generative adversarial network (GAN) loss. For example, joint training can be performed using a multi-resolution STFT loss, based on calculating STFT amplitude spectrograms from ground truth speech information and the synthesized speech information 455. The multi-resolution STFT loss can then be determined as a sum of mean absolute error and mean absolute error in the log domain. The STFT amplitude spectrograms can be calculated at different window lengths to capture representations of the error and/or loss at different time and/or frequency resolutions. In some aspects, joint training based on adversarial or GAN loss can be used to learn temporal fine structures in speech signals. An adversarial or GAN loss can be used to match the distribution of real speech and the synthesized speech 455. For example, the neural speech synthesizer 450 (and/or an NHV-based neural speech synthesizer) can be used as the generator network. A separate discriminator network can be used to train based on the adversarial loss, with both the NHV and the discriminator jointly trained using adversarial training techniques. In some examples, the discriminator network can be implemented as a classifier configured to classify whether an input audio represents real speech or synthesized speech. The generator network can attempt to fool the discriminator network to classify synthesized speech as real speech. Over the course of training, the generator network improves and begins producing (e.g., generating or synthesizing) speech that is very similar to the real speech. In some cases, the discriminator network can be implemented using WaveNet. Qualcomm Ref. No.2400252WO [0145] The neural speech synthesizer 450 can be implemented using one or more trained neural networks. For example, the one or more trained neural networks can be trained to perform speech synthesis and/or signal synthesis. In some examples, the neural speech synthesizer 450 can be implemented using or based on an LPCNet machine learning architecture, a WaveNet machine learning architecture, a WaveRNN machine learning architecture, etc. In one illustrative example, the neural speech synthesizer 450 can be provided as a neural homomorphic vocoder (NHV). [0146] FIG. 16 illustrates an example of an audio codec system 1600 (e.g., generative voice codec) that includes a decoder 1610 configured to implement an NHV-based neural speech synthesizer with a technique for random background noise synthesis, in accordance with aspects of the present disclosure. In some examples, an NHV is a type of neural vocoder that can synthesize speech with source-filter models controlled by one or more neural networks. An NHV-based neural vocoder may include one or more neural networks in a source-filter model that can synthesize speech based on filtering impulse trains and noise with linear time-varying (LTV) filters, with the one or more neural networks used to control the LTV filters by estimating complex cepstrums of time-varying impulse responses given acoustic features. Traditional or non-neural vocoders may operate based on decomposing speech into various parameters such as pitch, timbre, rhythm, etc., which can subsequently be manipulated and resynthesized to generate a desired output audio or voice signal. Neural vocoders (e.g., such as NHVs, etc.) can apply transformations and manipulations to speech signals directly within a learned feature space of the neural network. The learned feature space used by neural vocoders may capture more complex relationships and characteristics of speech than non-neural network-based signal processing and/or vocoder techniques. For example, neural vocoders can be trained to learn a mapping between raw speech waveforms and the spectral or cepstral representations of the speech. An NHV system can apply one or more homomorphic processing techniques within the learned space corresponding to the mapping. [0147] In some aspects, the decoder 1610 may also be referred to as an NHV synthesizer decoder. In some examples, the audio codec system 1600 of FIG.16 may be similar to the audio codec system 1500 of FIG.15. In some cases, the channel 1640 of FIG.16 can be the same as or similar to the channel 1540 of FIG.15, and may be associated with an encoder that is the same as or similar to the encoder 1505 of FIG.15, etc. Qualcomm Ref. No.2400252WO [0148] In some aspects, the decoder 1610 can include a neural speech synthesizer 1650 that is the same as or similar to the neural speech synthesizer 1550 included in the decoder 1510 of FIG. 15. For example, the neural speech synthesizer 1650 can receive a first input comprising decoded spectral envelope features (e.g., determined by a FRAE decoder 1645 that is the same as or similar to the FRAE decoder 1545 of FIG.15). The neural speech synthesizer 1650 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 1550 of FIG. 15. [0149] In one illustrative example, the NHV synthesizer decoder 1610 can include the FRAE decoder 1645, which may be the same as or similar to the FRAE decoder 1545 of FIG. 15, and may include a pitch dequantization engine 1638 (e.g., de Q()). The pitch dequantization engine 1638 can be used to process a received pitch encoding obtained by the decoder 1610 over the channel 1640. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 1610 over the channel 1640 by a corresponding encoder (e.g., an encoder associated with the decoder 1610, which may be the same as or similar to the encoder 1505 of FIG.15, etc.). In some aspects, the pitch dequantization engine 1638 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 1640 (e.g., to dequantize the quantized pitch encoding received by the decoder 1610 over the channel 1640). [0150] The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1638 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and/or pitch correlation information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1638 can include information indicative of a voiced or unvoiced (V/UV) classification. [0151] In one illustrative example, the neural speech synthesizer 1650 can be implemented as a neural homomorphic vocoder (NHV) speech synthesizer and/or an NHV-based neural speech synthesizer. For example, the neural speech synthesizer 1650 can include a neural filter estimator 1652 configured to generate respective filter specification or filter configuration information to parameterize one or more linear time-varying (LTV) filters of the NHV speech synthesizer 1650. Qualcomm Ref. No.2400252WO In some aspects, the neural filter estimator 1652 can be trained to generate respective filter coefficients for a noise linear time-varying filter (LTVF) 1675 and to generate respective filter coefficients for a harmonic LTVF 1665. [0152] The neural network model of the neural filter estimator 1652 can include any neural network architecture that can be trained to model the filter coefficients for the noise LTVF 1675 and the harmonic LTVF 1665 (e.g., and/or that can be trained to model the filter coefficients for one or more additional filters implemented by the NHV speech synthesizer 1650, in either the time-domain, the frequency-domain, or combinations thereof). Examples of neural network architectures that may be included in the neural network filter estimator 1652 can include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and/or other type(s) of neural networks, etc. [0153] In some examples, the neural filter estimator 1652 can be configured to receive a stream of decoded spectral envelope features from the FRAE decoder 1645. Based on the decoded spectral envelope features, the neural filter estimator 1652 can generate corresponding filter characterization parameters for each filter of one or more filters included in the NHV speech synthesizer 1650. The respective filter characterization parameters can be output from the neural filter estimator 1652 and used to parameterize and/or configure corresponding learned filters (e.g., learned linear filters, etc.) for each respective set of filter characterization parameters. For example, the neural filter estimator 1652 can generate a first set of filters corresponding to the noise LTVF 1675, can generate a second set of filter characterization parameters corresponding to the harmonic LTVF 1665, etc. The one or more filters included in the NHV speech synthesizer 1650 (e.g., the noise LTVF 1675, the harmonic LTVF 1665, etc.) can be specified in various different forms. For example, the noise LTVF 1675, the harmonic LTVF 1665, and/or various other filters that may be included in the NHV speech synthesizer 1650 can be specified or characterized (e.g., using the filter characterization parameters determined by the neural filter estimator 1652) as cepstrums, as frequency responses, as time-domain impulse responses, as difference equation coefficients, etc. [0154] In some aspects, the neural filter estimator 1652 can be configured to receive as input the decoded spectral envelope features from the FRAE decoder 1645, and may additionally receive as input at least a portion of the dequantized pitch information generated by the pitch dequantization engine 1638. For example, the neural filter estimator 1652 may receive as input the decoded Qualcomm Ref. No.2400252WO spectral envelope features and pitch dequantization information (e.g., such as f0 pitch information in the frequency domain, pitch lag or pitch delay in the time domain (e.g., in units of samples or milliseconds, etc.), etc.). In some examples, the pitch dequantization information provided as input to the neural filter estimator 1652 may include voiced/unvoiced (V/UV) classification information indicating whether the underlying audio represented in the encoded information received by the decoder 1610 over the channel 1640 corresponds to voiced or unvoiced sounds, speech, etc. [0155] In examples where the neural filter estimator 1652 receives spectral envelope features and dequantized pitch information as inputs, the neural filter estimator 1652 can generate the corresponding filter characterization parameters for each NHV filter (e.g., noise LTVF 1675, harmonic LTVF 1665, etc.) based on the decoded spectral envelope features and the dequantized pitch information. [0156] The noise LTVF 1675 and the harmonic LTVF 1665 can be implemented as time domain filters or frequency domain filters. In some examples, the noise LTVF 1675 and the harmonic LTVF 1665 can be implemented in the time domain or the frequency domain, independent of whether the neural filter estimator 1652 is configured to generate the corresponding filter characterization parameters in the time domain or the frequency domain. For example, in some aspects, the neural filter estimator 1652 can output impulse response-based filter characterization parameters, and the operation of noise LTVF 1675 and/or harmonic LTVF 1665 can be implemented in the time domain as convolutions with the impulse response. Using the same impulse response-based filter characterization parameters, the operation of noise LTVF 1675 and/or harmonic LTVF 1665 may be implemented in the frequency domain based on converting the impulse response to frequency response (e.g., using an FFT transform) and multiplying with the FFT of the input signal, and subsequently converting back to the time domain using an IFFT transform. [0157] In one illustrative example, the NHV speech synthesizer 1650 can include a pulse train generator 1660. The dequantized pitch information (e.g., generated using the pitch dequantization engine 1638) can be provided to the pulse train generator 1660. In some aspects, the pulse train generator 1660 can generate a pulse train based at least in part on the pitch frequency (e.g., f0) and/or pitch lag or pitch delay information obtained from the pitch dequantization engine 1638 of the decoder 1610. In some examples, the pulse train generator 1660 can be implemented as a cosine Qualcomm Ref. No.2400252WO sum pulse generator, which can be configured to process the dequantized pitch information (e.g., obtained from the pitch dequantization engine 1638) to generate a pulse train p[n]. The pulse train may also be referred to as an impulse train. For example, the pulse train generator 1660 can be a differentiable cosine sum pulse generator, and/or can be a non-differentiable cosine sum pulse generator (e.g., among various other pulse generators). [0158] The pulse train generated by the pulse train generator 1660 can be processed by the harmonic LTVF 1665, using the corresponding harmonic filter characterization parameters determined by the neural filter estimator 1652 for the harmonic LTVF 1665. In some aspects, the harmonic LTVF 1665 can generate a harmonic output based on processing the pulse train from the pulse train generator 1660. [0159] In some examples, the pulse train generator 1660 may receive an additional input from the pitch dequantization engine 1638, indicative of a voice or unvoiced (e.g., V/UV) classification. For example, based on receiving an unvoiced (UV) indication or classification from the pitch dequantization engine 1638, the pulse train generator 1660 can be configured to generate a 0 output or a noise output that is provided to the harmonic LTVF 1665 instead of the pulse train p[n]ep[n] provided from the pulse train generator 1660 to the harmonic LTVF 1665 in response to a voiced (V) indication or classification from the pitch dequantization engine 1638). [0160] The NHV speech synthesizer 1650 can include a noise generator 1670 that is configured to generate a noise signal (e.g., white noise, etc.) for processing by the noise LTVF 1675. For example, the noise LTVF 1675 can be parameterized based on the respective noise filter characterization parameters generated by the neural filter estimator 1652 for the noise LTVF 1675, and can subsequently be used to process the noise signal generated by the noise generator 1670. Based on processing the noise signal from the noise generator 1670, the noise LTVF 1675 can generate a noise-filtered output. [0161] In some cases, a hangover detector 1681 may detect that an audio frame is from the hangover region. For example, the hangover detector 1681 may be a portion of a decoder (e.g., decoder 701 of FIG.7) or classifier (e.g., classifier 703 of FIG.7) that may detect (and/or infer) an indication that an audio frame is from a hangover region. Based on the detection that the audio frame is from the hangover region, the hangover detector 1681 may indicate to a filter coefficients copier 1682 to copy the coefficients of the noise LTVF 1675 to noise filter 1684. In some cases, Qualcomm Ref. No.2400252WO noise filter 1684 may operate as an inactive synthesizer (e.g., inactive synthesizer 712 of FIG.7, inactive synthesizer 1104 of FIG. 11) to use the random noise from the random noise generator 1686 as the synthesized background noise. In some cases, the random noise generator 1686 may correspond to random noise generator 1102 of FIG. 11. In some cases, a hangover region may include inactive frames following speech and thus filter coefficients estimated in the hangover region may be used for noise generation and the noise filter 1684 may mix the random noise from the random noise generator 1686 with synthesized background noise based on the copied filter coefficients. In some cases, the random noise generator 1686 may be the same as noise generator 1670, or random noise generator 1686 may be a separate noise generator from noise generator 1670. [0162] The NHV speech synthesizer 1650 can include a combination function 1680 to combine the harmonic output (e.g., from the harmonic LTVF 1665) and the noise-filtered output (e.g., from the noise LTVF 1675) to generate a predicted sample (S't) that may be input to a LPC 1654 during active speech. During a hangover or inactive region, the combination function 1680 may combine the harmonic output (e.g., from the harmonic LTVF 1665) and the synthesized noise output from the noise filter 1684 to generate a predicted sample (S't) that may be input to a LPC 1654. In some examples, the combination function 1680 can include an adder, a multiplier, a divider, weighted sum, a weighted product, a weighted ratio, an average, a weighted average, a weighted mean, a weighted median, a weighted mode, or a combination thereof. The reconstructed audio 1655 may be an example of the reconstructed speech signal 105 of FIG.1, the reconstructed speech signal 205 of FIG.2, etc. [0163] FIG.17 is a block diagram illustrating an example of a decoder 1710 of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear prediction coding (LPC) network with a technique for random background noise synthesis, in accordance with aspects of the present disclosure, in accordance with aspects of the present disclosure. The decoder 1710 can be included in an audio codec system (e.g., generative voice codec system) 1700 that can be used to generate reconstructed audio (e.g., synthesized speech) 1755. [0164] In some aspects, the generative voice codec system 1700 of FIG.17 can be the same as or similar to the generative voice codec system 1500 of FIG. 15, 1600 of FIG. 16, etc. In some Qualcomm Ref. No.2400252WO cases, the channel 1740 of FIG.17 can be the same as or similar to the channel 1540 of FIG.15, the channel 1640 of FIG.16, etc. [0165] The decoder 1710 of FIG.17 can be the same as or similar to the decoder 1510 of FIG. 15, the decoder 1610 of FIG.16, etc. For example, the decoder 1710 can include an FRAE decoder 1745 the same as or similar to the FRAE decoder 1545 of FIG. 15, 1645 of FIG. 16, etc. The decoder 1710 can include a neural speech synthesizer 1750 that can be the same as or similar to the neural speech synthesizer 1550 of FIG.15, 1650 of FIG.16, etc. The decoder 1710 can generate reconstructed audio 1755 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 1555 of FIG.15, 1655 of FIG.16, etc. [0166] In some examples, the neural speech synthesizer 1750 can receive a first input comprising decoded spectral envelope features (e.g., determined by the FRAE decoder 1745). The neural speech synthesizer 1750 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 1550 of FIG.15, etc. For example, the decoder 1710 may include a pitch de-quantization engine 1738 (e.g., de Q()). The pitch de-quantization engine 1738 can be used to process a received pitch encoding obtained by the decoder 1710 over the channel 1740. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 1710 over the channel 1740 by a corresponding encoder (e.g., an encoder associated with the decoder 1710, which may be the same as or similar to the encoder 1505 of FIG.15, etc.). In some aspects, the pitch dequantization engine 1738 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 1740 (e.g., to dequantize the quantized pitch encoding received by the decoder 1710 over the channel 1740). The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1738 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and/or pitch correlation information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1738 can include information indicative of a voiced or unvoiced (V/UV) classification. Qualcomm Ref. No.2400252WO [0167] In one illustrative example, the neural speech synthesizer 1750 can be implemented as a linear predictive coding (LPC) network. For example, the LPC network-based neural speech synthesizer 1750 can be used to implement a linear prediction (LP) synthesis filter, based on ^^^^^^ ൌ ∑ ^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^. Here, ^^^^^^ represents the input signal to the LP filter, ^^^^^^ represents the output signal, ^^^ represents the linear prediction coefficients associated with implementing the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values. [0168] For example, the LPC-based neural speech synthesizer 1750 can include a frame rate network 1752 configured to process inputs comprising the decoded features obtained using the FRAE decoder 1745 and the decoded pitch information obtained using the pitch dequantization engine 1738. An LPC estimation engine 1762 can process the decoded features from the FRAE decoder 1745 to determine one or more estimated LP coefficients. In some aspects, the LPC estimation engine 1762 can generate estimated LP coefficients ^^^, based on an input comprising the decoded features determined using the FRAE decoder 1745. [0169] In some examples, the LP coefficients estimated using the LPC estimation engine 1762 can be provided to an LP prediction engine 1764, configured to generate as output a prediction p[n], where ^^^^^^ ൌ ∑ ^ ^ୀ^ ^^^^^^^^ െ ^^^ . In some aspects, the LP prediction engine 1764 can generate the LP prediction p[n] based on a first input comprising the LP coefficients ^^^ estimated by the LPC estimation engine 1762, and a feedback input s[n-1]. [0170] The estimated or predicted LP coefficients p[n] can be provided as input to a sample rate network 1772 and a downstream combination or summation operation 1780. The sample rate network 1772 can receive additional inputs comprising the output of the frame rate network 1752, the feedback s[n-1] from the previous step n-1, and an intermediate feedback value e[n-1] from the same previous step n-1. [0171] The output of the sample rate network 1772 can be the probability distribution P(e[n]), which is a probability distribution for e[n]. A sampling engine 1774 can perform sampling from the probability distribution P(e[n]) to obtain a realization or representation of e[n]. Qualcomm Ref. No.2400252WO [0172] The representation of e[n] determined by the sampling engine 1774 can be combined with the LP prediction p[n] (e.g., determined by the LP prediction engine 1764), using the combination or summation operation 1780 to thereby generate as output the signal s[n]. In some aspects, the output signal s[n] can be the same as the reconstructed audio 1755 of the LPC network- based neural speech synthesizer 1750 and/or decoder 1710. [0173] The representation of e[n] determined by the sampling engine 1774 can additionally be provided to a first feedback calculation 1778, which generates as output e[n-1] provided as an additional input to the sample rate network 1772. [0174] The output signal s[n] of the combination or summation operation 1780 can be output as the reconstructed audio 1755 and may additionally be provided to a second feedback calculation 1779, which generates the representation s[n-1] based on the input s[n]. The representation s[n-1] can be provided as a feedback input to the LP prediction engine 1764 and to the sample rate network 1772. [0175] In some cases, a hangover detector 1781 may detect that an audio frame is from the hangover region. For example, the hangover detector 1781 may be a portion of a decoder that may detect (and/or infer) an indication that an audio frame is from a hangover region in a manner similar to that discussed above with respect to decoder 701 of FIG.7. Based on the detection that the audio frame is from the hangover region, the hangover detector 1781 may indicate to a noise filter 1784 that the frame is from the hangover region. In some cases, noise filter 1784 may operate as an inactive synthesizer (e.g., inactive synthesizer 712 of FIG.7, inactive synthesizer 1104 of FIG.11) to use the random noise from the random noise generator 1786 and output the random noise to the combination or summation operation 1780 to mix in as the synthesized background noise. In some cases, the random noise generator 1786 may correspond to random noise generator 1102 of FIG. 11. In some cases, a hangover region may include inactive frames following speech and thus filter coefficients estimated in the hangover region may be used for noise generation. [0176] FIG.18 is a flow diagram illustrating an example of a process 1800 for audio playback, in accordance with aspects of the present disclosure. The process 1800 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, processor 484 of FIG.4, DSP 482 of FIG.4, processor 1910 of FIG.19, etc.) of the computing device (e.g., UE 104 or UE 190 of FIGs.1-3, wireless device 407 of FIG.7, receiving UE 1406 of FIG.14, computing system Qualcomm Ref. No.2400252WO 1900 of FIG. 19, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device. In some cases, the computing device may be or may include UE device, such as the UE 104 or UE 190 of FIGs.1-3. The operations of the process 1800 may be implemented as software components that are executed and run on one or more processors. [0177] At block 1802, the computing device (or component thereof) may receive a first decoded audio frame. In some cases, the first decoded audio frame includes active speech. For example, decoder 701 of FIG. 7 may decode an audio frame of the encoded audio data stream to obtain decoded active speech. In some cases, the computing device (or component thereof) may receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device. [0178] At block 1804, the computing device (or component thereof) may generate information associated with background noise (e.g., information about the background noise 710 of FIG.7) in the first decoded audio frame. In some cases, the computing device (or component thereof) may generate the information associated with the background noise in the first decoded audio frame by: denoising (e.g., by denoiser 704 of FIG 7) the first decoded audio frame to generate a denoised audio frame; subtracting (e.g., subtracted 706 of FIG.7) the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generating the information associated with the background noise based on the background noise frame. In some examples, the information about the background noise in the first decoded audio frame is generated by a machine learning model. In some cases, the information associated with background noise is determined based on at least one of a noise dictionary (e.g., noise dictionary 960 of FIG.9B) or a speech dictionary (e.g., speech dictionary 962 of FIG. 9B). In some examples, the information associated with background noise is generated by a neural synthesizer (e.g., neural speech synthesizer 1550 of FIG.15, neural speech synthesizer 1650 of FIG.16, etc.) of a decoder (e.g., decoder 1510 of FIG.15, decoder 1610 of FIG.16, etc.). In some cases, the information associated with the background noise is received from a noise linear time-varying filter (e.g., noise LTVF 1675 of FIG.16) of the neural speech synthesizer. In some examples, the synthesized background noise is combined with an output of a harmonic linear time-varying filter (e.g., harmonic LTVF Qualcomm Ref. No.2400252WO 1665 of FIG.16) of the neural speech synthesizer to generate a predicted sample. In some cases, the predicted sample is input to a linear predictive coding (LPC) network (e.g., LPC 1552 of FIG. 15, LPC 1654 of FIG.16, etc.) to generate reconstructed audio (e.g., reconstructed audio 1555 of FIG.15, reconstructed audio 1655 of FIG.16. [0179] At block 1806, the computing device (or component thereof) may detect a transition to an inactive region after the first decoded audio frame. For example, the encoder may indicate a transition to the inactive region, and the decoder 701 of FIG.7 and/or classifier 703 of FIG.7 may determine that there is a transition to the inactive region. In some cases, the computing device (or component thereof) may to detect the transition to the inactive region after the first decoded audio frame by receiving a radio access network message (e.g., message 1420 of FIG.14) indicating the transition to the inactive region. [0180] At block 1808, the computing device (or component thereof) may synthesize background noise (e.g., synthesized background noise 714 of FIG.7) based on the information associated with the background noise in response to the detected transition to the inactive region. In some cases, the computing device (or component thereof) may receive a third decoded audio frame (e.g., third input frame 870 of FIG.8B), the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. In some examples, the background noise is synthesized by a machine learning model (e.g., GRU 802, 810, 818 of FIG.8A, inactive synthesizer 850 of FIG.8B, ambient noise prediction engine 1004 of FIG.10, ML model 1304 of FIG.13, etc.). In some cases, the computing device (or component thereof) may receive an encoded audio frame from a second device; receive location information (e.g., location information 1002 of FIG.10) for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information. For example, the ambient noise prediction engine 1004 of FIG. 10 may predict a type of ambient noise (e.g., background noise) 1006 FIG. 10 that may be present based on the location information 1002 FIG. 10. In some examples, the background noise is synthesized based on random noise. For example, a random Qualcomm Ref. No.2400252WO noise generator 1102 of FIG.11 may be used to generate random noise based on, for example, a random number generator. [0181] In some examples, the techniques or processes described herein may be performed by a computing device, an apparatus, and/or any other computing device. In some cases, the computing device or apparatus may include a processor, microprocessor, microcomputer, or other component of a device that is configured to carry out the steps of processes described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. For example, the computing device may include a camera device, which may or may not include a video codec. As another example, the computing device may include a mobile device with a camera (e.g., a camera device such as a digital camera, an IP camera or the like, a mobile phone or tablet including a camera, or other type of device with a camera). In some cases, the computing device may include a display for displaying images. In some examples, a camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives the captured video data. The computing device may further include a network interface, transceiver, and/or transmitter configured to communicate the video data. The network interface, transceiver, and/or transmitter may be configured to communicate Internet Protocol (IP) based data or other network data. [0182] The processes described herein can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer- executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes. [0183] In some cases, the devices or apparatuses configured to perform the operations of the process 1800 and/or other processes described herein may include a processor, microprocessor, micro-computer, or other component of a device that is configured to carry out the steps of the process 1800 and/or other process. In some examples, such devices or apparatuses may include one or more sensors configured to capture image data and/or other sensor measurements. In some Qualcomm Ref. No.2400252WO examples, such computing device or apparatus may include one or more sensors and/or a camera configured to capture one or more images or videos. In some cases, such device or apparatus may include a display for displaying images. In some examples, the one or more sensors and/or camera are separate from the device or apparatus, in which case the device or apparatus receives the sensed data. Such device or apparatus may further include a network interface configured to communicate data. [0184] The components of the device or apparatus configured to carry out one or more operations of the process 1800 and/or other processes described herein can be implemented in circuitry. For example, the components can include and/or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and/or other suitable electronic circuits), and/or can include and/or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The computing device may further include a display (as an example of the output device or in addition to the output device), a network interface configured to communicate and/or receive the data, any combination thereof, and/or other component(s). The network interface may be configured to communicate and/or receive Internet Protocol (IP) based data or other type of data. [0185] The process 1800 is illustrated as a logical flow diagram, the operations of which represent sequences of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer- executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer- executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes. [0186] Additionally, the processes described herein (e.g., process 1800 and/or other processes) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer Qualcomm Ref. No.2400252WO programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer- readable or machine-readable storage medium, for example, in the form of a computer program including a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory. [0187] Additionally, the processes described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non- transitory. [0188] FIG.19 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. In particular, FIG.19 illustrates an example of computing system 1900, which may be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 1905. Connection 1905 may be a physical connection using a bus, or a direct connection into processor 1910, such as in a chipset architecture. Connection 1905 may also be a virtual connection, networked connection, or logical connection. [0189] In some aspects, computing system 1900 is a distributed system in which the functions described in this disclosure may be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components may be physical or virtual devices. [0190] Example system 1900 includes at least one processing unit (CPU or processor) 1910 and connection 1905 that communicatively couples various system components including system memory 1915, such as read-only memory (ROM) 1920 and random access memory (RAM) 1925 to processor 1910. Computing system 1900 may include a cache 1912 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1910. Qualcomm Ref. No.2400252WO [0191] Processor 1910 may include any general purpose processor and a hardware service or software service, such as services 1932, 1934, and 1936 stored in storage device 1930, configured to control processor 1910 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1910 may essentially be a completely self- contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric. [0192] To enable user interaction, computing system 1900 includes an input device 1945, which may represent any number of input mechanisms, such as a microphone for speech, a touch- sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1900 may also include output device 1935, which may be one or more of a number of output mechanisms. In some instances, multimodal systems may enable a user to provide multiple types of input/output to communicate with computing system 1900. [0193] Computing system 1900 may include communications interface 1940, which may generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and/or transmission wired or wireless communications using wired and/or wireless transceivers, including those making use of an audio jack/plug, a microphone jack/plug, a universal serial bus (USB) port/plug, an AppleTM LightningTM port/plug, an Ethernet port/plug, a fiber optic port/plug, a proprietary wired port/plug, 3G, 4G, 5G and/or other cellular data network wireless signal transfer, a BluetoothTM wireless signal transfer, a BluetoothTM low energy (BLE) wireless signal transfer, an IBEACONTM wireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof. The communications interface 1940 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to Qualcomm Ref. No.2400252WO determine a location of the computing system 1900 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed. [0194] Storage device 1930 may be a non-volatile and/or non-transitory and/or computer- readable memory device and may be a hard disk or other types of computer readable media which may store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip/stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini/micro/nano/pico SIM card, another integrated circuit (IC) chip/card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache), resistive random-access memory (RRAM/ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and/or a combination thereof. [0195] The storage device 1930 may include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1910, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function may include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1910, connection 1905, output device 1935, Qualcomm Ref. No.2400252WO etc., to carry out the function. The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and/or data. A computer-readable medium may include a non-transitory medium in which data may be stored and that does not include carrier waves and/or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and/or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and/or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like. [0196] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described. [0197] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional Qualcomm Ref. No.2400252WO components may be used other than those shown in the figures and/or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects. [0198] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure. [0199] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re- arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function. [0200] Processes and methods according to the above-described examples may be implemented using computer-executable instructions that are stored or otherwise available from computer- readable media. Such instructions may include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used may be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of Qualcomm Ref. No.2400252WO computer-readable media that may be used to store instructions, information used, and/or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on. [0201] In some aspects the computer-readable storage devices, mediums, and memories may include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se. [0202] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc. [0203] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also may be embodied in peripherals or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example. [0204] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure. Qualcomm Ref. No.2400252WO [0205] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and/or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non- volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that may be accessed, read, and/or executed by a computer, such as propagated signals or waves. [0206] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein. Qualcomm Ref. No.2400252WO [0207] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein may be replaced with less than or equal to (“^”) and greater than or equal to (“^”) symbols, respectively, without departing from the scope of this description. [0208] Where components are described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof. [0209] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and/or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and/or other suitable communication interface) either directly or indirectly. [0210] Claim language or other language reciting “at least one of” a set and/or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and/or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein. [0211] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of Qualcomm Ref. No.2400252WO operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z. [0212] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and/or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions. [0213] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and/or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and/or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function). [0214] Illustrative aspects of the disclosure include: [0215] Aspect 1. A device for audio playback, comprising: one or more memories; and one or more processors coupled to the one or more memories and configured to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information Qualcomm Ref. No.2400252WO associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0216] Aspect 2. The device of Aspect 1, wherein, to generate the information associated with the background noise in the first decoded audio frame, the one or more processors are configured to: denoise the first decoded audio frame to generate a denoised audio frame; subtract the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generate the information associated with the background noise based on the background noise frame. [0217] Aspect 3. The device of any of Aspects 1-2, wherein the one or more processors are configured to: receive a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. [0218] Aspect 4. The device of any of Aspects 1-3, wherein the background noise is synthesized by a machine learning model. [0219] Aspect 5. The device of any of Aspects 1-4, wherein the one or more processors are configured to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information. [0220] Aspect 6. The device of any of Aspects 1-5, wherein the background noise is synthesized based on random noise. [0221] Aspect 7. The device of any of Aspects 1-6, wherein the one or more processors are configured to receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device. Qualcomm Ref. No.2400252WO [0222] Aspect 8. The device of any of Aspects 1-7, wherein the background noise is synthesized based on random noise. [0223] Aspect 9. The device of any of Aspects 1-8, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model. [0224] Aspect 10. The device of any of Aspects 1-9, wherein, to detect the transition to the inactive region after the first decoded audio frame, the one or more processors are configured to receive a radio access network message indicating the transition to the inactive region. [0225] Aspect 11. The device of any of Aspects 1-10, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary. [0226] Aspect 12. A method for audio playback comprising: receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0227] Aspect 13. The method of Aspect 12, wherein generating the information associated with the background noise in the first decoded audio frame comprises: denoising the first decoded audio frame to generate a denoised audio frame; subtracting the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generating the information associated with the background noise based on the background noise frame. [0228] Aspect 14. The method of any of Aspects 12-13, comprising: receiving a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generating information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. [0229] Aspect 15. The method of any of Aspects 12-14, wherein the background noise is synthesized by a machine learning model. [0230] Aspect 16. The method of any of Aspects 12-15, comprising: receiving an encoded audio frame from a second device; receiving location information for the second device; and decoding Qualcomm Ref. No.2400252WO the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information. [0231] Aspect 17. The method of any of Aspects 12-16, wherein the background noise is synthesized based on random noise. [0232] Aspect 18. The method of any of Aspects 12-17, comprising receiving, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device. [0233] Aspect 19. The method of any of Aspects 12-18, wherein the background noise is synthesized based on random noise. [0234] Aspect 20. The method of any of Aspects 12-19, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model. [0235] Aspect 21. The method of any of Aspects 12-20, wherein detecting the transition to the inactive region after the first decoded audio frame comprises receiving a radio access network message indicating the transition to the inactive region. [0236] Aspect 22. The method of any of Aspects 12-21, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary. [0237] Aspect 23. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. [0238] Aspect 24. The non-transitory computer-readable medium of Aspect 23, wherein generating the information associated with the background noise in the first decoded audio frame, the instructions cause the one or more processors to: denoise the first decoded audio frame to generate a denoised audio frame; subtract the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generate the information associated with the background noise based on the background noise frame. Qualcomm Ref. No.2400252WO [0239] Aspect 25. The non-transitory computer-readable medium of any of Aspects 23-24, wherein instructions case the one or more processors to: receive a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. [0240] Aspect 26. The non-transitory computer-readable medium of any of Aspects 23-25, wherein the background noise is synthesized by a machine learning model. [0241] Aspect 27. The non-transitory computer-readable medium of any of Aspects 23-26, wherein instructions case the one or more processors to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information. [0242] Aspect 28. The non-transitory computer-readable medium of any of Aspects 23-27, wherein the background noise is synthesized based on random noise. [0243] Aspect 29. The non-transitory computer-readable medium of any of Aspects 23-28, wherein the instructions case the one or more processors to receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device. [0244] Aspect 30. The non-transitory computer-readable medium of any of Aspects 23-29, wherein the background noise is synthesized based on random noise. [0245] Aspect 31. The non-transitory computer-readable medium of any of Aspects 23-30, wherein the information about the background noise in the first decoded audio frame is generated by a machine learning model. [0246] Aspect 32. The non-transitory computer-readable medium of any of Aspects 23-31, wherein, to detect the transition to the inactive region after the first decoded audio frame, the instructions case the one or more processors to receive a radio access network message indicating the transition to the inactive region. Qualcomm Ref. No.2400252WO [0247] Aspect 33. The non-transitory computer-readable medium of any of Aspects 23-32, wherein information associated with background noise is determined based on at least one of a noise dictionary or a speech dictionary. [0248] Aspect 34. The device of any of Aspects 1-11, wherein the information associated with background noise is generated by a neural synthesizer of a decoder. [0249] Aspect 35. The device of Aspect 35, wherein the information associated with the background noise is received from a noise linear time-varying filter of the neural speech synthesizer. [0250] Aspect 36. The device of Aspect 35, wherein the synthesized background noise is combined with an output of a harmonic linear time-varying filter of the neural speech synthesizer to generate a predicted sample. [0251] Aspect 37. The device of Aspect 36, wherein the predicted sample is input to a linear predictive coding (LPC) network to generate reconstructed audio. [0252] Aspect 38. The method of Aspect 12, further comprising performing the method according to any of Aspects 34-37. [0253] Aspect 39. The non-transitory computer-readable medium of Aspect 23, wherein the instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any of Aspects 34-37. [0254] Aspect 34. An apparatus comprising means for performing a method according to any of Aspects 12 to 22 and Aspects 34-37.

Claims

Qualcomm Ref. No.2400252WO CLAIMS WHAT IS CLAIMED IS: 1. A device for audio playback, comprising: one or more memories; and one or more processors coupled to the one or more memories and configured to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. 2. The device of claim 1, wherein, to generate the information associated with the background noise in the first decoded audio frame, the one or more processors are configured to: denoise the first decoded audio frame to generate a denoised audio frame; subtract the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generate the information associated with the background noise based on the background noise frame. 3. The device of claim 1, wherein the one or more processors are configured to: receive a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. 4. The device of claim 1, wherein the one or more processors are configured to synthesize the background noise using a machine learning model. Qualcomm Ref. No.2400252WO 5. The device of claim 1, wherein the one or more processors are configured to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the one or more processors are configured to synthesize the background noise based on the received location information. 6. The device of claim 1, wherein the one or more processors are configured to synthesize the background noise based on random noise. 7. The device of claim 1, wherein the one or more processors are configured to receive, from a wireless node, an indication that a second device will not send a silence descriptor (SID) to the device. 8. The device of claim 1, wherein the one or more processors are configured to generate the information associated with the background noise in the first decoded audio frame using a machine learning model. 9. The device of claim 1, wherein, to detect the transition to the inactive region after the first decoded audio frame, the one or more processors are configured to receive a radio access network message indicating the transition to the inactive region. 10. The device of claim 1, wherein information associated with the background noise is determined based on at least one of a noise dictionary or a speech dictionary. 11. The device of claim 1, wherein the one or more processors are configured to generate the information associated with the background noise using a neural speech synthesizer of a decoder. Qualcomm Ref. No.2400252WO 12. The device of claim 11, wherein the one or more processors are configured to receive the information associated with the background noise from a noise linear time-varying filter of the neural speech synthesizer. 13. The device of claim 11, wherein the one or more processors are configured to combine the synthesized background noise with an output of a harmonic linear time-varying filter of the neural speech synthesizer to generate a predicted sample. 14. The device of claim 13, wherein the predicted sample is input to a linear predictive coding (LPC) network to generate reconstructed audio. 15. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a first decoded audio frame, the first decoded audio frame including active speech; generate information associated with background noise in the first decoded audio frame; detect a transition to an inactive region after the first decoded audio frame; and synthesize background noise based on the information associated with the background noise in response to the detected transition to the inactive region. 16. The non-transitory computer-readable medium of claim 15, wherein, to generate the information associated with the background noise in the first decoded audio frame, the instructions, when executed by one or more processors, cause the one or more processors to: denoise the first decoded audio frame to generate a denoised audio frame; subtract the denoised audio frame from the first decoded audio frame to obtain a background noise frame; and generate the information associated with the background noise based on the background noise frame. 17. The non-transitory computer-readable medium of claim 15, wherein the instructions, when executed by one or more processors, cause the one or more processors to: Qualcomm Ref. No.2400252WO receive a third decoded audio frame, the third decoded audio frame from a hangover region following the first decoded audio frame; and generate information associated with the background noise in the third decoded audio frame, and wherein the synthesized background noise is based on the information associated with the background noise in the third decoded audio frame from the hangover region. 18. The non-transitory computer-readable medium of claim 15, wherein the background noise is synthesized by a machine learning model. 19. The non-transitory computer-readable medium of claim 15, wherein the instructions, when executed by one or more processors, cause the one or more processors to: receive an encoded audio frame from a second device; receive location information for the second device; and decode the encoded audio frame to generate the first decoded audio frame, and wherein the background noise is synthesized based on the received location information. 20. A method for audio playback comprising: receiving a first decoded audio frame, the first decoded audio frame including active speech; generating information associated with background noise in the first decoded audio frame; detecting a transition to an inactive region after the first decoded audio frame; and synthesizing background noise based on the information associated with the background noise in response to the detected transition to the inactive region.
PCT/US2025/028522 2024-05-14 2025-05-08 Decoder silence generation without coded silence description Pending WO2025240232A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GR20240100345 2024-05-14
GR20240100345 2024-05-14

Publications (1)

Publication Number Publication Date
WO2025240232A1 true WO2025240232A1 (en) 2025-11-20

Family

ID=96094648

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/028522 Pending WO2025240232A1 (en) 2024-05-14 2025-05-08 Decoder silence generation without coded silence description

Country Status (1)

Country Link
WO (1) WO2025240232A1 (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20130332175A1 (en) * 2011-02-14 2013-12-12 Fraunhofer-Gesellschaft Zur Forderung Der Angewandten Forschung E.V. Audio codec using noise synthesis during inactive phases
US11526734B2 (en) 2019-09-25 2022-12-13 Qualcomm Incorporated Method and apparatus for recurrent auto-encoding
WO2023064735A1 (en) * 2021-10-14 2023-04-20 Qualcomm Incorporated Audio coding using machine learning based linear filters and non-linear neural sources
US20230215420A1 (en) * 2020-07-21 2023-07-06 Ai Speech Co., Ltd. Speech synthesis method and system

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20130332175A1 (en) * 2011-02-14 2013-12-12 Fraunhofer-Gesellschaft Zur Forderung Der Angewandten Forschung E.V. Audio codec using noise synthesis during inactive phases
US11526734B2 (en) 2019-09-25 2022-12-13 Qualcomm Incorporated Method and apparatus for recurrent auto-encoding
US20230215420A1 (en) * 2020-07-21 2023-07-06 Ai Speech Co., Ltd. Speech synthesis method and system
WO2023064735A1 (en) * 2021-10-14 2023-04-20 Qualcomm Incorporated Audio coding using machine learning based linear filters and non-linear neural sources

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
"Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders", ARXIV:2102.06610V1, 12 February 2021 (2021-02-12)
REO YONEYAMA ET AL: "Unified Source-Filter GAN: Unified Source-filter Network Based On Factorization of Quasi-Periodic Parallel WaveGAN", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 27 June 2021 (2021-06-27), XP081980797 *
ZHANG: "Autoencoder and its various variants", 2018 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS

Similar Documents

Publication Publication Date Title
EP4029016B1 (en) Artificial intelligence based audio coding
EP4416726B1 (en) Audio coding using machine learning based linear filters and non-linear neural sources
KR101785885B1 (en) Adaptive bandwidth extension and apparatus for the same
EP4416722B1 (en) Audio coding using combination of machine learning based time-varying filter and linear predictive coding filter
RU2636685C2 (en) Decision on presence/absence of vocalization for speech processing
KR20130138362A (en) Audio codec using noise synthesis during inactive phases
O’Shaughnessy Review of methods for coding of speech signals
CN109243478A (en) System, method, equipment and the computer-readable media sharpened for the adaptive resonance peak in linear prediction decoding
WO2025240232A1 (en) Decoder silence generation without coded silence description
WO2025240195A1 (en) Generative audio codec for signal synthesis based on spectral envelope features and pitch information
WO2025240227A1 (en) Generative audio codec for signal synthesis based on joint coding of spectral envelope features and pitch information
WO2025240231A1 (en) Generative audio codec for signal synthesis based on groupwise joint coding of spectral envelope features and pitch information
WO2025240194A1 (en) Signal synthesis using temporal upsampling of input features
WO2025240193A1 (en) Signal synthesis using dynamic decoder complexity scaling based on device load
WO2025240228A1 (en) Systems and methods for delay reduction using recurrent layers
Sun et al. Speech compression
WO2025240197A1 (en) Systems and methods for differentiable pulse generation
WO2025240233A1 (en) Systems and methods for differentiable pulse generation
TW202605806A (en) Generative audio codec for signal synthesis based on spectral envelope features and pitch information
WO2025240230A1 (en) Systems and methods for pitch estimation
WO2025240229A1 (en) Systems and methods for spectral feature learning for audio
TW202601634A (en) Improved pitch content in neural speech synthesis
Zhang et al. Low Bit-Rate Speech Coding Based upon GMD-LPCNet
WO2025240223A1 (en) Systems and methods of selecting one or more machine learning model processing branches based on audio data classification

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25732598

Country of ref document: EP

Kind code of ref document: A1