WO2024258481A1 - Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding - Google Patents
Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding Download PDFInfo
- Publication number
- WO2024258481A1 WO2024258481A1 PCT/US2024/023863 US2024023863W WO2024258481A1 WO 2024258481 A1 WO2024258481 A1 WO 2024258481A1 US 2024023863 W US2024023863 W US 2024023863W WO 2024258481 A1 WO2024258481 A1 WO 2024258481A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- sample
- language
- training
- training data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/69—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for evaluating synthetic or decoded voice signals
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/005—Language recognition
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/16—Speech classification or search using artificial neural networks
Definitions
- the present disclosure relates to processing and evaluating synthesized speech.
- Synthesized speech can refer to machine-generated speech that is used to communicate or convey information to listeners.
- the quality of synthesized speech can depend on the content of the speech as well as audible characteristics such as tone or emphasis. There is a need to predict a quality or a comprehensibility of synthesized speech in order to evaluate the computerized system that generates the synthesized speech.
- the present disclosure is related to a method for evaluating synthesized speech, comprising receiving, via processing circuitry, a speech sample in a first language; and determining, via the processing circuitry, a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech audio and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
- the present disclosure is related to a device comprising processing circuitry configured to receive a speech sample in a first language, and determine a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
- the present disclosure is related to a non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method, the method comprising receiving a speech sample in a first language; and determining a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
- FIG. 1 is a schematic of a networked system for evaluating speech
- FIG. 2 is a method for training and using a statistical model to evaluate speech
- FIG. 3 is a schematic of a hardware system for performing a method, according to an exemplary embodiment of the present disclosure.
- FIG. 4 is a schematic of a hardware configuration of a device for performing a method, according to an exemplary embodiment of the present disclosure.
- Natural language generation is a process by which machines can synthesize and output artificial speech.
- Artificial speech, or synthesized speech can be implemented in a variety of devices and environments to facilitate speech-based and audio-based interactions between humans and machines and can be used as an alternative for, or a supplement to, visual content.
- Text-to-speech can refer to the pipeline or system for generating artificial speech wherein a text is generated, analyzed, and converted to an audio waveform that is output by a machine as synthesized speech.
- the effectiveness of a TTS system can depend on a variety of factors related to consumption and processing of audio content, such as the content, intelligibility, prosody, pitch, and dynamics of the synthesized speech.
- the audible characteristics of the synthesized speech need to match the content (e.g., the text) as well as the context of the synthesized speech. In some instances, it can also be desirable to mask the automation involved in synthesizing speech so that a listener believes the speech may have been generated (spoken) by another human rather than by a machine.
- the quality of synthesized speech is evaluated to determine a naturalness of the synthesized speech.
- naturalness can be related to a similarity to a real or a hypothetical human utterance.
- naturalness can be related to how effective or clear the synthesized speech is.
- Quality of synthesized speech can also be evaluated based on a correctness of the synthesized speech (e.g., whether the synthesized speech includes mispronounced words) and an appropriateness of the synthesized speech.
- the quality of the synthesized speech is determined by human evaluators with or without a reference to human utterance.
- the assessment can be subjective, and the use of human evaluators can be inefficient and impractical when scaling large datasets of natural language generation outputs. There is therefore a need to develop an automated model that can effectively and accurately evaluate quality of synthesized speech.
- the present disclosure is directed towards systems and methods for evaluating synthesized speech using a statistical model that has been trained in stages on a combination of multilingual text and speech data.
- the statistical model as referenced herein can refer to an artificial intelligence (Al) model or a model that has been trained for evaluation using machine learning techniques on the combination of multilingual text and speech data.
- the statistical model can include a neural network of one or more layers of neurons between an input layer and an output layer.
- the statistical model can be an encoding model, wherein the encoding model can encode an input (e.g., a sample of synthesized speech) to generate an output (e.g., an evaluation metric).
- the statistical model can be stored, trained, and hosted or executed by one or more servers.
- the server can store and/or access audio data, including synthesized speech, and can evaluate the synthesized speech for a similarity to human speech using the statistical model.
- the server can run the statistical model in order to assign a score to the speech to quantify the relationship between the speech and human speech.
- the server can extract characteristics of the speech that are related to, or indicative of, a similarity to human speech. According to some examples, the training and execution of the statistical model can be distributed across more than one server.
- FIG. 1 is a schematic of a networked device, such as a server 1500, in communication with one or more devices wherein the server can evaluate speech samples, according to one embodiment of the present disclosure.
- the user devices 1100, 1101, 1 lOn can be, for example, computers, such as those that will be described in further detail herein with reference to FIG. 3 and FIG. 4.
- the server 1500 can be a hardware device such as a computer, as described herein with reference to FIG. 3 and FIG. 4.
- the server 1500 can be in communication with the one or more devices via a communication network 1200 such that data can be transmitted between the devices and the server.
- the server can store the statistical model and can train and run the statistical model.
- the server can store the training data or can retrieve the training data from a remote device, such as a second server 1501.
- the server can retrieve the training data from the remote device via the communication network 1200.
- the server can be more than one server, wherein the more than one server can store and/or run parts (e.g., layers) of the statistical model.
- a device 1100 can access the statistical model by transmitting an input to the server 1500 over the communication network 1200.
- the input can be, for example, a speech sample to be evaluated.
- the server 1500 can receive the speech sample and run the statistical model using the speech sample as an input to generate an output, wherein the output is an evaluation of the input speech sample.
- the server 1500 can transmit the output over the communication network 1200 to the device 1100.
- the device can store and run the statistical model or a portion of the statistical model locally on the device hardware.
- the statistical model can be referred to herein as a speech evaluation model.
- the statistical model can include an encoder, wherein the encoder can generate a map, vector, or similar structure to characterize an input, such as a speech sample.
- the map or vector can represent the features of the input that the statistical model is trained to identify.
- the output of the encoder can be used as an input into a decoder, wherein the decoder can map the input to a second output based on the features of the input that are encoded by the encoder.
- the second output can be, for example, a rating of the speech sample.
- the statistical model can include, for example, at least one deep neural network, such as a convolutional neural network (CNN) or re current neural network (RNN), or a transformer.
- CNN convolutional neural network
- RNN re current neural network
- the structure of the statistical model can include combinations of known neural networks and/or can include modified layers in known neural networks to optimize the learning ability and/or the evaluation accuracy of the statistical model.
- the parameters of the statistical model can be set during the training stages, as well be described herein.
- the statistical model can be trained using self-supervised learning (SSL) or semi-supervised learning.
- the statistical model can be trained for a combination of tasks related to speech evaluation, including, but not limited to, phoneme recognition, speaker or source identification, evaluation of emotional states, or speech comprehension.
- FIG. 2 is a method for training the statistical model and evaluating speech, according to one embodiment of the present disclosure.
- the server can train the statistical model in a first training stage 105 and a second training stage 106.
- the first training stage can be referred to herein as a pre-training stage, while the second training stage can be referred to herein as a fine- tuning stage.
- the server can train the statistical model in the first training stage using a first set of training data and in the second training stage using a second set of training data.
- the server can perform one or more pre-training tasks by running the statistical model on the first set of training data in step 110 to output a prediction of speech or text.
- the pre-training tasks can include, for example, predicting masked language or mapping speech to text.
- the server can compare the output of the statistical model with the first set of training data to determine an accuracy of the statistical model, e.g., a prediction accuracy.
- the server can calculate a loss function in step 120, wherein the loss function models the difference, deviation, or distance between the output of the statistical model, e.g., a prediction, and the first set of training data, which includes the actual value that was predicted by the statistical model.
- the server can determine one or more parameters of the statistical model that reduce or minimize the output of the loss function.
- the one or more parameters can include, for example, neuron weights, biases, dropout rates, etc.
- the server can set the one or more parameters of the statistical model in step 130.
- the server can perform the pre-training task again by running the modified statistical model on the first set of training data.
- the server can repeat the steps of the first training stage 105, including performing a pre-training task in step 110, calculating the loss function in step 120, and adjusting the statistical model in step 130 until the output of the loss function is below a threshold of acceptability.
- the server can perform one or more fine-tuning tasks after pre-training is complete by running the statistical model on the second set of training data.
- the fine-tuning tasks can be, for example, related to evaluating the quality of a speech sample.
- the server can modify the statistical model before performing the one or more fine- tuning tasks.
- the modifications can include, for example, modifying one or more layers of the statistical model or modifying a hyperparameter of the statistical model.
- the server can freeze a portion of the statistical model before the one or more fine-tuning tasks.
- Freezing can refer to fixing the parameters (e.g., weights) and architecture of the one or more layers that have already been configured during the first training stage in order to preserve the results of the first training stage.
- the second set of training data can include speech samples that have been rated in a quality evaluation test.
- the server can perform the finetuning task by providing the second set of training data as an input to the statistical model.
- the server can run the statistical model on the second set of training data to generate a rating output for each sample in the second set of training data in step 140.
- the server can compare the rating output for each sample to the actual rating of the speech sample, which is part of the second set of training data.
- the server can calculate a second loss function in step 150, wherein the second loss function models the difference, deviation, or distance between the rating output for each sample and the actual rating of the sample.
- the server can determine one or more parameters of the statistical model that minimize the output of the second loss function.
- the one or more parameters can include, for example, neuron weights, biases, dropout rates, etc.
- the server can set the one or more parameters of the statistical model in step 160.
- the server can perform the fine-tuning task again by running the modified statistical model on the second set of training data.
- the server can repeat the steps of the second training stage 106, including performing a fine-tuning task in step 140, calculating the second loss function in step 150, and adjusting the statistical model in step 160 until the output of the second loss function is below a threshold of acceptability.
- the server can evaluate a speech sample using the trained statistical model in step 107.
- the server can receive a speech sample in step 170.
- the speech sample can be, for example, a synthesized speech sample from a TTS system.
- the server can use the statistical model to evaluate the speech sample in step 180.
- the evaluation of the speech sample can include encoding the speech sample to extract features of the speech sample.
- the server can assign a rating to the speech sample based on the extracted features of the speech sample and output the rating in step 190.
- the rating can be assigned according to the fine-tuning training of the statistical model and can correspond to a quality of the speech sample, such as correctness, naturalness, appropriateness, similarity to human speech, etc.
- the server can perform one or more pretraining tasks using various representations of language in the first set of training data in order to train the statistical model for language recognition and processing.
- the first set of training data can include unlabeled speech data, unlabeled text data, and paired speech and text data.
- the unlabeled speech data can include recordings of speech, recordings of text that is read aloud (without the corresponding text), and recordings of conversations between speakers.
- the unlabeled text data can include text from various sources, such as books, articles, written correspondence, etc.
- the paired speech and text data can include speech data and corresponding text representations of the speech data.
- the paired speech and text data can include speech data that originates from text data as well as text data that originates from speech data.
- the paired speech and text data can include an audio recording of a reading from a book paired with the text of the book as a first sample and a recording of a conversation paired with a transcript of the conversation as a second sample.
- the speech data and the text data used in the paired speech and text data can be different from, or can overlap with, the unlabeled speech data and the unlabeled text data.
- the training data can include speech and text data of varying length, formality, and complexity.
- the audio of the unlabeled speech data and the paired speech data can include non-speech audio and noise from various sources.
- the server can train the statistical model using character-based tokenization by splitting the training data into individual characters as inputs.
- Character-based tokenization can be useful when the training data is multilingual, as will be discussed in further detail herein.
- Word-based tokenization and sentence piece tokenization which is language-independent, are also compatible with the present disclosure.
- the first set of training data can include multilingual speech data and text data.
- the unlabeled speech data can include readings in various languages and the unlabeled text data can include text in various languages, e.g., up to 65 languages, 65 languages, or more than 65 languages.
- the languages in the first set of training data can include low-resource languages and high-resource languages.
- the server can up-sample low-resource languages to provide more samples for training, e.g., using temperature sampling.
- the first set of training data can include native speech and text that has been generated in a first language as well as translated speech and text, e.g., samples that have been translated from a second language to the first language.
- the use of a multilingual dataset can improve the recognition ability of the statistical model for each individual language included in the multilingual dataset. For example, training the statistical model on languages such as Spanish or German can improve the performance of the statistical model in evaluating English speech samples when compared with a model that is only trained on English data.
- the server can train the statistical model on a number of pretraining tasks related to language comprehension in the first training stage. Examples of the pretraining tasks can include, but are not limited to, mapping speech and text, predicting words in speech or text, and determining a sequence of words in speech or text.
- the server can perform the pretraining tasks to train the statistical model to recognize and represent language in speech as well as in text.
- the server can train the statistical model to map speech data to text data and/or text data to speech data.
- the statistical model can determine relationships between the two language modalities, as well as relationships within speech data and text data.
- the server can generate vectors in a vector space to represent speech data and/or text data based on the first set of training data.
- the vectors can include mapping vectors between speech data and text data.
- the paired speech and text data can be used as training examples for how audio data (speech) corresponds to text data.
- the server can map speech data to text data and similarly text data to speech data in order to represent a relationship between the two language modalities.
- the mapping can include building contextualized representations of speech data and/or text data.
- the server can train the statistical model using more than one pretraining task in parallel.
- the server can train the statistical model to carry out the pretraining tasks on each of the types of data in the first set of training data (unlabeled speech, unlabeled text, paired speech and text) or on a subset of data in the first set of training data.
- the server can train the statistical model to map speech data to text data using connectionist temporal classification (CTC) during the first stage of training.
- CTC can refer to a classification of sequences of data based on a likelihood of alignment between the sequences at a point in time.
- an audio data sample can be paired with a corresponding text representation (transcript) of the audio in the paired speech and text data.
- the server can train the statistical model to predict an alignment between paired speech data and text data using CTC.
- the alignment can include synchronization of a point in time in the speech data (e.g., when a syllable is spoken) with a corresponding character or group of characters representing the point of time (e.g., the syllable) in the text data.
- the server can input the speech data into the statistical model and can train the statistical model using the paired text data as a target.
- the server can train the statistical model to map the speech data input to the paired text data with the proper alignment.
- CTC can provide probability distributions for the likelihood of alignment between two sequences of data at a point in time.
- the server can use the statistical model to predict one or more possible alignments between two sequences, wherein the server can calculate a score or probability of each alignment.
- the server can calculate the likelihood probability of each alignment by computing a loss function for alignments within a sequence.
- a neural network such as an RNN, can be used to estimate probabilities of alignments within a sequence.
- the alignments can then be merged to determine a probability of an alignment as a whole.
- the server can train the statistical model to determine an increased or maximal likelihood of alignment between two sequences.
- the server can assign one or more classification labels to an input data sequence at a point in time.
- the classification label can correspond to a point in a target data sequence, indicating an alignment with the input data sequence at that point.
- the classification label can be associated with a probability of the point in the input data sequence being aligned with the point in the target data sequence.
- the classification label can be a phoneme identified in a speech sample, the phoneme being part of a word or group of words in a paired text sample (transcript).
- the server can use CTC to determine a probability that the phoneme is an accurate representation of what was said in the audio sample.
- the server can train the statistical model using CTC with a coefficient for paired CTC loss of approximately 0.03. The training of the statistical model for CTC using the first set of training data can improve the performance of the statistical model in aligning speech and text representations of language.
- the server can mask the training data in order to train the statistical model to predict speech and/or text.
- Masking can refer to the removal (masking) of components in a sequence of text or speech.
- the components can be, for example, a word or a sequence of words.
- the server can mask a sample from the training data and use the statistical model to predict the word or sequence of words that would fill in the masked components in the sample.
- the server can mask both the speech data and the text data such that the statistical model is trained to predict missing speech as well as missing text.
- the server can mask the paired speech and text data.
- the server can thus train the statistical model to predict how sentences are formed and how missing speech or text data can be replaced.
- Masking can improve the ability of the statistical model to process context and meaning in language. Training the statistical model on masked speech data as well as masked text data results in cross-modal transfer of learning for the different representations of language to improve prediction ability for both speech data and text data.
- the server can mask the training data and use the predicted output for CTC.
- the server can mask paired speech and text data.
- the server can train the statistical model to predict the missing speech and text for the respective data samples and generate predicted speech data and predicted text data.
- the server can use the predicted speech data as an input into the statistical model for CTC training, with the paired text data as the target data.
- the server can train the statistical model to align the predicted speech data with the predicted text data.
- the text data can be a character-level transcript.
- the combination of predicting masked data and alignment using CTC in sequential or linked training steps can improve the accuracy of alignment of the statistical model.
- the server can use the statistical model to predict whether a first sample of training data is followed by a second sample of training data. For example, a recorded speech can be split into multiple audio samples. In one embodiment, the server can use the statistical model to predict whether a second sample directly follows a first sample. In one embodiment, the server can use the statistical model to reconstruct the full speech by determining an order of the audio samples based on the content of the audio samples.
- the content of the audio samples can include the content of the spoken language as well as additional speech cues, such as the tone of the speaker’s voice, changes in volume or pitch, etc.
- the server can identify the content of the audio samples based on the training of the statistical model.
- the server can use the statistical model to recognize continuity as well as incongruities (contrast) between and within the audio samples.
- the first stage of training can include iterative training steps.
- the server can modify the statistical model during training in order to improve performance on the pretraining tasks that have been described herein.
- the server can generate or can retrieve paired speech and text samples for training.
- the paired samples can have known characteristics, such as a language, an alignment between speech and text, an audio quality, etc.
- the server can then use a portion of the generated sample as an input for a pretraining task of the statistical model.
- the server can mask a portion of the sample and use the statistical model to predict the word or sequence of words that is missing from the sample.
- the server can compare a prediction output of the statistical model with the unmasked sample and can determine a loss value as a measure of the difference or deviation between the predicted output and the unmasked sample.
- the loss value can be, for example, a multi-class loss, a regression loss, or any similar loss value that can be calculated using a loss function.
- the server can modify the statistical model in order to minimize the output of the loss function.
- the modifications can include changes to parameters of the statistical model, such as nodes, weights, or layers.
- the parameters can define the statistical model and can modify the mapping and output of the statistical model.
- the server can modify hyperparameters of the statistical model. The hyperparameters can affect the training of the statistical model.
- the hyperparameters can include a learning rate or a batch size of training data that is input to the statistical model.
- the server can repeat the pre-training task (e.g., predicting a masked word) until the output of the loss function is below a threshold of acceptability.
- the server can train the statistical model to recognize representations of language in the first training stage.
- the recognition of speech can include distinguishing spoken language from other sounds and noise in an audio sample.
- the recognition of speech can include determining the language of the speech and the linguistic content of the speech.
- the server can train the statistical model to determine characteristics of the speech or the speaker, such as a context or a tone.
- the characteristics of the speech can be dynamic and time-varying.
- the server can use the statistical model to determine if speech is present in an audio sample, the location of the speech in the audio sample, and the language of the speech in the audio sample.
- the server can use the statistical model to generate a text representation of the speech in the audio sample.
- the text representation can be generated based on the vector space that was generated by the statistical model during the training stage.
- the server can use an Adam optimizer for the first training stage with a Transformer learning rate schedule.
- the server can increase the learning rate of the statistical model during the first training stage, followed by an inverse square root decay of the learning rate during the first training stage.
- the server can apply a loss (e.g., dropout) to the statistical model to avoid overfitting.
- the coefficient of speech loss can be 1.0
- the coefficient of text loss can be 0.3
- the coefficient of paired CTC loss can be 0.03.
- the server can fine-tune the statistical model in the second training stage to evaluate the quality of a speech sample.
- the statistical model can be fine-tuned to evaluate a speech sample and assign a score to the speech sample.
- the score can be a measure of the quality of the speech sample.
- the score can be a measure of a relationship between the sample and human speech.
- the score can be a measure of a relationship between a first speech sample and a second speech sample.
- a first speech sample can be synthesized by a first speech synthesis system and a second speech sample can be synthesized by a second speech synthesis system.
- the second system can be, in some cases, an alternative system or a different version of the first system.
- the second system can be a known or trusted speech synthesis system with an output that has been previously evaluated or verified.
- the statistical model can be fine-tuned to evaluate and compare the quality of the first speech sample and the second speech sample.
- the score can be related to or based on a mean opinion score (MOS), which is a metric of quality evaluation used for telecommunications.
- MOS mean opinion score
- Mean opinion scores are typically determined in a quality evaluation test that includes categorical rating scales. Subjects listen to speech samples and assign ratings to the speech samples based on how similar the speech sounds to human speech. The ratings can indicate whether the listener believes that the speech sample is produced by a machine or by a human.
- a rating can be a quantity that corresponds to a categorical rating scale.
- the MOS rating can be a mean of the ratings assigned by human listeners to a speech sample in the quality evaluation test. Obtaining MOS ratings from human listeners can be inefficient and expensive.
- the prediction of MOS ratings can provide an accurate and consistent metric for evaluating the quality of speech samples.
- the ratings can be an evaluation of the similarity between the sample and human speech or a likelihood that the sample is human speech.
- the server can use the statistical model of the present disclosure to process the speech samples and predict an evaluation using the MOS ratings.
- the prediction of MOS ratings can indicate whether machine-synthesized speech samples can effectively imitate human speech.
- a speech sample can be synthesized by a TTS model.
- the server can evaluate the synthesized speech sample using the statistical model to determine whether the TTS model is capable of artificially synthesizing humanlike speech.
- the server can use the statistical model to evaluate the quality of a speech sample using a different metric known to one of ordinary skill in the art.
- the second set of training data for the second training stage can include speech samples and ratings of the speech samples.
- the speech samples can include synthesized speech samples from a machine as well as natural speech samples from a human.
- the training data can include a library of speech samples and the MOS ratings that were assigned to the speech samples during quality evaluation tests.
- the synthesized speech samples can originate from more than one speech generation system. The length of the speech samples can vary.
- the speech samples can be utterances, wherein each utterance is assigned a rating from a quality evaluation test.
- An utterance can refer to a continuous piece of speech.
- an utterance can be a sequence that is spoken without pauses interrupting the sequence.
- a length of an utterance can vary from a single syllable to a sequence of words.
- the server can train the statistical model for utterancelevel evaluation. In some embodiments, the server can train the statistical model for frame-level or system-level evaluation.
- the speech samples can include speech samples in various languages. In one example, the languages used in the second training stage can be a subset of the languages used in the first training stage and/or can include languages that were not included in the first training stage.
- the server can modify the speech samples in the second set of training data before fine-tuning the statistical model, e.g., by changing the speed of utterance in the sample, trimming the sample, or buffering the sample with silence.
- the second set of training data can include text data that has been rated for quality evaluation. In one embodiment, the second set of training data can include language identifiers for the utterances in the second set of training data.
- the server can fine-tune the statistical model that has been trained in the first training stage by freezing one or more layers of the statistical model and using the statistical model to perform one or more fine-tuning tasks using the second set of training data.
- one or more layers of the model are not frozen.
- the server can modify the parameters of the layers that are not frozen as a result of the training on the second set of training data.
- the server can fine-tune the statistical model that has been trained in the first training stage by adding one or more layers to the statistical model, wherein the additional layers have not been previously trained or configured. The server can then train the statistical model using the second set of training data.
- the additional layers will be trained from scratch on the second set of training data, while the original layers can be frozen or can be modified based on the fine-tuning tasks.
- the additional layers can be one or more output layers.
- the additional layers can be pooling layers.
- the additional layers can include fully connected layers for linear projection and rescaling. The parameters of linear projection functions can be determined during the second training stage.
- the server can fine-tune the statistical model that has been trained in the first training stage by changing the learning rate of the model during the second training stage.
- the server can set the learning rate of the statistical model to increase linearly during the first training stage, followed by a decay in the learning rate during the second training stage.
- the decrease in the learning rate during the second training stage can preserve the results of the first training stage while tuning the statistical model on the second set of training data, which may be smaller than the first set of training data.
- the methods for fine-tuning the statistical model can be implemented by the server individually or in combination with each other or with any other method for fine-tuning a statistical model known to one of ordinary skill in the art.
- the server can fine-tune the statistical model on one or more tasks related to speech evaluation in a similar manner to the first training stage. For example, the server can input a speech sample that has been assigned a quality evaluation rating to the statistical model and determine a rating of the speech sample using the statistical model. The server can compare the rating output of the statistical model to the actual rating of the speech sample to determine a loss value based on a loss function. The server can modify parameters of the statistical model to minimize the output of the loss function. The server can repeat the fine-tuning tasks until the output of the loss function is below a threshold of acceptability.
- the use of multilingual training data can result in cross-lingual transfer of learning, wherein the prediction accuracy of the server for each language can improve as a result of the statistical model being trained on the number of languages.
- the server can expose the statistical model to samples of varying quality across the number of languages in the training data and can train the statistical model to recognize characteristics of speech in each language.
- Poor speech quality can refer to speech that is robotic or sounds like machine-synthesized speech.
- Samples with poor speech quality can include characteristics such as mispronunciation, lack of auditory clarity or intelligibility, or lack of tone and affect in the speech.
- Incoherence, or lack of meaning, in the words or word groupings in a sample can also be content-based characteristics of poor speech quality.
- Good speech quality can refer to speech that sounds like human speech. Samples with good speech quality can include characteristics, especially temporal characteristics, such as dynamic volume, tone, emphasis, pacing, etc.
- the server can train the statistical model to identify and encode the characteristics of speech across the number of languages. Identification and encoding of the characteristics in a first language can be transferrable and can enable identification and encoding of the characteristics in a second language. The identification and encoding of characteristics of speech in each language can affect how the statistical model can be used to encode and evaluate speech samples in any language.
- the use of a combination of unlabeled speech, unlabeled text, and paired speech and text data in the training data can result in cross-modal transfer of learning, wherein the prediction accuracy of the server in evaluating speech samples can improve based on the training with unlabeled text and paired speech and text data when compared with a statistical model that is only trained on speech data.
- the training on the unlabeled text and the paired speech and text data can provide additional representations of language and the characteristics of language. Training the statistical model to process and encode the unlabeled text and the paired speech and text can improve the accuracy of the server in processing unlabeled speech.
- training the statistical model to process and encode unlabeled speech can improve the accuracy of the server in processing unlabeled text.
- the server can train the statistical model in two stages to process a speech sample and assign a rating, such as an MOS rating, to the speech sample as a measure of the quality of the speech sample.
- a rating such as a MOS rating
- the systems and methods of the present disclosure can also be used to evaluate speech samples based on qualities including, but not limited to, naturalness, correctness, and appropriateness.
- the statistical model can be trained on speech samples that have been assigned indicators (e.g., ratings) of any of these qualities.
- the server can use the statistical model to predict whether the speech sample is produced by a machine or by a human.
- the rating of the speech sample can correspond to a likelihood or probability that the speech sample is produced by a machine.
- the rating of the speech sample can correspond to a quality of the speech sample, such as naturalness, correctness, accuracy.
- the server can use the statistical model to predict whether the speech sample will be processed and interpreted as human speech by a listener.
- the server can use the statistical model to recognize and evaluate speech samples in a number of languages. The server does not need to encode the language of the speech sample as an input into the statistical model.
- the server can receive an input, wherein the input is a speech sample.
- the input can include metadata related to the speech sample.
- the metadata can include, but is not limited to, language identifiers of utterances in the speech sample or an input text corresponding to the speech sample (e.g., a transcript or a source text used for TTS synthesis).
- the language identifier can include a locale tag indicating language and other characteristics of the speech sample, such as cultural or location-specific references or modes of speech.
- the server can read a language identifier and can generate or correct predictions based on the identified language.
- the server can evaluate a speech sample without a language identifier.
- the server can receive the input from a device over a communication network.
- the server can determine an MOS rating for the speech sample by running the statistical model on the speech sample.
- the MOS rating can be a measure of the quality of the speech sample.
- the server can output the MOS rating as an evaluation of the speech sample input.
- the server can transmit the output to the device that transmitted the speech sample.
- the server can use the statistical model for split testing or comparative evaluation. For example, the server can input a first speech sample and a second speech sample to the statistical model. The server can use the statistical model to evaluate each speech sample and determine a relative quality of each speech sample.
- the relative quality can include a first MOS rating for the first speech sample and a second MOS rating for the second speech sample.
- the relative quality can include a determination of whether the first speech sample or the second speech sample sounds more natural, e.g., more like human speech.
- the server can thus use the statistical model to evaluate speech samples from different speech generation systems, such as different versions of a speech generation system or a control speech generation system and an experimental speech generation system.
- the server can train the statistical model to output a determination of which speech sample is of higher quality when the first speech sample and the second speech sample are input to the statistical model.
- the server can train the statistical model to generate and evaluate intermediate representations or transformations of utterances.
- the server can use the statistical model to generate a representation of the speech sample based on the speech sample.
- the server can then use the statistical model to evaluate the representation of the speech sample to determine a quality of the speech sample.
- the server can use the statistical model to evaluate the quality of text samples.
- the server can train the statistical model using the combination of unlabeled speech, unlabeled text, and paired speech and text data, as has been described herein.
- the server can fine-tune the statistical model using rated text data in addition to or in place of the speech data.
- the server can input a text sample into the statistical model, wherein the text sample was synthesized by a natural language processor.
- the server can use the statistical model to predict a rating for the text sample, the rating being a measure of a likelihood that the text sample was written by a human rather than synthesized by a machine.
- the server can determine using the statistical model that the text sample is of good quality and that it resembles human language. If the natural language processor is not effective, the server can determine using the statistical model that the text sample is of poor quality and that it does not resemble human language.
- the characteristics of the text sample that can indicate the quality of the text sample can include word usage, structure, grammar, organization, meaning, etc.
- the server can use the statistical model to evaluate text samples in various languages without encoding or inputting the language of the text sample to the statistical model.
- the server can use the statistical model to evaluate the quality of a paired speech and text sample.
- the paired speech and text sample can be, for example, a synthesized text sample and the speech that was produced based on the synthesized text sample by a TTS system.
- Embodiments of the subject matter and the functional operations described in this specification can be implemented by digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
- Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non- transitory program carrier for execution by, or to control the operation of data processing apparatus, such as the networked device or server 1500 and 1501, the devices 1100, 1101, 1 lOn, and the like.
- the computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
- data processing apparatus refers to data processing hardware and may encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers.
- the apparatus can also be or further include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
- the apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
- a computer program which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, Subroutine, or other unit suitable for use in a computing environment.
- a computer program may, but need not, correspond to a file in a file system.
- a program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code.
- a computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
- Computers suitable for the execution of a computer program include, by way of example, general or special purpose microprocessors or both, or any other kind of central processing unit.
- a CPU will receive instructions and data from a read-only memory or a random access memory or both.
- Elements of a computer are a CPU for performing or executing instructions and one or more memory devices for storing instructions and data.
- a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.
- mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.
- a computer need not have such devices.
- a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
- Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD-ROM and DVD-ROM disks.
- the processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
- a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer.
- a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor
- a keyboard and a pointing device e.g., a mouse or a trackball
- Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
- a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to
- Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more Such back-end, middleware, or front-end components.
- the components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
- LAN local area network
- WAN wide area network
- the computing system can include clients (user devices) and servers.
- a client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
- a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the user device, which acts as a client.
- Data generated at the user device e.g., a result of the user interaction, can be received from the user device at the server.
- FIG. 3 An example of a type of computer is shown in FIG. 3.
- the computer 500 can be used for the operations described in association with any of the computer-implement methods described previously, according to one implementation.
- the computer 500 can be an example of a networked device such as the servers 1500 and 1501 including processing circuitry, as discussed herein.
- the computer 500 can be an example of any of the user devices 1100, 1101, 1 lOn in communication with the networked device via the network 1200.
- the processing circuitry includes one or more of the elements discussed next with reference to FIG. 3.
- the computer 500 includes a processor 510, a memory 520, a storage device 530, and an input/output device 540.
- the processor 510 is capable of processing instructions for execution within the system 500.
- the processor 510 is a single-threaded processor.
- the processor 510 is a multi -threaded processor.
- the processor 510 is capable of processing instructions stored in the memory 520 or on the storage device 530 to display graphical information for a user interface on the input/output device 540.
- the memory 520 stores information within the computer 500.
- the memory 520 is a computer-readable medium.
- the memory 520 is a volatile memory unit.
- the memory 520 is a non-volatile memory unit.
- the storage device 530 is capable of providing mass storage for the computer 500.
- the storage device 530 is a computer-readable medium.
- the storage device 530 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device.
- the input/output device 540 provides input/output operations for the computer 500.
- the input/output device 540 includes a keyboard and/or pointing device.
- the input/output device 540 includes a display unit for displaying graphical user interfaces.
- the device 601 which can be any of the above described devices, including the servers 1500 and 1501 or any of the user devices 1100, 1101, 1 lOn, includes processing circuitry.
- the processing circuitry includes one or more of the elements discussed next with reference to FIG. 4.
- the process data and instructions may be stored in memory 602. These processes and instructions may also be stored on a storage medium disk 604 such as a hard drive (HDD) or portable storage medium or may be stored remotely.
- a storage medium disk 604 such as a hard drive (HDD) or portable storage medium or may be stored remotely.
- the claimed advancements are not limited by the form of the computer-readable media on which the instructions of the inventive process are stored.
- the instructions may be stored on CDs, DVDs, in FLASH memory, RAM, ROM, PROM, EPROM, EEPROM, hard disk or any other information processing device with which the device 601 communicates, such as a server or computer.
- the claimed advancements may be provided as a utility application, background daemon, or component of an operating system, or combination thereof, executing in conjunction with CPU 600 and an operating system such as Microsoft Windows, UNIX, Solaris, LINUX, Apple MAC-OS and other systems known to those skilled in the art.
- an operating system such as Microsoft Windows, UNIX, Solaris, LINUX, Apple MAC-OS and other systems known to those skilled in the art.
- CPU 600 may be a Xenon or Core processor from Intel of America or an Opteron processor from AMD of America, or may be other processor types that would be recognized by one of ordinary skill in the art.
- the CPU 600 may be implemented on an FPGA, ASIC, PLD or using discrete logic circuits, as one of ordinary skill in the art would recognize.
- CPU 600 may be implemented as multiple processors cooperatively working in parallel to perform the instructions of the processes described above.
- the device 601 in FIG. 4 also includes a network controller 606, such as an Intel Ethernet PRO network interface card from Intel Corporation of America, for interfacing with network 650. and to communicate with the other devices.
- the network 650 can be a public network, such as the Internet, or a private network such as an LAN or WAN network, or any combination thereof and can also include PSTN or ISDN sub-networks.
- the network 650 can also be wired, such as an Ethernet network, or can be wireless such as a cellular network including EDGE, 3G, 4G and 5G wireless cellular systems.
- the wireless network can also be WiFi, Bluetooth, or any other wireless form of communication that is known.
- the device 601 further includes a display controller 608, such as a NVIDIA GeForce GTX or Quadro graphics adaptor from NVIDIA Corporation of America for interfacing with display 610, such as an LCD monitor.
- a general purpose I/O interface 612 interfaces with a keyboard and/or mouse 614 as well as a touch screen panel 616 on or separate from display 610.
- General purpose I/O interface also connects to a variety of peripherals 618 including printers and scanners.
- a sound controller 620 is also provided in the device 601 to interface with speakers/microphone 622 thereby providing sounds and/or music.
- the general purpose storage controller 624 connects the storage medium disk 604 with communication bus 626, which may be an ISA, EISA, VESA, PCI, or similar, for interconnecting all of the components of the device 601.
- communication bus 626 may be an ISA, EISA, VESA, PCI, or similar, for interconnecting all of the components of the device 601.
- a description of the general features and functionality of the display 610, keyboard and/or mouse 614, as well as the display controller 608, storage controller 624, network controller 606, sound controller 620, and general purpose I/O interface 612 is omitted herein for brevity as these features are known.
- Embodiments of the present disclosure may also be set forth in the following parentheticals.
- a method for evaluating synthesized speech comprising receiving, via processing circuitry, a speech sample in a first language; and determining, via the processing circuitry, a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech audio and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
- the method of (1) further comprising receiving a language identifier or a text sample corresponding to the speech sample and determining the rating of the speech sample based on the language identifier or the text sample.
- a device comprising processing circuitry configured to receive a speech sample in a first language, and determine a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
- the first training stage of the artificial intelligence encoding model includes classifying the first set of training data using connectionist temporal classification.
- a non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method, the method comprising receiving a speech sample in a first language; and determining a rating of the speech sample based on an encoding of the speech sample by an artificial intelligence encoding model, the rating of the speech sample corresponding to a naturalness of the speech sample, wherein the encoding of the speech sample is based on a first training stage of the encoding model using a first set of training data and a second training stage of the encoding model using a second set of training data, the first set of training data includes unlabeled speech audio, unlabeled text, and paired speech and text data in the first language and at least one additional language, and the second set of training data includes rated speech audio.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Signal Processing (AREA)
- Evolutionary Computation (AREA)
- Machine Translation (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24724740.6A EP4713919A1 (en) | 2023-06-14 | 2024-04-10 | Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/334,876 | 2023-06-14 | ||
| US18/334,876 US12567433B2 (en) | 2023-06-14 | 2023-06-14 | Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024258481A1 true WO2024258481A1 (en) | 2024-12-19 |
Family
ID=91030184
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/023863 Ceased WO2024258481A1 (en) | 2023-06-14 | 2024-04-10 | Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12567433B2 (en) |
| EP (1) | EP4713919A1 (en) |
| WO (1) | WO2024258481A1 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11704506B2 (en) | 2020-08-26 | 2023-07-18 | Google Llc | Learned evaluation model for grading quality of natural language generation outputs |
-
2023
- 2023-06-14 US US18/334,876 patent/US12567433B2/en active Active
-
2024
- 2024-04-10 WO PCT/US2024/023863 patent/WO2024258481A1/en not_active Ceased
- 2024-04-10 EP EP24724740.6A patent/EP4713919A1/en active Pending
Non-Patent Citations (2)
| Title |
|---|
| ANKUR BAPNA ET AL: "mSLAM: Massively multilingual joint pre-training for speech and text", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 3 February 2022 (2022-02-03), XP091149624 * |
| SELLAM THIBAULT ET AL: "SQuId: Measuring Speech Naturalness in Many Languages", ICASSP 2023 - 2023 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), IEEE, 4 June 2023 (2023-06-04), pages 1 - 5, XP034450305, DOI: 10.1109/ICASSP49357.2023.10094909 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4713919A1 (en) | 2026-03-25 |
| US12567433B2 (en) | 2026-03-03 |
| US20240420726A1 (en) | 2024-12-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Fendji et al. | Automatic speech recognition using limited vocabulary: A survey | |
| KR102757438B1 (en) | Method and computer readable storage medium for performing text-to-speech synthesis using machine learning based on sequential prosody feature | |
| US11929059B2 (en) | Method, device, and computer readable storage medium for text-to-speech synthesis using machine learning on basis of sequential prosody feature | |
| AU2019395322B2 (en) | Reconciliation between simulated data and speech recognition output using sequence-to-sequence mapping | |
| Harere et al. | Quran recitation recognition using end-to-end deep learning | |
| Keshet | Automatic speech recognition: A primer for speech-language pathology researchers | |
| CN102651217A (en) | Method and equipment for voice synthesis and method for training acoustic model used in voice synthesis | |
| Bhatt et al. | Continuous speech recognition technologies—a review | |
| CN109697988B (en) | A voice evaluation method and device | |
| Yanagita et al. | Neural iTTS: Toward Synthesizing Speech in Real-time with End-to-end Neural Text-to-Speech Framework. | |
| Coto‐Solano | Computational sociophonetics using automatic speech recognition | |
| Dumyn et al. | Review of automatic speech recognition systems for Ukrainian and english language | |
| Sinha et al. | Empirical analysis of linguistic and paralinguistic information for automatic dialect classification | |
| CN113053409A (en) | Audio evaluation method and device | |
| Adibian et al. | DeepMine-multi-TTS: a Persian speech corpus for multi-speaker text-to-speech: M. Adibian et al. | |
| US12567433B2 (en) | Automated evaluation of synthesized speech using cross-modal and cross-lingual transfer of language encoding | |
| Varatharaj et al. | Supporting teacher assessment in Chinese language learning using textual and tonal features | |
| Kostek et al. | Synthesizing medical terms–quality and naturalness of the deep Text-to-Speech algorithm | |
| Klessa et al. | Annpro: A desktop module for automatic segmentation and transcription | |
| Nandal et al. | Pronunciation accuracy calculator using machine learning | |
| Oralbekova et al. | Current advances and algorithmic solutions in speech generation | |
| Hassan | A character gram modeling approach towards Bengali Speech to Text with Regional Dialects | |
| Laryea et al. | Automatic Speech Recognition System for Somali in the interest of reducing Maternal Morbidity and Mortality. | |
| Xu | Evaluation of English pronunciation interaction quality based on deep learning | |
| Upadhyay et al. | Pronunciation similarity matching using deep learning |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24724740 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202517128715 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024724740 Country of ref document: EP |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| WWP | Wipo information: published in national office |
Ref document number: 202517128715 Country of ref document: IN |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| ENP | Entry into the national phase |
Ref document number: 2024724740 Country of ref document: EP Effective date: 20251219 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024724740 Country of ref document: EP |