WO2023088091A1 - 语音分离方法、装置、电子设备及可读存储介质 - Google Patents

语音分离方法、装置、电子设备及可读存储介质 Download PDF

Info

Publication number
WO2023088091A1
WO2023088091A1 PCT/CN2022/129118 CN2022129118W WO2023088091A1 WO 2023088091 A1 WO2023088091 A1 WO 2023088091A1 CN 2022129118 W CN2022129118 W CN 2022129118W WO 2023088091 A1 WO2023088091 A1 WO 2023088091A1
Authority
WO
WIPO (PCT)
Prior art keywords
speech
voice
feature
bottleneck
processed
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/129118
Other languages
English (en)
French (fr)
Inventor
汪鑫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Priority to US18/712,467 priority Critical patent/US20250006215A1/en
Publication of WO2023088091A1 publication Critical patent/WO2023088091A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/02Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/04Training, enrolment or model building
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/06Decision making techniques; Pattern matching strategies
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272Voice signal separating
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272Voice signal separating
    • G10L21/0308Voice signal separating characterised by the type of parameter measurement, e.g. correlation techniques, zero crossing techniques or predictive techniques
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks

Definitions

  • the present disclosure relates to the technical field of voice processing, and in particular to a voice separation method, device, electronic equipment and readable storage medium.
  • VAD Voice Activity Detection
  • VAD system + speech recognition separation system speech separation is realized through "VAD system + speech recognition separation system”. Specifically, firstly, the endpoint detection of the speech is performed through the VAD system, and then the speech recognition separation system is used to separate the target speech based on the endpoint detection result.
  • the disclosure provides a voice separation method, device, electronic equipment and readable storage medium.
  • the present disclosure provides a voice separation method, including:
  • a target speech segment matching the reference speech in the speech to be processed is determined according to the speech detection result, wherein the speech feature of the target speech segment matches the bottleneck feature of the reference speech.
  • the input of the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech to the speech separation model, and obtaining the speech detection result output by the speech separation model includes:
  • the speech feature corresponding to the speech to be processed is input to the first neural network included in the speech separation model, and the vector expression corresponding to the speech feature output by the first neural network is obtained;
  • the fusion feature is input to the second neural network included in the speech separation model, a matrix output by the second neural network is obtained, and the speech detection result is obtained based on the matrix.
  • the acquiring the speech detection result based on the matrix includes:
  • the probability values that each of the audio frames belong to the first category and the second category are obtained; the voice features corresponding to the audio frames included in the first category are the same as the reference voice The corresponding bottleneck feature matches, and the speech feature corresponding to the audio frame included in the second category does not match the bottleneck feature corresponding to the reference speech;
  • the voice detection result corresponding to each of the audio frames is used to indicate the audio Whether the speech feature corresponding to the frame matches the bottleneck feature corresponding to the reference speech.
  • the speech separation model is obtained by training based on the speech features corresponding to the sample speech, the bottleneck features corresponding to the sample speech, and the marked speech detection results of the sample speech, and the sample speech Include the reference voice.
  • the speech features include one or more of FBank features, Mel spectrum features, and timbre features.
  • a voice separation device including:
  • An acquisition module configured to acquire speech features corresponding to the speech to be processed and bottleneck features corresponding to the reference speech
  • a speech detection module configured to input the speech features corresponding to the speech to be processed and the bottleneck features corresponding to the reference speech to the speech separation model, and obtain the speech detection result output by the speech separation model;
  • a separation module configured to determine a target speech segment in the speech to be processed that matches the reference speech according to the speech detection result, wherein the speech feature of the target speech segment is similar to the bottleneck feature of the reference speech match.
  • the speech detection module is specifically configured to input the speech features corresponding to the speech to be processed to the first neural network included in the speech separation model, and obtain the output of the first neural network.
  • a vector expression corresponding to the speech feature splicing the vector expression corresponding to the speech feature and the bottleneck feature corresponding to the reference speech to obtain a fusion feature; inputting the fusion feature to the second neural network included in the speech separation model , obtaining a matrix output by the second neural network, and obtaining the speech detection result based on the matrix.
  • the speech detection module is specifically configured to obtain the probability values that each of the audio frames belongs to the first category and the second category respectively according to the elements corresponding to the audio frames included in the matrix;
  • the speech features corresponding to the audio frames included in the first category match the bottleneck features corresponding to the reference speech, and the speech features corresponding to the audio frames included in the second category do not match the bottleneck features corresponding to the reference speech; according to each The audio frames respectively belong to the maximum value among the probability values of the first category and the second category, and determine the speech detection result corresponding to each of the audio frames;
  • the speech detection result corresponding to the audio frame is used to indicate whether the speech feature corresponding to the audio frame matches the bottleneck feature corresponding to the reference speech.
  • the speech separation model is obtained by training based on the speech features corresponding to the sample speech, the bottleneck features corresponding to the sample speech, and the marked speech detection results of the sample speech, and the sample speech Include the reference voice.
  • the speech features include one or more of FBank features, Mel spectrum features, and timbre features.
  • the present disclosure provides an electronic device, including: a memory and a processor;
  • the memory is configured to store computer program instructions
  • the processor is configured to execute the computer program instructions, so that the electronic device implements the speech separation method according to any one of the first aspect.
  • the present disclosure provides a readable storage medium, including: computer program instructions;
  • At least one processor of the electronic device executes the computer program instructions to realize the speech separation method described in any one of the first aspect.
  • the present disclosure provides a computer program product.
  • the computer program product When the computer program product is executed by a computer, the computer is enabled to implement the speech separation method according to any one of the first aspect.
  • Embodiments of the present disclosure provide a voice separation method, device, electronic equipment, and readable storage medium.
  • the method includes: acquiring the voice feature corresponding to the voice to be processed and the bottleneck feature corresponding to the reference voice;
  • the bottleneck feature corresponding to the feature and the reference speech is input to the speech separation model, and the speech detection result output by the speech separation model is obtained; based on the speech detection result, a target speech segment matching the reference speech in the speech to be processed is determined, wherein the The speech features of the target speech segment match the bottleneck features of the reference speech.
  • the system for realizing speech separation is implemented in an end-to-end manner, and by using the speech features corresponding to the speech to be processed and the bottleneck features corresponding to the reference speech as the joint input of the speech separation system, it is used to extract the speech from the speech to be processed Separate the target speech.
  • FIG. 1 is a schematic flowchart of a voice separation method provided by an embodiment of the present disclosure
  • FIG. 2 is a schematic structural diagram of a speech separation model provided by an embodiment of the present disclosure
  • FIG. 3 is a schematic flowchart of a speech separation method provided by another embodiment of the present disclosure.
  • FIG. 4 is a schematic structural diagram of a speech separation device provided by an embodiment of the present disclosure.
  • Fig. 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
  • the speech separation method provided in the present disclosure may be implemented by the speech separation device provided in the present disclosure.
  • the speech separation device can be realized by any software and/or hardware.
  • the voice separation device may be: a tablet computer, a mobile phone (such as a folding screen mobile phone, a large-screen mobile phone, etc.), a wearable device, a vehicle-mounted device, an augmented reality (augmented reality, AR)/virtual reality (virtual reality, VR) Devices, Notebook PCs, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart TVs, smart screens, HDTVs, 4K TVs, smart speakers, smart projectors Such as Internet of Things (the internet of things, IOT) equipment, this disclosure does not make any restriction on the specific type of electronic equipment.
  • IOT Internet of Things
  • the voice separation method provided by the present disclosure will be described in detail below by taking the method for voice separation performed by an electronic device as an example, combined with several specific embodiments.
  • FIG. 1 is a schematic flowchart of a speech separation method provided by an embodiment of the present disclosure. Referring to Figure 1, the method provided in this embodiment includes:
  • the disclosure does not limit parameters such as duration, storage format, and content of the speech to be processed.
  • the electronic device can acquire speech features corresponding to the speech to be processed, wherein the speech features may include one or more of FBank features (filter bank features), Mel spectrum features, bottleneck features (bottleneck features) and timbre features (pitch features) item.
  • FBank features filter bank features
  • Mel spectrum features Mel spectrum features
  • bottleneck features bottleneck features
  • timbre features timbre features
  • the filters for obtaining FBank features overlap each other, the correlation between the features of each dimension in the FBank features is relatively high, and the FBank feature is used as the speech feature of the speech to be processed, and the speech separation system can use each of the FBank features.
  • the correlation between the features of each dimension can output more accurate speech detection results.
  • the mel spectrum feature is a feature extraction performed in the mel domain, which is closer to the human auditory system and thus can more accurately represent the sound.
  • Pitch feature is a perceptual property that allows to order sounds on a frequency-dependent scale. Pitch can be quantified as a frequency, called the fundamental frequency (F0). Pitch variation forms the tone of a tonal language and is an important feature for speaker recognition and speech recognition.
  • the electronic device can use a speech feature extraction model to perform feature extraction on the speech to be processed.
  • the speech feature extraction model is, for example, an ASR model, or,
  • the electronic device can also use digital processing technology to convert the signal of the speech to be processed, and obtain the speech feature corresponding to the speech to be processed.
  • the bottleneck is a nonlinear feature transformation technique and an effective dimensionality reduction technique.
  • Bottleneck features can include information in dimensions such as prosody and content.
  • the bottleneck feature corresponding to the reference speech is mainly used to distinguish which audio frames in the speech to be processed are the target audio frames to be separated, wherein the electronic device can perform feature extraction on the reference speech based on a separate neural network, and obtain the reference Speech-corresponding bottleneck features.
  • the audio frame included in the speech to be processed matches the reference speech, it means that the audio frame is a target audio frame that needs to be separated. If the audio frame included in the speech to be processed does not match the reference speech, it means that the audio frame is not a target audio frame that needs to be separated.
  • the electronic device may perform voice detection on the voice to be processed through a pre-trained voice separation model, and output a voice detection result.
  • the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech can be used as the input of the speech separation model, and the speech separation model can output a classification label of whether each audio frame of the speech to be processed matches the reference speech, the classification The label is the speech detection result.
  • the bottleneck feature of each audio frame included in the speech to be processed is extracted by using a similar processing flow and method as the bottleneck feature, and then by calculating the bottleneck feature of each audio frame and The similarity between bottleneck features determines the classification result corresponding to the audio frame.
  • the calculation method of the similarity can be but not limited to consine distance, inner product, etc., and then, through the preset distance threshold, compare the distance value corresponding to each audio frame with the preset distance threshold to obtain the corresponding distance of the audio frame classification results.
  • the speech feature corresponding to the speech to be processed is the bottleneck feature corresponding to the speech to be processed.
  • the electronic device may determine a target speech segment matching the reference speech from the speech to be processed according to the speech detection result corresponding to each audio frame.
  • the speech feature corresponding to the target speech segment matches the bottleneck feature corresponding to the reference speech, which means that the matching degree of the timbre of the target speech segment and the timbre of the reference speech meets the requirements.
  • the method provided in this embodiment is to obtain the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech; input the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech into the speech separation model, and obtain the output of the speech separation model The voice detection result; Based on the voice detection result, determine the target voice segment matching the reference voice in the voice to be processed.
  • the system for realizing speech separation is implemented in an end-to-end manner, and by using the speech features corresponding to the speech to be processed and the bottleneck features corresponding to the reference speech as the joint input of the speech separation system, it is used to extract the speech from the speech to be processed Separating the target speech can also improve the accuracy of the separated target speech.
  • Fig. 2 is a schematic structural diagram of a speech separation model provided by an embodiment of the present disclosure.
  • the speech separation model 200 provided by this embodiment includes: a first neural network 201 , a feature fusion layer 202 , a second neural network 203 and a classification layer 204 .
  • the output terminal of the first neural network 201 is connected with the input terminal of the feature fusion layer 202
  • the output terminal of the feature fusion layer 202 is connected with the input terminal of the second neural network 203
  • the output terminal of the second neural network 203 is connected with the classification layer 204 connect.
  • the present disclosure does not limit parameters such as network structures and types of the first neural network 201 and the second neural network 203 .
  • the first neural network 201 and the second neural network 203 can be a feedforward neural network (FeedForward Neural Network), a convolutional neural network (Convlution Neural Network, CNN), a deep feedforward memory network (Deep Feed Forward Sequential Memory Networks, DFSMN), transformer, conformer, etc.
  • the network structure with the best performance can be selected as the first neural network 201 and the second neural network 203 through experiments.
  • the types of the first neural network 201 and the second neural network 203 may be the same or different, which is not limited in the present disclosure.
  • the first neural network 201 is mainly used to receive speech features of the speech to be processed as input, and perform dimensionality reduction and other processing on the speech features to obtain vector expressions corresponding to the speech features.
  • the speech feature is represented as F1, F1 ⁇ R N ⁇ k1 dimension
  • the bottleneck feature corresponding to the reference speech is represented as F2, F2 ⁇ R 1 ⁇ k2 dimension, where N represents the total number of audio frames of the speech to be processed.
  • the method provided by the present disclosure needs to use the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech as joint input, if the speech feature corresponding to the speech to be processed is directly spliced with the bottleneck feature corresponding to the reference speech, due to two If the feature dimensions are different, splicing cannot be performed, and the subsequent speech detection process cannot be performed. Therefore, it is necessary to perform dimensionality reduction and other processing on the speech features corresponding to the speech to be processed.
  • the output of the first neural network 201 and the bottleneck feature corresponding to the reference speech are input to the feature fusion layer 202, so as to splicing the vector expression of the speech feature and the bottleneck feature corresponding to the reference speech to obtain the fusion of the output of the feature fusion layer 202 feature.
  • the present disclosure does not limit the implementation manner of splicing.
  • the fusion feature is used as the input of the second neural network 203, and the second neural network 203 obtains a matrix of the target dimension by spatially mapping the fusion feature, wherein, the matrix output by the second neural network 203 is N*2 dimensions, where N is The frame number of audio frames to process speech.
  • Each 1*2 sub-matrix in the matrix corresponds to 1 audio frame, based on the sub-matrix corresponding to each audio frame.
  • the matrix output by the above-mentioned second neural network 203 is input to the classification layer 204, so that the classification layer uses a preset classification function to calculate according to the matrix, and output the speech detection result of each audio frame.
  • the classification layer 204 uses the softmax function to calculate the sub-matrix corresponding to each audio frame to obtain the probabilities that the audio frames belong to the first category and the second category respectively, and then based on the audio frames belonging to the first category and the second category respectively The maximum value in the probability of determines whether the audio frame matches the reference speech. Wherein, the audio frames included in the first category match the reference speech, and the audio frames included in the second category do not match the reference speech.
  • the speech detection result corresponding to each audio frame output by the classification layer 204 can be expressed by 0, 1 vector, when the speech detection result is 0, it means that the speech feature corresponding to the audio frame does not match the bottleneck feature corresponding to the reference speech, That is, the audio frame does not match the reference speech; when the speech detection result is 1, it means that the speech feature corresponding to the audio frame matches the bottleneck feature corresponding to the reference speech, that is, the audio frame matches the reference speech.
  • a target audio segment matching the reference speech can be determined from the speech to be processed. For example, all audio frames whose speech detection results are 1 in the speech to be processed may be determined as the audio frames included in the target audio segment, and then the target audio segment may be separated.
  • the speech to be processed may include speech segments of multiple voice roles, and each voice role may also correspond to multiple voice segments, and the multiple voice segments are discontinuous, therefore, the target audio segment may include one or more voice segment.
  • the speech separation model provided by the present disclosure can realize speech separation in an end-to-end manner , the training process of the speech separation model provided by the present disclosure is simpler.
  • the speech separation method provided by the present disclosure uses the speech features of the speech to be processed and the bottleneck feature of the reference speech as the joint input of the speech separation model to realize speech separation, which can improve the accuracy of speech separation.
  • Fig. 3 is a schematic flowchart of a speech separation method provided by another embodiment of the present disclosure. Referring to Figure 3, the method provided in this embodiment includes:
  • the sample speech includes the above reference speech, that is, during the training process, it is necessary to ensure that the speech separation model has learned the characteristics of the reference speech, so as to ensure that the trained speech separation model can correctly identify the audio frame that matches the reference speech .
  • the present disclosure does not limit parameters such as the quantity and content of the sample speech and the reference speech.
  • the implementation manner of acquiring the speech features of the sample speech and the bottleneck feature of the sample speech may refer to the relevant description of S101 in the embodiment shown in FIG. 1 .
  • some related databases pre-store the sample speech and the speech features of the sample speech, the bottleneck features corresponding to the sample speech, and the speech detection results of the marked sample speech, etc., it is also possible to separate the speech When the model is training, it is read from the database.
  • a preset loss function can be used to calculate the corresponding loss value, and determine whether the preset convergence condition is met according to the loss value. If the preset convergence condition is met, Then the training is ended, and if the preset convergence condition is not satisfied, the relevant parameters of the speech separation model are optimized based on the loss value. Then, the next round of training is carried out through the sample speech until the preset convergence condition is met, and the trained speech separation model is obtained.
  • the present disclosure does not limit the implementation manner of the preset convergence condition.
  • the preset convergence condition may be an evaluation index such as the number of training iterations and a loss threshold. It should also be noted that the present disclosure does not limit the preset loss function.
  • the speech separation model can be trained by using the relevant data of the non-reference speech in the sample speech, so that the speech separation model has a certain ability of speech separation. On this basis, based on the relevant data of the reference speech, The speech separation model is trained so that the speech separation model has the ability to separate audio frames matching the reference speech.
  • the speech features of the reference speech and the bottleneck features of other sample speeches can also be combined as the input of the speech separation model, so that the speech separation model can learn.
  • Such training can not only enable the speech separation model to learn correct samples, but also learn wrong samples, thereby improving the performance of the speech separation model.
  • the speech detection result of the predicted sample speech of the speech separation model is obtained; based on the speech detection result of the marked sample speech and the speech detection result of the predicted sample speech, the speech separation model is optimized until the preset convergence condition is met , to obtain the trained speech separation model.
  • the present disclosure also provides a voice separation device.
  • Fig. 4 is a schematic structural diagram of a speech separation device provided by an embodiment of the present disclosure.
  • the speech separation device 400 provided in this embodiment includes:
  • the obtaining module 401 is used to obtain the speech feature corresponding to the speech to be processed and the bottleneck feature corresponding to the reference speech.
  • the voice detection module 402 is configured to input the voice feature corresponding to the voice to be processed and the bottleneck feature corresponding to the reference voice into the voice separation model, and obtain a voice detection result output by the voice separation model.
  • a separation module 403 configured to determine a target speech segment in the speech to be processed that matches the reference speech according to the speech detection result, wherein the speech characteristics of the target speech segment are different from the bottleneck characteristics of the reference speech match.
  • the voice detection result is used to indicate whether the voice feature of each audio frame included in the voice to be processed matches the bottleneck feature of the reference voice.
  • the speech detection module 402 is specifically configured to input the speech features corresponding to the speech to be processed into the first neural network included in the speech separation model, and obtain all the speech features output by the first neural network.
  • the vector expression corresponding to the speech feature; the vector expression corresponding to the speech feature and the bottleneck feature corresponding to the reference speech are spliced to obtain the fusion feature; the fusion feature is input to the second neuron included in the speech separation model A network that acquires a matrix output by the second neural network, and acquires the voice detection result based on the matrix.
  • the speech detection module 402 is specifically configured to acquire the probability values that each audio frame belongs to the first category and the second category respectively according to the elements corresponding to each audio frame included in the matrix;
  • the voice features corresponding to the audio frames included in the first category match the bottleneck features corresponding to the reference voice, and the voice features corresponding to the audio frames included in the second category do not match the bottleneck features corresponding to the reference voice; according to
  • Each of the audio frames respectively belongs to the maximum value of the probability values of the first category and the second category, and the speech detection result corresponding to each of the audio frames is determined.
  • the speech detection result corresponding to the audio frame is used to indicate whether the bottleneck feature corresponding to the audio frame matches.
  • the speech separation model is obtained by training based on the speech features corresponding to the sample speech, the bottleneck features corresponding to the sample speech, and the marked speech detection results of the sample speech, and the sample speech Include the reference voice.
  • the speech features include one or more of FBank features, Mel spectrum features, bottleneck features, and timbre features.
  • the voice separation device provided in this embodiment can be used to implement the technical solutions of any of the above method embodiments.
  • the implementation principles and technical effects are similar. You can refer to the detailed description of the foregoing method embodiments. For the sake of brevity, details are not repeated here.
  • the present disclosure also provides an electronic device.
  • FIG. 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
  • an electronic device 500 provided in this embodiment includes: a memory 501 and a processor 502 .
  • the memory 501 may be an independent physical unit, and may be connected with the processor 502 through the bus 503 .
  • the memory 501 and the processor 502 may also be integrated together, implemented by hardware, and the like.
  • the memory 501 is used to store program instructions, and the processor 502 invokes the program instructions to execute the technical solution of any one of the above method embodiments.
  • the foregoing electronic device 500 may also include only the processor 502 .
  • the memory 501 for storing programs is located outside the electronic device 500, and the processor 502 is connected to the memory through circuits/wires, and is used to read and execute the programs stored in the memory.
  • the processor 502 may be a central processing unit (central processing unit, CPU), a network processor (network processor, NP) or a combination of CPU and NP.
  • CPU central processing unit
  • NP network processor
  • the processor 502 may further include a hardware chip.
  • the aforementioned hardware chip may be an application-specific integrated circuit (application-specific integrated circuit, ASIC), a programmable logic device (programmable logic device, PLD) or a combination thereof.
  • the aforementioned PLD may be a complex programmable logic device (complex programmable logic device, CPLD), a field-programmable gate array (field-programmable gate array, FPGA), a general array logic (generic array logic, GAL) or any combination thereof.
  • the memory 501 may include a volatile memory (volatile memory), such as a random-access memory (random-access memory, RAM); the memory may also include a non-volatile memory (non-volatile memory), such as a flash memory (flash memory) ), a hard disk (hard disk drive, HDD) or a solid-state drive (solid-state drive, SSD); the memory can also include a combination of the above-mentioned types of memory.
  • volatile memory such as a random-access memory (random-access memory, RAM
  • non-volatile memory such as a flash memory (flash memory)
  • HDD hard disk drive
  • solid-state drive solid-state drive
  • the present disclosure also provides a readable storage medium, including: computer program instructions; when the computer program instructions are executed by at least one processor of the electronic device, the speech separation method shown in any one of the above method embodiments is implemented.
  • the present disclosure also provides a computer program product.
  • the computer program product When the computer program product is executed by a computer, the computer implements the speech separation method described in any one of the above method embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Quality & Reliability (AREA)
  • Evolutionary Computation (AREA)
  • Business, Economics & Management (AREA)
  • Game Theory and Decision Science (AREA)
  • Telephonic Communication Services (AREA)

Abstract

一种语音分离方法、装置、电子设备及可读存储介质,该方法包括:获取待处理语音对应的语音特征和参考语音对应的瓶颈特征(S101);将待处理语音对应的语音特征和参考语音对应的瓶颈特征输入至语音分离模型,获取语音分离模型输出的语音检测结果(S102);基于语音检测结果,确定待处理语音中与参考语音相匹配的目标语音段,其中,目标语音段的语音特征与参考语音的瓶颈特征相匹配(S103)。该方法将待处理语音对应的语音特征和参考语音对应的瓶颈特征作为语音分离系统的联合输入,从待处理语音中分离出目标语音,能够提高分离出的目标语音的准确度。

Description

语音分离方法、装置、电子设备及可读存储介质
相关申请的交叉引用
本申请是以申请号为202111386550.8,申请日为2021年11月22日的中国申请为基础,并主张其优先权,该中国申请的公开内容在此作为整体引入本申请中。
技术领域
本公开涉及语音处理技术领域,尤其涉及一种语音分离方法、装置、电子设备及可读存储介质。
背景技术
语音活动检测(Voice Activity Detection,VAD)技术被广泛应用于语音识别的前端,用于检测语音与非语音。在一些场景下,不仅需要检测语音与非语音,还需要分离出目标语音。
相关技术中,进行语音分离是通过“VAD系统+语音识别分离系统”实现,具体地,首先,通过VAD系统对语音进行端点检测,再利用语音识别分离系统基于端点检测结果,分离出目标语音。
发明内容
本公开提供了一种语音分离方法、装置、电子设备及可读存储介质。
第一方面,本公开提供了一种语音分离方法,包括:
获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;
将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果;
根据所述语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段的语音特征与所述参考语音的瓶颈特征相匹配。
作为一种可能的实施方式,所述将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果,包括:
将所述待处理语音对应的语音特征输入至所述语音分离模型包括的第一神经 网络,获取所述第一神经网络输出的所述语音特征对应的向量表达;
将所述语音特征对应的向量表达和所述参考语音对应的瓶颈特征进行拼接,获得融合特征;
将所述融合特征输入至所述语音分离模型包括的第二神经网络,获取所述第二神经网络输出的矩阵,且基于所述矩阵获取所述语音检测结果。
作为一种可能的实施方式,所述基于所述矩阵获取所述语音检测结果,包括:
根据所述矩阵包括的各音频帧对应的元素,获取各所述音频帧分别属于第一类别和第二类别的概率值;所述第一类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征匹配,所述第二类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征不匹配;
根据各所述音频帧分别属于第一类别和第二类别的概率值中的最大值,确定各所述音频帧对应的语音检测结果,所述音频帧对应的语音检测结果用于指示所述音频帧对应的语音特征与所述参考语音对应的瓶颈特征是否匹配。
作为一种可能的实施方式,所述语音分离模型是基于样本语音对应的语音特征、所述样本语音对应的瓶颈特征以及标注的所述样本语音的语音检测结果进行训练获得的,所述样本语音包括所述参考语音。
作为一种可能的实施方式,所述语音特征包括FBank特征、梅尔频谱特征以及音色特征中的一项或多项。
第二方面,本公开提供了一种语音分离装置,包括:
获取模块,用于获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;
语音检测模块,用于将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果;
分离模块,用于根据所述语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段的语音特征与所述参考语音的瓶颈特征相匹配。
作为一种可能的实施方式,语音检测模块,具体用于将所述待处理语音对应的语音特征输入至所述语音分离模型包括的第一神经网络,获取所述第一神经网络输出的所述语音特征对应的向量表达;将所述语音特征对应的向量表达和所述 参考语音对应的瓶颈特征进行拼接,获得融合特征;将所述融合特征输入至所述语音分离模型包括的第二神经网络,获取所述第二神经网络输出的矩阵,且基于所述矩阵获取所述语音检测结果。
作为一种可能的实施方式,语音检测模块,具体用于根据所述矩阵包括的各所述音频帧对应的元素,获取各所述音频帧分别属于第一类别和第二类别的概率值;所述第一类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征匹配,所述第二类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征不匹配;根据各所述音频帧分别属于第一类别和第二类别的概率值中的最大值,确定各所述音频帧对应的语音检测结果;
其中,所述音频帧对应的语音检测结果用于指示所述音频帧对应的语音特征与所述参考语音对应的瓶颈特征是否匹配。
作为一种可能的实施方式,所述语音分离模型是基于样本语音对应的语音特征、所述样本语音对应的瓶颈特征以及标注的所述样本语音的语音检测结果进行训练获得的,所述样本语音包括所述参考语音。
作为一种可能的实施方式,所述语音特征包括FBank特征、梅尔频谱特征以及音色特征中的一项或多项。
第三方面,本公开提供了一种电子设备,包括:存储器和处理器;
所述存储器被配置为存储计算机程序指令;
所述处理器被配置为执行所述计算机程序指令,使得所述电子设备实现如第一方面任一项所述的语音分离方法。
第四方面,本公开提供了一种可读存储介质,包括:计算机程序指令;
电子设备的至少一个处理器执行所述计算机程序指令,已实现第一方面任一项所述的语音分离方法。
第五方面,本公开提供一种计算机程序产品,当所述计算机程序产品被计算机执行时,使得所述计算机实现如第一方面任一项所述的语音分离方法。
本公开实施例提供一种语音分离方法、装置、电子设备及可读存储介质,该方法包括:获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型, 获取语音分离模型输出的语音检测结果;基于语音检测结果,确定待处理语音中与参考语音相匹配的目标语音段,其中,所述目标语音段的语音特征与所述参考语音的瓶颈特征相匹配。本公开实施例,用于实现语音分离的系统采用端到端的方式实现,且通过将待处理语音对应的语音特征和参考语音对应的瓶颈特征作为语音分离系统的联合输入,用于从待处理语音中分离出目标语音。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理。
为了更清楚地说明本公开实施例或相关技术中的技术方案,下面将对实施例或相关技术描述中所需要使用的附图作简单地介绍,显而易见地,对于本领域普通技术人员而言,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开一实施例提供的语音分离方法的流程示意图;
图2为本公开一实施例提供的语音分离模型的结构示意图;
图3为本公开另一实施例提供的语音分离方法的流程示意图;
图4为本公开一实施例提供的语音分离装置的结构示意图;
图5为本公开一实施例提供的电子设备的结构示意图。
具体实施方式
为了能够更清楚地理解本公开的上述目的、特征和优点,下面将对本公开的方案进行进一步描述。需要说明的是,在不冲突的情况下,本公开的实施例及实施例中的特征可以相互组合。
在下面的描述中阐述了很多具体细节以便于充分理解本公开,但本公开还可以采用其他不同于在此描述的方式来实施;显然,说明书中的实施例只是本公开的一部分实施例,而不是全部的实施例。
VAD系统的性能以及语音识别分离系统的性能均对语音分离结果的准确度影响极大,因此,采用上述语音分离系统时,VAD系统和语音识别分离系统需要分离训练,较为复杂。
示例性地,本公开提供的语音分离方法可以由本公开提供的语音分离装置实现。语音分离装置可以通过任意的软件和/或硬件的方式实现。示例性地,语音分离装置可以为:平板电脑、手机(如折叠屏手机、大屏手机等)、可穿戴设备、车载设备、增强现实(augmented reality,AR)/虚拟现实(virtual reality,VR)设备、笔记本电脑、超级移动个人计算机(ultra-mobile personal computer,UMPC)、上网本、个人数字助理(personal digital assistant,PDA)、智能电视、智慧屏、高清电视、4K电视、智能音箱、智能投影仪等物联网(the internet of things,IOT)设备,本公开对电子设备的具体类型不作任何限制。
下面以电子设备执行语音分离方法为例,结合几个具体实施例,详细介绍本公开提供的语音分离方法。
图1为本公开一实施例提供的语音分离方法的流程示意图。参照图1所示,本实施例提供的方法包括:
S101、获取待处理语音对应的语音特征和参考语音对应的瓶颈特征。
本公开对于待处理语音的时长、存储格式、内容等等参数不做限定。电子设备能够获取待处理语音对应的语音特征,其中,语音特征可以包括FBank特征(filter bank特征)、梅尔频谱特征、瓶颈特征(bottleneck特征)以及音色特征(pitch特征)中的一项或多项。
由于获取FBank特征的各滤波器之间相互重叠,因此,FBank特征中每个维度的特征之间相关性较高,将FBank特征作为待处理语音的语音特征,语音分离系统能够利用FBank特征中每个维度的特征之间的相关性,输出较为准确的语音检测结果。
梅尔频谱特特征是在mel域进行的特征提取,mel域更接近于人类听觉系统,从而更能够准确地表示声音。
音高特征(pitch特征),音高是一种感知属性,允许在与频率相关的尺度上对声音进行排序。音高可以量化为频率,称为基频(F0)。音高变化形成了声调语言的声调,是重要的说话人识别、语音识别特征。
此外,本公开对于电子设备获取待处理语音对应的语音特征的实现方式不做限定,例如,电子设备可以采用语音特征提取模型对待处理语音进行特征提取, 语音特征提取模型例如为ASR模型,或者,电子设备也可以利用数字处理技术对待处理语音的信号进行转换,获取待处理语音对应的语音特征。
瓶颈(bottleneck)是一种非线性的特征转换技术以及有效的降维技术。瓶颈特征可以包括韵律、内容等维度的信息。本公开中,参考语音对应的瓶颈特征主要用于实现区分待处理语音中哪些音频帧为要分离的目标音频帧,其中,电子设备可以基于一个单独的神经网络对参考语音进行特征提取,获取参考语音对应的瓶颈特征。
S102、将待处理语音对应的语音特征和参考语音对应的瓶颈特征输入至语音分离模型,获取语音分离模型输出的语音检测结果。
其中,语音检测结果
其中,若待处理语音包括的音频帧与参考语音匹配,则表示该音频帧是需要分离出来的目标音频帧。若待处理语音包括的音频帧与参考语音不匹配,则表示该音频帧不是需要分离出来的目标音频帧。
一种可能的实施方式,电子设备可以通过预先训练好的语音分离模型对待处理语音进行语音检测,并输出语音检测结果。具体地,可将待处理语音对应的语音特征和参考语音对应的瓶颈特征作为语音分离模型的输入,语音分离模型可以输出待处理语音的每一个音频帧是否与参考语音匹配的分类标签,该分类标签即为语音检测结果。
另一种可能的实施方式,对于待处理语音,采用与瓶颈特征相似的处理流程和方式,提取到待处理语音包括的每一音频帧的瓶颈特征,再通过计算每一音频帧的瓶颈特征与瓶颈特征之间的相似度,确定音频帧对应的分类结果。其中,相似度的计算方式可以但不限于consine距离,内积等,之后,再通过预预设距离阈值,将每一个音频帧对应的距离值与预设距离阈值进行对比得到该音频帧对应的分类结果。
即,在该实施方式中,待处理语音对应的语音特征即为待处理语音对应的瓶颈特征。
S103、根据语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段对应的语音特征与所述参考语音对应的瓶颈 特征相匹配。
电子设备可以根据每一音频帧对应的语音检测结果,从待处理语音中确定出与参考语音相匹配的目标语音段。其中,目标语音段对应的语音特征与参考语音对应的瓶颈特征相匹配,即表示目标语音段的音色与参考语音的音色匹配度满足要求。
本实施例提供的方法,获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;将待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入语音分离模型,获取语音分离模型输出的语音检测结果;基于语音检测结果,确定待处理语音中与参考语音匹配的目标语音段。本公开实施例,用于实现语音分离的系统采用端到端的方式实现,且通过将待处理语音对应的语音特征和参考语音对应的瓶颈特征作为语音分离系统的联合输入,用于从待处理语音中分离出目标语音,还能够提高分离出的目标语音的准确度。
图2为本公开一实施例提供的语音分离模型的结构示意图。参照图2所示,本实施例提供的语音分离模型200包括:第一神经网络201、特征融合层202、第二神经网络203以及分类层204。
其中,第一神经网络201的输出端与特征融合层202的输入端连接,特征融合层202的输出端与第二神经网络203的输入端连接,第二神经网络203的输出端与分类层204连接。
本公开对于第一神经网络201以及第二神经网络203的网络结构、类型等参数不做限定。例如,第一神经网络201和第二神经网络203可以为前馈神经网络(FeedForward Neural Network)、卷积神经网络(Convlution Neural Network,CNN)、深层前馈记忆网络(Deep Feed Forward Sequential Memory Networks,DFSMN)、transformer、conformer等等。
在实际应用中,可以通过试验选取性能最优的网络结构作为第一神经网络201以及第二神经网络203。此外,还需要说明的是,第一神经网络201和第二神经网络203的类型可以相同,也可以不同,本公开对此不做限定。
第一神经网络201主要用于接收待处理语音的语音特征作为输入,并对语音特征进行降维等处理,获得语音特征对应的向量表达。其中,语音特征表示为F1, F1∈R N×k1维;参考语音对应的瓶颈特征表示为F2,F2∈R 1×k2维,其中,N表示待处理语音的音频帧的总数。
由于本公开提供的方法是需要将待处理语音对应的语音特征与参考语音对应的瓶颈特征作为联合输入,如果直接将待处理语音对应的语音特征与参考语音对应的瓶颈特征进行拼接,由于两个特征维度不同,则无法拼接,也无法进行后续的语音检测流程,因此,需要对待处理语音对应的语音特征进行降维等处理。
之后,将第一神经网络201的输出与参考语音对应的瓶颈特征输入至特征融合层202,以将上述语音特征的向量表达与参考语音对应的瓶颈特征进行拼接,获得特征融合层202输出的融合特征。本公开对于拼接的实现方式不做限定。
融合特征作为第二神经网络203的输入,第二神经网络203通过对融合特征进行空间映射,获得目标维度的矩阵,其中,第二神经网络203输出的矩阵为N*2维,其中,N为待处理语音的音频帧的帧数。
矩阵中每个1*2的子矩阵对应1个音频帧,基于每个音频帧对应的子矩阵。将上述第二神经网络203输出的矩阵输入至分类层204,以使分类层采用预设的分类函数,依据矩阵进行计算,输出每个音频帧的语音检测结果。
示例性地,分类层204采用softmax函数对每个音频帧对应的子矩阵进行计算,获得音频帧分别属于第一类别和第二类别的概率,再基于音频帧分别属于第一类别和第二类别的概率中的最大值确定音频帧是否与参考语音匹配。其中,第一类别包括的音频帧与参考语音匹配,第二类别包括的音频帧与参考语音不匹配。
其中,分类层204输出的每个音频帧对应的语音检测结果可以通过0、1向量表达,当语音检测结果为0,则表示该音频帧对应的语音特征与参考语音对应的瓶颈特征不匹配,即,该音频帧与参考语音不匹配;当语音检测结果为1,则表示该音频帧对应的语音特征与参考语音对应的瓶颈特征匹配,即,该音频帧与参考语音匹配。
之后,可以基于每个音频帧的语音检测结果,从待处理语音中确定出与参考语音匹配的目标音频段。例如,可将待处理语音中所有语音检测结果为1的音频帧确定为目标音频段包括的音频帧,继而可以分离出目标音频段。
还需要说明的是,由于待处理语音可能包括多个语音角色的语音段,每个语 音角色也可能对应多个语音段,且多个语音段不连续,因此,目标音频段可以包括一个或多个语音段。
在图2所示实施例的基础上,可知,与采用“VAD系统+语音识别分离系统”相结合的方式进行语音分离相比,本公开提供的语音分离模型能够通过端到端的方式实现语音分离,本公开提供的语音分离模型的训练过程更加简单。且本公开提供的语音分离方法是将待处理语音的语音特征和参考语音的瓶颈特征作为语音分离模型的联合输入,用于实现语音分离,能够提高语音分离的准确度。
结合前述图1以及图2所示实施例可知,本公开提供的语音分离方法可以通过语音分离模型实现,而语音分离模型是需要预先经过训练的,下面通过图3所示实施例介绍如何训练获取语音分离模型
图3为本公开另一实施例提供的语音分离方法的流程示意图。参照图3所示,本实施例提供的方法包括:
S301、获取样本语音对应的语音特征、样本语音对应的瓶颈特征以及标注的样本语音的语音检测结果。
本方案中,样本语音包括上述参考语音,即在训练的过程中,要保证语音分离模型学习过参考语音的特征,从而保证训练好的语音分离模型能够正确地识别出与参考语音匹配的音频帧。
本公开对于样本语音和参考语音的数量、内容等参数不做限定。
一种可能的实现方式,获取样本语音的语音特征以及样本语音的瓶颈特征的实现方式可以参考前述图1所示实施例中S101的相关描述。
另一种可能的实现方式,若一些相关的数据库中预先存储有样本语音以及样本语音的语音特征、样本语音对应的瓶颈特征以及标注的样本语音的语音检测结果等等,也可以在对语音分离模型进行训练时,从数据库中读取。
S302、将样本语音对应的语音特征和样本语音对应的瓶颈特征,输入至语音分离模型中,获取语音分离模型的预测的样本语音的语音检测结果。
S303、基于标注的样本语音的语音检测结果以及预测的样本语音的语音检测结果,对语音分离模型进行优化,直至满足预设收敛条件,获取训练好的语音分离模型。
可以基于标注的样本语音的语音检测结果以及预测的样本语音的语音检测结果,采用预设的损失函数计算相应的损失值,根据损失值确定是否满足预设收敛条件,若满足预设收敛条件,则结束训练,若不满足预设收敛条件,则基于损失值对语音分离模型的相关参数进行优化。接着,再通过样本语音进行下一轮训练,直至满足预设收敛条件,则获得训练好的语音分离模型。
本公开对于预设收敛条件的实现方式不做限定。示例性地,预设收敛条件可以是训练迭代次数、损失阈值等评价指标。还需要说明的是,本公开对于预设损失函数不做限定。
在训练的过程中,可以先利用样本语音中非参考语音的相关数据对语音分离模型进行训练,使语音分离模型具备一定的语音分离的能力,在此基础上,再基于参考语音的相关数据,对语音分离模型进行训练,使语音分离模型具有能够分离出与参考语音相匹配的音频帧的能力。
在对语音模型进行训练的过程中,还可以将参考语音的语音特征和其他样本语音的瓶颈特征进行组合,作为语音分离模型的输入,以使语音分离模型进行学习。
这样训练不仅能够使语音分离模型学习到正确的样本,还能够学习到错误的样本,从而提高语音分离模型的性能。
本实施例提供的方法,通过获取样本语音对应的语音特征、样本语音对应的瓶颈特征以及标注的样本语音的语音检测结果;将样本语音对应的语音特征和样本语音对应的瓶颈特征,输入至语音分离模型中,获取语音分离模型的预测的样本语音的语音检测结果;基于标注的样本语音的语音检测结果以及预测的样本语音的语音检测结果,对语音分离模型进行优化,直至满足预设收敛条件,获取训练好的语音分离模型。
示例性地,本公开还提供一种语音分离装置。
图4为本公开一实施例提供的语音分离装置的结构示意图。参照图4所示,本实施例提供的语音分离装置400包括:
获取模块401,用于获取待处理语音对应的语音特征和参考语音对应的瓶颈特征.
语音检测模块402,用于将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取语音分离模型输出的语音检测结果。
分离模块403,用于根据所述语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段的语音特征与所述参考语音的瓶颈特征相匹配。
作为一种可能的实施方式,所述语音检测结果用于指示所述待处理语音包括的各音频帧的语音特征是否与所述参考语音的瓶颈特征相匹配。
作为一种可能的实施方式,语音检测模块402,具体用于将所述待处理语音对应的语音特征输入至所述语音分离模型包括的第一神经网络,获取所述第一神经网络输出的所述语音特征对应的向量表达;将所述语音特征对应的向量表达和所述参考语音对应的瓶颈特征进行拼接,获得融合特征;将所述融合特征输入至所述语音分离模型包括的第二神经网络,获取所述第二神经网络输出的矩阵,且基于所述矩阵获取所述语音检测结果。
作为一种可能的实施方式,语音检测模块402,具体用于根据所述矩阵包括的各所述音频帧对应的元素,获取各所述音频帧分别属于第一类别和第二类别的概率值;所述第一类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征匹配,所述第二类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征不匹配;根据各所述音频帧分别属于第一类别和第二类别的概率值中的最大值,确定各所述音频帧对应的语音检测结果。
其中,所述音频帧对应的语音检测结果用于指示该音频帧对应的瓶颈特征是否匹配。
作为一种可能的实施方式,所述语音分离模型是基于样本语音对应的语音特征、所述样本语音对应的瓶颈特征以及标注的所述样本语音的语音检测结果进行训练获得的,所述样本语音包括所述参考语音。
作为一种可能的实施方式,所述语音特征包括FBank特征、梅尔频谱特征、瓶颈特征以及音色特征中的一项或多项。
本实施例提供的语音分离装置可以用于实现如上任一方法实施例的技术方案, 其实现原理以及技术效果类似,可以参照前述方法实施例的详细描述,简明起见,此处不再赘述。
示例性地,本公开还提供一种电子设备。
图5为本公开一实施例提供的电子设备的结构示意图。参照图5所示,本实施例提供的电子设备500包括:存储器501和处理器502。
其中,存储器501可以是独立的物理单元,与处理器502可以通过总线503连接。存储器501、处理器502也可以集成在一起,通过硬件实现等。
存储器501用于存储程序指令,处理器502调用该程序指令,执行以上任一方法实施例的技术方案。
可选地,当上述实施例的方法中的部分或全部通过软件实现时,上述电子设备500也可以只包括处理器502。用于存储程序的存储器501位于电子设备500之外,处理器502通过电路/电线与存储器连接,用于读取并执行存储器中存储的程序。
处理器502可以是中央处理器(central processing unit,CPU),网络处理器(network processor,NP)或者CPU和NP的组合。
处理器502还可以进一步包括硬件芯片。上述硬件芯片可以是专用集成电路(application-specific integrated circuit,ASIC),可编程逻辑器件(programmable logic device,PLD)或其组合。上述PLD可以是复杂可编程逻辑器件(complex programmable logic device,CPLD),现场可编程逻辑门阵列(field-programmable gate array,FPGA),通用阵列逻辑(generic array logic,GAL)或其任意组合。
存储器501可以包括易失性存储器(volatile memory),例如随机存取存储器(random-access memory,RAM);存储器也可以包括非易失性存储器(non-volatile memory),例如快闪存储器(flash memory),硬盘(hard disk drive,HDD)或固态硬盘(solid-state drive,SSD);存储器还可以包括上述种类的存储器的组合。
本公开还提供一种可读存储介质,包括:计算机程序指令;计算机程序指令被电子设备的至少一个处理器执行时,实现上述任一方法实施例所示的语音分离方法。
本公开还提供一种计算机程序产品,所述计算机程序产品被计算机执行时, 使得所述计算机实现上述任一方法实施例所述的语音分离方法。
需要说明的是,在本文中,诸如“第一”和“第二”等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
以上所述仅是本公开的具体实施方式,使本领域技术人员能够理解或实现本公开。对这些实施例的多种修改对本领域的技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开的精神或范围的情况下,在其它实施例中实现。因此,本公开将不会被限制于本文所述的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (10)

  1. 一种语音分离方法,包括:
    获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;
    将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果;
    根据所述语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段对应的语音特征与所述参考语音对应的瓶颈特征相匹配。
  2. 根据权利要求1所述的方法,其中,所述将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果,包括:
    将所述待处理语音对应的语音特征输入至所述语音分离模型包括的第一神经网络,获取所述第一神经网络输出的所述语音特征对应的向量表达;
    将所述语音特征对应的向量表达和所述参考语音对应的瓶颈特征进行拼接,获得融合特征;
    将所述融合特征输入至所述语音分离模型包括的第二神经网络,获取所述第二神经网络输出的矩阵,且基于所述矩阵获取所述语音检测结果。
  3. 根据权利要求2所述的方法,其中,所述基于所述矩阵获取所述语音检测结果,包括:
    根据所述矩阵包括的各音频帧对应的元素,获取各所述音频帧分别属于第一类别和第二类别的概率值;所述第一类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征匹配,所述第二类别包括的音频帧对应的语音特征与所述参考语音对应的瓶颈特征不匹配;
    根据各所述音频帧分别属于第一类别和第二类别的概率值中的最大值,确定各所述音频帧对应的语音检测结果。
  4. 根据权利要求1至3任一项所述的方法,其中,所述语音分离模型是基于样本语音对应的语音特征、所述样本语音对应的瓶颈特征以及标注的所述样本语 音的语音检测结果进行训练获得的,所述样本语音包括所述参考语音。
  5. 根据权利要求1至4任一项所述的方法,其中,所述语音特征包括FBank特征、梅尔频谱特征以及音色特征中的一项或多项。
  6. 一种语音分离装置,包括:
    获取模块,被配置为获取待处理语音对应的语音特征和参考语音对应的瓶颈特征;
    语音检测模块,被配置为将所述待处理语音对应的语音特征和所述参考语音对应的瓶颈特征输入至语音分离模型,获取所述语音分离模型输出的语音检测结果;
    分离模块,被配置为根据所述语音检测结果,确定所述待处理语音中与所述参考语音相匹配的目标语音段,其中,所述目标语音段的语音特征与所述参考语音的瓶颈特征相匹配。
  7. 根据权利要求6所述的装置,其中,所述语音检测模块,进一步被配置为将所述待处理语音对应的语音特征输入至所述语音分离模型包括的第一神经网络,获取所述第一神经网络输出的所述语音特征对应的向量表达;将所述语音特征对应的向量表达和所述参考语音对应的瓶颈特征进行拼接,获得融合特征;将所述融合特征输入至所述语音分离模型包括的第二神经网络,获取所述第二神经网络输出的矩阵,且基于所述矩阵获取所述语音检测结果。
  8. 一种电子设备,包括:存储器和处理器;
    所述存储器被配置为存储计算机程序指令;
    所述处理器被配置为执行所述计算机程序指令,使得所述电子设备实现如权利要求1至5任一项所述的语音分离方法。
  9. 一种可读存储介质,包括:计算机程序指令;
    所述计算机程序指令被电子设备的至少一个处理器执行时,使得所述电子设备实现如权利要求1至5任一项所述的语音分离方法。
  10. 一种计算机程序产品,当所述计算机程序产品被计算机运行时,使得所述计算机实现如权利要求1至5任一项所述的语音分离方法。
PCT/CN2022/129118 2021-11-22 2022-11-02 语音分离方法、装置、电子设备及可读存储介质 Ceased WO2023088091A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/712,467 US20250006215A1 (en) 2021-11-22 2022-11-02 Voice separation method and apparatus, electronic device and readable storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202111386550.8A CN116153326A (zh) 2021-11-22 2021-11-22 语音分离方法、装置、电子设备及可读存储介质
CN202111386550.8 2021-11-22

Publications (1)

Publication Number Publication Date
WO2023088091A1 true WO2023088091A1 (zh) 2023-05-25

Family

ID=86337594

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/129118 Ceased WO2023088091A1 (zh) 2021-11-22 2022-11-02 语音分离方法、装置、电子设备及可读存储介质

Country Status (3)

Country Link
US (1) US20250006215A1 (zh)
CN (1) CN116153326A (zh)
WO (1) WO2023088091A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118248133B (zh) * 2024-05-27 2024-09-20 暗物智能科技(广州)有限公司 二阶段语音识别方法、装置、计算机设备及可读存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108461085A (zh) * 2018-03-13 2018-08-28 南京邮电大学 一种短时语音条件下的说话人识别方法
US20190325880A1 (en) * 2018-04-24 2019-10-24 ID R&D, Inc. System for text-dependent speaker recognition and method thereof
CN112241467A (zh) * 2020-12-18 2021-01-19 北京爱数智慧科技有限公司 一种音频查重的方法和装置
CN112466336A (zh) * 2020-11-19 2021-03-09 平安科技(深圳)有限公司 基于语音的情绪识别方法、装置、设备及存储介质
CN113113044A (zh) * 2021-03-23 2021-07-13 北京小米移动软件有限公司 音频处理方法及装置、终端及存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10650807B2 (en) * 2018-09-18 2020-05-12 Intel Corporation Method and system of neural network keyphrase detection
CN110223680B (zh) * 2019-05-21 2021-06-29 腾讯科技(深圳)有限公司 语音处理方法、识别方法及其装置、系统、电子设备
CN110648656A (zh) * 2019-08-28 2020-01-03 北京达佳互联信息技术有限公司 语音端点检测方法、装置、电子设备及存储介质
US20210272573A1 (en) * 2020-02-29 2021-09-02 Robert Bosch Gmbh System for end-to-end speech separation using squeeze and excitation dilated convolutional neural networks
CN113066506B (zh) * 2021-03-12 2023-01-17 北京百度网讯科技有限公司 音频数据分离方法、装置、电子设备以及存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108461085A (zh) * 2018-03-13 2018-08-28 南京邮电大学 一种短时语音条件下的说话人识别方法
US20190325880A1 (en) * 2018-04-24 2019-10-24 ID R&D, Inc. System for text-dependent speaker recognition and method thereof
CN112466336A (zh) * 2020-11-19 2021-03-09 平安科技(深圳)有限公司 基于语音的情绪识别方法、装置、设备及存储介质
CN112241467A (zh) * 2020-12-18 2021-01-19 北京爱数智慧科技有限公司 一种音频查重的方法和装置
CN113113044A (zh) * 2021-03-23 2021-07-13 北京小米移动软件有限公司 音频处理方法及装置、终端及存储介质

Also Published As

Publication number Publication date
US20250006215A1 (en) 2025-01-02
CN116153326A (zh) 2023-05-23

Similar Documents

Publication Publication Date Title
CN109389971B (zh) 基于语音识别的保险录音质检方法、装置、设备和介质
WO2020253060A1 (zh) 语音识别方法、模型的训练方法、装置、设备及存储介质
WO2020168752A1 (zh) 一种基于对偶学习的语音识别与语音合成方法及装置
CN114678027B (zh) 语音识别结果的纠错方法、装置、终端设备及存储介质
WO2020073509A1 (zh) 一种基于神经网络的语音识别方法、终端设备及介质
CN111816166A (zh) 声音识别方法、装置以及存储指令的计算机可读存储介质
CN108305618B (zh) 语音获取及搜索方法、智能笔、搜索终端及存储介质
CN107229627A (zh) 一种文本处理方法、装置及计算设备
US20250029594A1 (en) Timbre selection method and apparatus, electronic device, readable storage medium, and program product
CN113593606A (zh) 音频识别方法和装置、计算机设备、计算机可读存储介质
CN110796231A (zh) 数据处理方法、装置、计算机设备和存储介质
CN115512692B (zh) 语音识别方法、装置、设备及存储介质
CN111540363A (zh) 关键词模型及解码网络构建方法、检测方法及相关设备
CN111126084B (zh) 数据处理方法、装置、电子设备和存储介质
CN111858936A (zh) 一种意图识别方法、装置、识别设备及可读存储介质
CN112562640A (zh) 多语言语音识别方法、装置、系统及计算机可读存储介质
CN111477212B (zh) 内容识别、模型训练、数据处理方法、系统及设备
WO2023088091A1 (zh) 语音分离方法、装置、电子设备及可读存储介质
CN116072147A (zh) 音乐检测模型训练方法、装置、电子设备及存储介质
CN114625860A (zh) 一种合同条款的识别方法、装置、设备及介质
CN115455142A (zh) 文本检索方法、计算机设备和存储介质
CN111785259A (zh) 信息处理方法、装置及电子设备
CN118865940A (zh) 一种说话人提取方法及系统
WO2024179519A1 (zh) 语义识别方法及其装置
CN111723204B (zh) 语音质检区域的校正方法、装置、校正设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22894628

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 18712467

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 29.08.2024)

122 Ep: pct application non-entry in european phase

Ref document number: 22894628

Country of ref document: EP

Kind code of ref document: A1