WO2020252935A1 - 声纹验证方法、装置、设备及存储介质 - Google Patents
声纹验证方法、装置、设备及存储介质 Download PDFInfo
- Publication number
- WO2020252935A1 WO2020252935A1 PCT/CN2019/103843 CN2019103843W WO2020252935A1 WO 2020252935 A1 WO2020252935 A1 WO 2020252935A1 CN 2019103843 W CN2019103843 W CN 2019103843W WO 2020252935 A1 WO2020252935 A1 WO 2020252935A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- initial
- coverage rate
- phoneme set
- final
- phoneme
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
Definitions
- This application relates to the field of biometrics, and in particular to a voiceprint verification method, device, equipment and storage medium.
- the registered voice of the voiceprint usually allows the user to speak at will, and the speaking time exceeds a certain threshold.
- the speaker's pronunciation characteristics are extracted, and the machine learning method is used to extract one Series feature vector.
- the SNR is required to be above a certain threshold.
- This application provides a voiceprint verification method, device, equipment and storage medium, which provide an important reference for voiceprint identity verification.
- this application provides a voiceprint verification method, which includes:
- the initial coverage rate and the final coverage rate perform voiceprint verification on the voice information to generate a verification result.
- the present application also provides a voiceprint verification device, which includes:
- the text conversion unit is used to convert voice information into text to obtain corresponding text information
- a phoneme acquiring unit configured to acquire a phoneme set corresponding to the text information according to a preset phoneme model, the phoneme set including the initials and finals corresponding to each character in the text information;
- the coverage calculation unit is configured to calculate the coverage rate of the initials of the phoneme set according to the initials table and each initial of the phoneme set; calculate the finals of the phoneme set according to the finals table and each final of the phoneme set Coverage
- the voiceprint verification unit is configured to perform voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate to generate a verification result.
- the present application also provides a computer device, the computer device includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and execute the The computer program implements the above-mentioned voiceprint verification method.
- the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor realizes the voiceprint verification as described above. method.
- the present application discloses a voiceprint verification method, device, equipment and storage medium.
- the voice information is converted into text to obtain the corresponding text information; according to a preset phoneme model, the phoneme set corresponding to the text information is obtained; Calculate the initial coverage rate of the phoneme set according to the initial table and each initial in the phoneme set; calculate the initial coverage rate of the phoneme set according to the final table and each final in the phoneme set; according to the initials
- the coverage rate and the vowel coverage rate are used to perform voiceprint verification on the voice information to generate a verification result, so that it can be known whether the voice information has the voiceprint characteristics of the user’s pronunciation, and whether it can cover most of the characteristics of the user’s voice, Then find out the voice information that covers most of the user's voice features and the user's voice features are highly complete, which provides an important reference for voiceprint authentication.
- FIG. 1 is a schematic flowchart of a voiceprint verification method provided by an embodiment of the present application
- FIG. 2 is a schematic flowchart of sub-steps of the voiceprint verification method in FIG. 1;
- Fig. 3 is a schematic flowchart of steps for obtaining a phoneme set according to an embodiment of the present application
- FIG. 4 is a schematic flowchart of steps for obtaining a phoneme set according to another embodiment of the present application.
- FIG. 5 is a schematic flow chart of the steps for calculating the initial coverage and the final coverage provided by an embodiment of the present application
- FIG. 6 is a schematic flowchart of steps for calculating initial coverage and final coverage provided by another embodiment of the present application.
- FIG. 7 is a schematic flowchart of a voiceprint verification method provided by another embodiment of the present application.
- FIG. 8 is a schematic flowchart of sub-steps of the voiceprint verification method in FIG. 7;
- FIG. 9 is a schematic flowchart of sub-steps of a voiceprint verification method provided by an embodiment of the present application.
- FIG. 10 is a schematic flowchart of sub-steps of a voiceprint verification method provided by another embodiment of the present application.
- FIG. 11 is a schematic flowchart of a voiceprint verification method provided by still another embodiment of the present application.
- FIG. 12 is a schematic block diagram of a voiceprint verification device provided by an embodiment of the present application.
- FIG. 13 is a schematic block diagram of a subunit of the voiceprint verification device in FIG. 12;
- Figure 14 is a schematic block diagram of the sub-modules of the Chinese phoneme acquisition sub-unit of Figure 13;
- FIG. 15 is a schematic block diagram of subunits of the voiceprint verification device in FIG. 12;
- Fig. 16 is a schematic block diagram of a subunit of the voiceprint verification device in Fig. 12;
- FIG. 17 is a schematic block diagram of the structure of a computer device according to an embodiment of the application.
- the embodiments of the present application provide a voiceprint verification method, device, computer equipment, and storage medium.
- the voiceprint verification method can be used to find out the voice information with high completeness of the user's voice feature when registering the user's voiceprint, and provides an important reference for the user's voiceprint identity verification.
- FIG. 1 is a schematic flowchart of steps of a voiceprint verification method provided by an embodiment of the present application.
- the voiceprint verification method specifically includes: step S110 to step S140.
- step S110 specifically includes: uploading the voice information to the cloud platform when it is connected to the external network; receiving the cloud platform according to the voice information The converted text message.
- the voice information is compressed and packaged, and then uploaded to a cloud platform, and the voice information is recognized and converted into text information through the cloud platform.
- the cloud platform refers to a network platform composed of multiple computers for providing voice recognition services.
- the specific process of converting voice information into text namely step S110 specifically includes: when the voice information is not connected to the external network, the voice information is recognized locally and converted into text information.
- the voice information is recognized locally and converted into text information.
- an application program for recognizing voice is installed locally, and a database for recognizing voice is stored.
- the method further includes: receiving the voice information.
- the voice information input by the user is received through an audio input device such as a microphone or a microphone.
- the user can speak freely or read a preset text aloud, and the terminal or server receives the user's voice information through the audio input device. After receiving the voice information, the phoneme set corresponding to the voice information is obtained, and the initial coverage rate and the final coverage rate are directly calculated, so as to perform voiceprint verification on the voice information.
- S120 Acquire a phoneme set corresponding to the text information according to a preset phoneme model.
- the phoneme set includes initials and finals corresponding to each character in the text information.
- the acquiring the phoneme set corresponding to the text information according to the preset phoneme model specifically includes: sub-steps S121, S122, and S123.
- S121 Perform word segmentation processing on the text information to obtain multiple word strings.
- step S121 specifically includes: performing sentence segmentation on the text information to obtain a segmented sentence; performing word segmentation processing on each of the segmented sentences to obtain a word string corresponding to each of the segmented sentences.
- sentence segmentation can be performed on the converted text information. For example, each text can be divided into complete sentences according to punctuation, so as to obtain several segments corresponding to the text information. Sub-statement. Then, perform word segmentation processing on each segmented sentence, thereby obtaining multiple word strings.
- the method for word segmentation processing for each segmentation sentence may adopt a word segmentation method of string matching, such as forward maximum matching method, reverse maximum matching method, shortest path word segmentation, and two-way maximum matching method.
- the forward maximum matching method refers to dividing the character string in a segmented sentence from left to right.
- the reverse maximum matching method is to segment a character string in a segmented sentence from right to left.
- the two-way maximum matching method refers to the forward and reverse (from left to right, right to left) simultaneous word segmentation matching.
- the shortest path segmentation method means that the number of words required to be cut out in a string in a segmented sentence is the least.
- the method for performing word segmentation processing on each segmented sentence may be any other suitable word segmentation method, for example, performing word segmentation processing on each segmented sentence through word meaning segmentation.
- word sense segmentation is a word segmentation method for machine phonetic judgment, which uses syntactic and semantic information to process ambiguity to segment words.
- a Chinese dictionary database with a word set is acquired, and the text information and words in the Chinese dictionary database are traversed, segmented and matched by a two-way maximum matching method, thereby realizing word segmentation of the text information.
- the commonly used words in the Chinese dictionary are sorted by first letter.
- the Chinese dictionary database may be "Modern Chinese Dictionary”.
- the text information S is segmented into a number of segmented sentences.
- the continuous characters of the phrase length m in the segmented sentence are matched with the words in the Chinese dictionary database. If the segmentation sentence fails to match each word in the Chinese dictionary database, the length of consecutive characters is successively reduced for multiple scan matching until the sentence matches a word in the Chinese dictionary database successfully, and finally the text information S Decompose into multiple word strings to obtain word strings FS1, FS2,..., FSN.
- S122 Perform pinyin conversion on each of the word strings to obtain a pinyin string corresponding to each of the word strings.
- N word strings are obtained, which are respectively FS1, FS2, ..., FSN.
- the corresponding pinyin strings of each word string, PS1, PS2,..., PSN are obtained.
- the sub-pinyin string "zhang1san1" is obtained, where the number 1 indicates that the tone is Yinping.
- the method before inputting each of the pinyin strings into a preset phoneme model to obtain a phoneme set, the method further includes: obtaining a standard pronunciation speech library; The Cove model performs model training to establish a phoneme model.
- obtaining the standard pronunciation speech library may specifically include: obtaining a plurality of original recording data and corresponding annotations; performing filtering and correction processing on each of the original recording data and the annotations of each original recording data to obtain the standard pronunciation speech Library.
- the original recording data can be sourced from the Internet, or can be obtained by recording in a recording device such as a voice recorder.
- the original recording data and the corresponding annotations of the original recording data are subjected to multiple rounds of inspection and screening correction processing through automatic or manual methods to obtain standard voice data.
- the collection of standard speech data is constructed as the standard pronunciation speech library.
- the annotation includes tone annotation.
- the standard pronunciation voice library can be obtained directly through the Internet.
- the phoneme model includes an initial sub-model and a final sub-model.
- the process of obtaining the phoneme set, step S123 specifically includes sub-steps S123a, S123b, and S123c.
- each syllable includes a final, and possibly an initial.
- the initials are consonants, and the finals start with a single or diphthong.
- the initials correspond to the initial part of a syllable, and the finals correspond to the final part of a syllable.
- the 23 initials include the 21 initials, w and y in Hanyu Pinyin.
- w and y are not used as initials in the "Hanyu Pinyin Plan", but according to people's customary spelling, w and y are spelled out with initials and vowels, for example, yan is spelled out with initials and vowels, that is y-an-yan, so w and y are also used as initials in this application. Specifically, the 23 initials are shown in Table 1.
- Table 1 is the consonant table of the Chinese dictionary database
- Table 2 is the final table of the Chinese dictionary database
- the first tone also known as Yinping or Ping Tiao
- the second tone also known as Yang Ping or tone
- the third tone also known as Shang Sheng or Zhe Tiao
- the fourth tone also known as falling tone or falling tone
- soft tone soft tone
- S123c Construct a phoneme set according to the initials and vowels corresponding to each character in the word string, and the tones corresponding to each of the vowels.
- the initials corresponding to each character in the word string, the finals corresponding to each character in the word string, and the tones corresponding to each of the finals are constructed as a phoneme set.
- the phoneme model includes a syllable sub-model and a final sub-model.
- the method further includes step S101 of inputting the pinyin string into the syllable sub-model to output the The syllable corresponding to each character in the word string.
- said constructing a phoneme set according to the initials and vowels corresponding to each character in the word string, and the tones corresponding to each of the finals specifically includes: according to the initials, vowels, syllables, and syllables corresponding to each word in the word string.
- the tones corresponding to the finals construct a phoneme set.
- the consonant table can be the consonant table in the "Hanyu Pinyin Plan”
- the final table can be the final table in the "Hanyu Pinyin Plan”.
- the initial coverage rate of the phoneme set is calculated according to the initials table and the initials in the phoneme set; according to the initials table and each initial in the phoneme set, Calculating the final coverage rate of the phoneme set, that is, step S130 specifically includes sub-steps S131, S132, and S133.
- the different initials in the phoneme set are statistically summed to obtain the number of initials corresponding to the phoneme set.
- the different finals in the phoneme set are statistically summed to obtain the number of finals corresponding to the phoneme set.
- the calculation process of the number of initials and finals includes: counting the number of initials appearing in the text information according to the syllables and initials corresponding to the word string; Syllables and vowels, counting the number of vowels appearing in the text information.
- the pinyin of the text message "Zhang San likes to run” is "zhang1san1xi3huan1pao3bu4".
- Table 3 is a display table of initials that do not appear in the pinyin of the text message "Zhang San likes to run"
- Table 4 is a display table of the finals that did not appear in the pinyin of the text message "Zhang San likes to run"
- step S131 calculates the number of initials and the number of finals in the phoneme set, which may also include:
- the de-duplication method of finals and syllables can refer to the de-duplication method of initials, which will not be repeated here.
- the calculating the number of initials and the number of finals in the phoneme set specifically includes: calculating the number of initials and the number of finals in the de-stressed phoneme set.
- the initial coverage rate formula is:
- ⁇ is the coverage rate of the initials
- S is the number of initials
- M is the total number of initials in the initials table in the Chinese dictionary database.
- the pinyin of the text message "Zhang San likes to run” is "zhang1san1xi3huan1pao3bu4".
- the six initials are "zh, s, x, h, d, q”, and the number of initials is 6.
- S133 Based on the vowel coverage rate formula, calculate the vowel coverage rate according to the number of vowels and the vowel table.
- the vowel coverage rate formula is:
- ⁇ is the coverage of the finals
- S is the number of finals
- M is the total number of finals in the Chinese dictionary database.
- the pinyin of the text message "Zhang San likes to run” is "zhang1san1xi3huan1pao3bu4".
- the six vowels are "ang, an, i, uan, ao, u"
- the number of vowels is 6.
- S140 Perform voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate to generate a verification result.
- the verification result may be two types: voiceprint verification passed or voiceprint verification failed.
- the verification result of the voiceprint verification can be considered that the voice information input by the user has the voiceprint characteristics of the user's pronunciation, can cover most of the characteristics of the user's voice, and meet the deep-level requirements of voiceprint registration. If the voiceprint verification fails, it can be considered that the voice information input by the user does not have the voiceprint characteristics of the user's pronunciation, and cannot cover most of the characteristics of the user's voice, so it does not meet the deep-level requirements of voiceprint registration.
- step S140 performs voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate to generate a verification result, before further including step S103,
- the syllable table and each syllable in the phoneme set are used to calculate the syllable coverage rate of the phoneme set.
- the specific process of performing voiceprint verification on the voice information that is, step S140 specifically includes: performing voiceprint verification on the voice information according to the initial coverage rate, the final coverage rate, and the syllable coverage rate to generate Validation results.
- step S103 calculates the syllable coverage of the phoneme set according to the syllable table and each syllable in the phoneme set, including sub-steps S103a and S103b.
- the different syllables in the phoneme set are statistically summed to obtain the number of syllables corresponding to the phoneme set.
- S103b Calculate the syllable coverage rate of the phoneme set according to the syllable table and the number of the syllables.
- the calculation process of the syllable coverage rate specifically includes: based on the syllable coverage rate formula, calculating the syllable coverage rate of the phoneme set according to the syllable table and the number of syllables in the phoneme set, so as to determine the input voice information Whether it can fully reflect the user's voice characteristics provides an important reference.
- the syllable coverage rate formula is:
- ⁇ is the initial coverage
- P is the number of syllables in the phoneme set
- U is the total number in the syllable table.
- the pinyin of the text message "Zhang San likes to run” is "zhang1san1xi3huan1pao3bu4".
- the step S140 performs voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate, specifically including sub-steps S141a, S141b, and S141c.
- S141a Determine whether the initial coverage rate is greater than the initial coverage threshold, and whether the final coverage is greater than the initial coverage threshold.
- the initial coverage threshold and the final coverage threshold can be designed as any appropriate values according to actual application scenarios, for example, the initial coverage threshold is designed to be 50%, and the final coverage threshold is designed to be 30%.
- the initial coverage threshold is 50%
- the final coverage threshold is 30%.
- the initial coverage is 55%
- the final coverage is 32%
- the initial coverage 55% is greater than the initial coverage threshold of 50%
- the final coverage 32% is greater than the initial coverage threshold of 30%.
- the voice information voiceprint verification fails, and the voiceprint verification fails. Passed verification results.
- the initial coverage threshold is 50%
- the final coverage threshold is 30%.
- the initial coverage rate is 48%
- the final coverage rate is 32%. Since the initial initial coverage rate of 48% is less than the initial initial coverage rate threshold of 50%, it is determined that the voiceprint verification of the voice information input by the user has failed, and the voiceprint verification has not been generated. Passed verification results.
- step S140 performing voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate includes: according to the initial coverage rate, the final coverage rate and the final coverage rate.
- the syllable coverage rate is used to perform voiceprint verification on the voice information to generate a verification result.
- step S140 includes sub-steps S142a, S142b, and S142c.
- S142a Determine whether the initial coverage rate is greater than the initial coverage threshold, whether the final coverage is greater than the final coverage threshold, and determine whether the syllable coverage is greater than the syllable coverage threshold.
- the initial coverage threshold, the final coverage threshold, and the syllable coverage threshold can be designed to any appropriate values according to actual application scenarios.
- the initial coverage threshold is designed to be 50%
- the final coverage threshold is designed to 30%
- the syllable coverage threshold is designed to be 0.100%.
- the initial coverage threshold is 50%
- the final coverage threshold is 30%
- the syllable coverage threshold is 0.100%.
- the initial coverage rate is 55%
- the final coverage rate is 32%
- the syllable coverage rate is 0.152%
- the initial coverage rate is 55% greater than the initial coverage threshold of 50%
- the final coverage rate is 32% greater than the final coverage threshold.
- the syllable coverage rate is 0.152% greater than the final coverage rate threshold value of 0.100%.
- the initial coverage rate is not greater than the initial coverage threshold
- the final coverage rate is not greater than the final coverage threshold
- the syllable coverage rate is not greater than the syllable coverage threshold
- the initial coverage threshold is 50%
- the final coverage threshold is 30%
- the syllable coverage threshold is 0.100%.
- the initial coverage rate is 48%
- the final coverage rate is 32%
- the syllable coverage rate is 0.152%. Since the initial coverage rate of 48% is less than the initial coverage threshold of 50%, it is determined that the user input voice information voiceprint verification Failed, a verification result that the voiceprint verification failed is generated.
- the method further includes:
- voiceprint registration application scenario that is, before receiving and storing the input voice information, perform voiceprint verification on the voice information. If the verification result indicates that the voiceprint verification of the voice information has passed, then receive and store the voice information. Voice information is registered voice information, used for voice recognition verification. In this way, the voice information verified by voiceprint can more completely reflect the user's voice characteristics, provide an important reference for subsequent voiceprint registration voice recognition, and improve the security of voiceprint registration.
- the verification voice information and the registered voice information are input into the pre-trained voice recognition model to output the voice recognition result.
- the pre-trained speech recognition model can be obtained by training the initial neural network with a large amount of speech-text sample data.
- the initial neural network can be various neural networks, for example, convolutional neural network, recurrent neural network, long-short-term memory neural network, and so on.
- the prompt message may be "voiceprint verification failed, please re-enter the voice message.” After seeing the prompt message, the user re-enters the voice message until the initial coverage and final coverage meet the requirements.
- the voiceprint verification method described above obtains corresponding text information by converting voice information into text; obtaining a phoneme set corresponding to the text information according to a preset phoneme model; and according to the initials table and each initial in the phoneme set Calculate the initial coverage rate of the phoneme set; calculate the final coverage rate of the phoneme set according to the final table and each final in the phoneme set; calculate the coverage rate of the finals according to the initial coverage rate and the final coverage rate
- Voice information is subjected to voiceprint verification to generate verification results, so that it can be known whether the voice information has the voiceprint characteristics of the user’s pronunciation, whether it can cover most of the characteristics of the user’s voice, and then find out whether the voice
- the voice information with high completeness of voice features provides an important reference for voiceprint identity verification and ensures that the voice information meets the in-depth requirements of voiceprint registration.
- FIG. 12 is a schematic block diagram of a voiceprint verification device according to an embodiment of the present application.
- the voiceprint verification device is used to execute any of the aforementioned voiceprint verification methods.
- the voiceprint verification device can be configured in a server or a terminal.
- the server can be an independent server or a server cluster.
- the terminal can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.
- the voiceprint verification device 200 includes: a text conversion unit 210, a phoneme acquisition unit 220, a coverage calculation unit 230, and a voiceprint verification unit 240.
- the text conversion unit 210 is configured to convert voice information to text to obtain corresponding text information.
- the phoneme obtaining unit 220 is configured to obtain a phoneme set corresponding to the text information according to a preset phoneme model, the phoneme set including the initials and finals corresponding to each character in the text information.
- the coverage calculation unit 230 is configured to calculate the coverage rate of the initials of the phoneme set according to the initials table and the initials in the phoneme set; calculate the initials of the phoneme set according to the finals table and the finals in the phoneme set Final coverage rate.
- the voiceprint verification unit 240 is configured to perform voiceprint verification on the voice information according to the initial coverage rate and the final coverage rate to generate a verification result.
- the phoneme obtaining unit 220 includes a word segmentation processing subunit 221, a pinyin conversion subunit 222 and a phoneme obtaining subunit 223.
- the word segmentation processing subunit 221 is configured to perform word segmentation processing on the text information to obtain multiple word strings.
- the pinyin conversion subunit 222 is used to perform pinyin conversion on each of the word strings to obtain the pinyin string corresponding to each of the word strings.
- the phoneme obtaining subunit 223 is used to input each of the Pinyin strings into a preset phoneme model to obtain a phoneme set.
- the phoneme acquiring subunit 223 includes a consonant output module 223a, a final output module 223c, and a collection construction module 223c.
- the initial consonant output module 223a is configured to input the pinyin string into the initial consonant sub-model to output the initial consonants corresponding to each character in the word string.
- the vowel output module 223c is configured to input the pinyin string into the vowel sub-model to output the vowel corresponding to each character in the word string and the tone corresponding to each vowel.
- the set construction module 223c is used to construct a phoneme set according to the initials and vowels corresponding to each character in the word string, and the tones corresponding to each of the vowels.
- the coverage calculation unit 230 includes a number calculation subunit 231, an initial calculation subunit 232 and a final calculation subunit 233.
- the number calculation subunit 231 is used to calculate the number of initials and the number of finals in the phoneme set.
- the initial initial calculation subunit 232 is configured to calculate the initial initial coverage based on the initial initial coverage formula and the number of the initial initials and the initial initial table.
- the vowel calculation subunit 233 is configured to calculate the vowel coverage rate based on the vowel coverage rate formula and the number of vowels and the vowel table.
- the voiceprint verification device 200 further includes a syllable calculation unit 201, configured to calculate the syllable coverage of the phoneme set according to the syllable table and each syllable in the phoneme set.
- a syllable calculation unit 201 configured to calculate the syllable coverage of the phoneme set according to the syllable table and each syllable in the phoneme set.
- the voiceprint verification unit 240 is configured to perform voiceprint verification on the voice information according to the initial coverage rate, the final coverage rate, and the syllable coverage rate to generate a verification result.
- the voiceprint verification unit 240 includes a coverage judgment subunit 241, a first judgment subunit 242, and a second judgment subunit 243.
- the coverage rate determining subunit 241 is configured to determine whether the initial coverage rate is greater than the initial initial coverage threshold, and whether the final coverage rate is greater than the initial coverage threshold;
- the first determination subunit 242 is configured to determine that the voice information voiceprint verification is passed if the initial coverage rate is greater than the initial coverage threshold, and the final coverage rate is greater than the final coverage threshold;
- the second judging subunit 243 is configured to determine the voice information voiceprint verification if the initial coverage rate is not greater than the initial coverage threshold; or, the final coverage threshold is not greater than the final coverage threshold. Did not pass.
- the voiceprint verification apparatus 200 further includes: an information storage unit 250 and an information generation unit 260.
- the information storage unit 250 is configured to receive and store the voice information if the verification result indicates that the voiceprint verification of the voice information is passed;
- the information generating unit 260 is configured to generate prompt information to prompt the user to re-enter the voice information if the verification result indicates that the voiceprint verification of the voice information has not passed.
- the above-mentioned voiceprint verification device can be implemented in the form of a computer program, and the computer program can be run on a computer device as shown in FIG. 17.
- FIG. 17 is a schematic block diagram of a computer device according to an embodiment of the present application.
- the computer equipment can be a server or a terminal.
- the computer device includes a processor, a memory, and a network interface connected through a system bus, where the memory may include a non-volatile storage medium and an internal memory.
- the non-volatile storage medium can store an operating system and a computer program.
- the computer program includes program instructions.
- the processor can execute a voiceprint verification method.
- the processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
- the internal memory provides an environment for the operation of the computer program in the non-volatile storage medium.
- the processor can execute a voiceprint verification method.
- the network interface is used for network communication, such as sending assigned tasks.
- the network interface is used for network communication, such as sending assigned tasks.
- FIG. 17 is only a block diagram of part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
- the specific computer device may Including more or fewer parts than shown in the figure, or combining some parts, or having a different arrangement of parts.
- the processor may be a central processing unit (Central Processing Unit, CPU), the processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), and application specific integrated circuits (Application Specific Integrated Circuits). Circuit, ASIC), Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc.
- the general-purpose processor may be a microprocessor or the processor may also be any conventional processor.
- the processor is used to run a computer program stored in the memory to implement the following steps:
- Transform voice information into text to obtain corresponding text information obtain a phoneme set corresponding to the text information according to a preset phoneme model, and the phoneme set includes the initials and finals corresponding to each word in the text information
- Calculate the initial coverage rate of the phoneme set according to the initials table and the initials in the phoneme set calculate the initial coverage rate of the phoneme set according to the finals table and the finals in the phoneme set;
- the initial coverage rate and the final coverage rate are used to perform voiceprint verification on the voice information to generate a verification result.
- the embodiments of the present application also provide a computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and the processor executes the program instructions to implement the present application Any of the voiceprint verification methods provided in the embodiments.
- the computer-readable storage medium may be the internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device.
- the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SMC), or a secure digital (Secure Digital, SD) equipped on the computer device. ) Card, Flash Card, etc.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Telephonic Communication Services (AREA)
- Document Processing Apparatus (AREA)
Abstract
一种声纹验证方法、装置、设备及存储介质,该方法包括:将语音信息转化为对应的文本信息;获取文本信息对应的音素集合;根据声母表和音素集合中的各声母,计算音素集合的声母覆盖率;根据韵母表和音素集合中的各韵母,计算音素集合的韵母覆盖率;根据声母覆盖率和韵母覆盖率,对语音信息进行声纹验证,以生成验证结果。
Description
本申请要求于2019年6月17日提交中国专利局、申请号为201910522762.0、发明名称为“声纹验证方法、装置、设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及生物识别领域,尤其涉及一种声纹验证方法、装置、设备及存储介质。
在言语无关的说话人识别系统中,声纹注册的语音通常会让用户随意说话,说话时长超过一定的阈值即可,通过这一段语音,提取说话人的发音特征,使用机器学习的方法提取一系列特征向量。一般对于这段语音,要求信噪比在一定阈值以上。然而,信噪比符合要求的语音难以完整地体现出用户的语音特征。比如,用户在说话的这段时间内,一直重复同一个单词,那么这段语音虽然时长和信噪比都可以达标,但是对于所反映的发音特征是非常有限的。
发明内容
本申请提供了一种声纹验证方法、装置、设备及存储介质,为声纹身份验证提供了重要参考。
第一方面,本申请提供了一种声纹验证方法,所述方法包括:
将语音信息进行文本转化,以得到对应的文本信息;
根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;
根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;
根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
第二方面,本申请还提供了一种声纹验证装置,所述装置包括:
文本转化单元,用于将语音信息进行文本转化,以得到对应的文本信息;
音素获取单元,用于根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;
覆盖率计算单元,用于根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;
声纹验证单元,用于根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
第三方面,本申请还提供了一种计算机设备,所述计算机设备包括存储器和处理器;所述存储器用于存储计算机程序;所述处理器,用于执行所述计算机程序并在执行所述计算机程序时实现如上述的声纹验证方法。
第四方面,本申请还提供了一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序被处理器执行时使所述处理器实现如上述的声纹验证方法。
本申请公开了一种声纹验证方法、装置、设备及存储介质,通过将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果,从而能够知悉该语音信息是否具有用户发音的声纹特征,是否能够涵盖该用户语音的大部分特征,进而找出具有涵盖用户大部分语音特征、用户语音特征完整度高的语音信息,为声纹身份验证提供了重要参考。
为了更清楚地说明本申请实施例技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是本申请的实施例提供的一种声纹验证方法的示意流程图;
图2是图1中的声纹验证方法的子步骤示意流程图;
图3是本申请一实施例提供的获得音素集合的步骤示意流程图;
图4是本申请另一实施例提供的获得音素集合的步骤示意流程图;
图5是本申请一实施例提供的计算声母覆盖率和韵母覆盖率的步骤示意流程图;
图6是本申请另一实施例提供的计算声母覆盖率和韵母覆盖率的步骤示意流程图;
图7是本申请的另一实施例提供的声纹验证方法的示意流程图;
图8是图7中的声纹验证方法的子步骤示意流程图;
图9是本申请一实施例提供的声纹验证方法的子步骤示意流程图;
图10是本申请另一实施例提供的声纹验证方法的子步骤示意流程图;
图11是本申请的再一实施例提供的声纹验证方法的示意流程图。
图12是本申请的实施例提供的声纹验证装置的示意性框图;
图13是图12中声纹验证装置的子单元的示意性框图;
图14是图13中国音素获取子单元的子模块的示意性框图;
图15是图12中声纹验证装置的子单元的示意性框图;
图16是图12中声纹验证装置的子单元的示意性框图;
图17为本申请一实施例提供的一种计算机设备的结构示意性框图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
附图中所示的流程图仅是示例说明,不是必须包括所有的内容和操作/步骤,也不是必须按所描述的顺序执行。例如,有的操作/步骤还可以分解、组合或部分合并,因此实际执行的顺序有可能根据实际情况改变。
本申请的实施例提供了一种声纹验证方法、装置、计算机设备及存储介质。该声纹验证方法可用于针对用户声纹注册时,找出用户语音特征完整度高的语音信息,为用户的声纹身份验证提供了重要的参考。
下面结合附图,对本申请的一些实施方式作详细说明。在不冲突的情况下,下述的实施例及实施例中的特征可以相互组合。
请参阅图1,图1是本申请实施例提供的一种声纹验证方法的步骤示意流程图。
如图1所示,该声纹验证方法,具体包括:步骤S110至步骤S140。
S110、将语音信息进行文本转化,以得到对应的文本信息。
在一实施例中,将语音信息进行文本转化的具体过程,即步骤S110具体包括:当处于连接外网状态时,将所述语音信息上传至云平台;接收所述云平台根据所述语音信息转化后的文本信息。
具体的,将该语音信息进行压缩打包处理,然后上传到云平台,通过云平台对语音信息进行识别转化为文本信息。其中,云平台是指由多台计算机组成的用于提供语音识别服务的网络平台。
在一实施例中,将语音信息进行文本转化的具体过程,即步骤S110具体包括:当处于未连接外网状态时,在本地对所述语音信息进行识别,并转化为文本信息。具体的,在本地安装有对语音进行识别的应用程序,且存储有识别语音的数据库。
在一实施例中,所述将语音信息进行文本转化,以得到对应的文本信息,即步骤S110之前还包括:接收所述语音信息。
具体的,通过麦克风或话筒等音频输入设备接收用户输入的语音信息。
在一实施例中,用户可以随意说话,也可以朗读预设文本,终端或服务器通过音频输入设备接收用户的语音信息。在接收该语音信息后,获取该语音信息对应的音素集合,直接计算声母覆盖率和韵母覆盖率,从而对该语音信息进行声纹验证。
S120、根据预设的音素模型,获取所述文本信息对应的音素集合。
具体的,所述音素集合包括所述文本信息中每个字所对应的声母和韵母。如图2所示,在一实施例中,所述根据预设的音素模型,获取所述文本信息对应的音素集合,具体包括:子步骤S121、S122和S123。
S121、对所述文本信息进行分词处理,以得到多个词串。
具体的,步骤S121具体包括:对所述文本信息进行语句切分,以得到切分语句;对各所述切分语句进行分词处理,以得到各所述切分语句对应的词串。
具体的,对所述语音信息进行文本转化后,可对转化后的文本信息进行语句切分,例如可根据标点符号将各个文本切分成一条条完整的语句,从而得到该文本信息对应的若干切分语句。然后,对各个切分语句进行分词处理,从而得到多个词串。
在一实施例中,对各个切分语句进行分词处理的方法可以采用字符串匹配的分词方法,例如正向最大匹配法、反向最大匹配法、最短路径分词法和双向最大匹配法等。其中,正向最大匹配法是指把一个切分的语句中的字符串从左至右来分词。反向最大匹配法是指把一个切分的语句中的字符串从右至左来分词。双向最大匹配法是指正反向(从左到右、从右到左)同时进行分词匹配。 最短路径分词法是指一个切分的语句中的字符串里面要求切出的词数是最少的。
在其他实施例中,对各个切分语句进行分词处理的方法可以为其他任意合适的分词方法,例如通过词义分词法对各个切分后的语句进行分词处理。其中,词义分词法是一种机器语音判断的分词方法,利用句法信息和语义信息来处理歧义现象来分词。
示例性的,获取具有词语集的汉语词典库,通过双向最大匹配法对文本信息与汉语词典库中的词语进行遍历分割匹配,从而实现对所述文本信息进行分词。其中,汉语词典库中的常用词语按首字母排序。例如,汉语词典库可以为《现代汉语词典》。
具体的,假设汉语词典库的最长词组的长度为m,文本信息S经语句切分后,得到若干切分语句。正反向同时将切分语句中词组长度为m的连续字符与汉语词典库中的词语进行匹配。若切分语句与汉语词典库中的各词语匹配不成功,则逐次减小连续字符的长度进行多次扫描匹配,直至该语句与汉语词典库中的某一词语匹配成功,最终将文本信息S分解为多个词串,即得到词串FS1、FS2、...、FSN。
S122、对各所述词串进行拼音转换,以得到各所述词串对应的拼音串。
示例性的,文本信息S经分词处理后,得到N个词串,分别为FS1、FS2、...、FSN。N个词串分别经拼音转换后,得到各词串对应的拼音串,PS1、PS2、...、PSN。例如,词串“张三”经拼音转换后,得到子拼音串“zhang1san1”,其中数字1表示声调为阴平。
S123、将各所述拼音串输入预设的音素模型,以得到音素集合。
在一实施例中,所述将各所述拼音串输入预设的音素模型,以得到音素集合之前,还包括:获取标准发音语音库;根据所述标准发音语音库,对预设的隐马尔科夫模型进行模型训练,以建立音素模型。
在一实施例中,获取标准发音语音库可以具体包括:获取多个原始录音数据以及对应的标注;对各所述原始录音数据和各原始录音数据的标注进行筛选修正处理,以得到标准发音语音库。
具体的,原始录音数据可以来源于互联网,也可以通过录音设备如录音笔录入获取。通过自动或人工方式对原始录音数据和原始录音数据对应的标注进行多轮检查和筛选修正处理,得到标准语音数据。各标准语音数据的集合构造为所述标准发音语音库。
其中,所述标注包括声调标注。对各所述原始录音数据和各原始录音数据的标注进行筛选修正处理,以得到标准发音语音库,具体可以包括:去除各所述原始录音数据中声调发音模糊的数据;根据汉语词典库,修正所述原始录音数据对应的声调标注。
可以理解的,在其他实施例中,获取标准发音语音库可以通过互联网直接获取。
如图3所示,在一实施例中,所述音素模型包括声母子模型和韵母子模型。音素集合的获得过程,即步骤S123,具体包括子步骤S123a、S123b和S123c。
S123a、将所述拼音串输入所述声母子模型,以输出所述词串中各字对应的声母。
具体的,每个音节包括一个韵母,可能还包括一个声母。声母为辅音,韵母由单元音或双元音开头。声母相应于音节的声母部分,韵母相应于音节的韵母部分。汉语词典库中共有23个声母。23个声母中包括汉语拼音中的21个声 母、w和y。w和y在《汉语拼音方案》中不被作为声母,但根据人们的习惯拼法,会将w和y使用声母拼韵母的方式拼出,比如将yan使用声母拼韵母的方式拼出,即y-an-yan,故本申请中把w和y也作为声母。具体的,23个声母具体如表1所示。
表1 是汉语词典库声母表
| b | p | m | f | d | t |
| n | l | g | k | h | j |
| q | x | zh | ch | sh | r |
| z | c | s | w | y |
S123b、将所述拼音串输入所述韵母子模型,以输出所述词串中各字对应的韵母和各所述韵母对应的声调。
其中,汉语词典库中共有35个声母,如表2所示。
表2 是汉语词典库韵母表
| i | u | ü | |
| a | ia | ua | |
| o | uo | ||
| e | ie | üe | |
| ai | uai | ||
| ei | uei | ||
| ao | iao | ||
| ou | iou | ||
| an | ian | uan | üan |
| en | in | un | ün |
| ang | iang | uang | |
| eng | ing | ueng | |
| ong | iong |
具体的,表2中有一部分韵母,在组成音节时会缩写。比如“iou”,“有”字的拼音写成“you”,有”字的韵母“iou”缩写为“ou”。在一实施例中,在输出韵母时,只考虑表2中出现的韵母,缩写的韵母将会被还原成完整的形式。
其中,汉语词典库中的声调包括五种,分别为第一声(亦称阴平或平调)、第二声(亦称阳平或声调)、第三声(亦称上声或折调)、第四声(亦称去声或降调)、轻声。
S123c、根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合。
具体的,将所述词串中各字对应的声母、所述词串中各字对应的韵母和各所述韵母对应的声调构建为音素集合。
如图4所示,在一实施例中,所述音素模型包括音节子模型和韵母子模型。所述根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合之前,还包括步骤S101、将所述拼音串输入所述音节子模型,以输出所述词串中各字对应的音节。
具体的,汉语拼音中共有潜在的3990个音节(声母和韵母的所有可能组合)。但是并非每个声母、韵母和音调的可能组合都能构成合法音节。实际上只有不含声调的大约416个合法音节,和大约1300多个有意义的带调音节。
其中,所述根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合,具体包括:根据所述词串中各字对应的声母、韵母、音节以及各所述韵母对应的声调,构建音素集合。
S130、根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率。
具体的,声母表可以是《汉语拼音方案》中的声母表,韵母表可以是《汉语拼音方案》中的韵母表。
如图5所示,在一实施例中,所述根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率,即步骤S130具体包括子步骤S131、S132和S133。
S131、计算所述音素集合中声母的数量和韵母的数量。
具体的,将音素集合中互不相同的声母进行统计求和,得到音素集合对应的声母的数量。同样的,将音素集合中互不相同的韵母进行统计求和,得到音素集合对应的韵母的数量。
在一实施例中,声母和韵母的数量的计算过程,即步骤S131包括:根据所述词串对应的音节和声母,统计所述文本信息中出现的声母的数量;根据所述词串对应的音节和韵母,统计所述文本信息中出现的韵母的数量。
比如,文本信息“张三喜欢跑步”的拼音为“zhang1san1xi3huan1pao3bu4”。该文本信息中出现了六个声母和六个韵母,六个声母为“zh、s、x、h、d、q”,六个韵母为“ang、an、i、uan、ao、u”。文本信息“张三喜欢跑步”的拼音中没有出现的声母有17个,具体如表3所示。
表3 是文本信息“张三喜欢跑步”的拼音中未出现的声母展示表
| b | p | m | f | t | n |
| l | g | k | j | ch | sh |
| r | z | c | w | y |
其中,文本信息“张三喜欢跑步”的拼音中没有出现的韵母有29个,具体如表4所示。
表4 是文本信息“张三喜欢跑步”的拼音中未出现的韵母展示表
| ü | |||
| a | ia | ua | |
| o | uo | ||
| e | ie | üe | |
| ai | uai | ||
| ei | uei | ||
| iao | |||
| ou | iou | ||
| ian | üan | ||
| en | in | un | ün |
| iang | uang | ||
| eng | ing | ueng | |
| ong | iong |
如图6所示,在一实施例中,步骤S131计算所述音素集合中声母的数量和韵母的数量,之前还可以包括:
S102、对所述音素集合中的声母、韵母和音节进行去重处理,以得到去重音素集合。
具体的,音素集合中某一声母多次重复出现,将该声母重复的部分舍弃,使得该声母在音素集合中只出现一次。同样的,韵母和音节的去重方法可以参照声母的去重方法,在此不再赘述。所述计算所述音素集合中声母的数量和韵母的数量,具体包括:计算所述去重音素集合中声母的数量和韵母的数量。
S132、基于声母覆盖率公式,根据所述声母的数量和所述声母表,计算所述声母覆盖率。
其中,所述声母覆盖率公式为:
其中,α为所述声母覆盖率,S为声母的数量,M为汉语词典库中声母表的声母的总数量。
比如,文本信息“张三喜欢跑步”的拼音为“zhang1san1xi3huan1pao3bu4”。该文本信息中出现了六个声母,六个声母为“zh、s、x、h、d、q”,声母的数量为6。汉语词典库中声母表的声母的总数量为23,声母覆盖率=6/23=26.09%。
S133、基于韵母覆盖率公式,根据所述韵母的数量和所述韵母表,计算所述韵母覆盖率。
具体的,所述韵母覆盖率公式为:
其中,β为所述韵母覆盖率,S为韵母的数量,M为汉语词典库中韵母表的韵母的总数量。
比如,文本信息“张三喜欢跑步”的拼音为“zhang1san1xi3huan1pao3bu4”。该文本信息中出现了六个韵母,六个韵母为“ang、an、i、uan、ao、u”,韵母的数量为6。汉语词典库中韵母表的韵母的总数量为35,韵母覆盖率=6/35=17.14%。
S140、根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
具体的,验证结果可以为声纹验证通过或声纹验证未通过两种。声纹验证通过的验证结果可认为用户输入的语音信息具有用户发音的声纹特征,能够涵盖该用户语音的大部分特征,符合声纹注册的深层次需求。声纹验证未通过的验证结果可认为用户输入的语音信息不具有用户发音的声纹特征,不能涵盖该用户语音的大部分特征,因而不符合声纹注册的深层次需求。
如图7所示,在一实施例中,步骤S140根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之前,还可以包括步骤S103、根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率。对所述语音信息进行声纹验证的具体过程,即步骤S140具体包括:根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
如图8所示,步骤S103根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率包括子步骤S103a和S103b。
S103a、计算所述音素集合中音节的数量。
具体的,将音素集合中互不相同的音节进行统计求和,得到音素集合对应的音节的数量。
S103b、根据音节表和所述音节的数量,计算所述音素集合的音节覆盖率。
具体的,音节覆盖率的计算过程,具体包括:基于音节覆盖率公式,根据音节表和所述音素集合中音节的数量,计算所述音素集合的音节覆盖率,从而为判断所输入的语音信息是否能够完整地体现用户的语音特征提供了重要参考。
其中,所述音节覆盖率公式为:
其中,γ为所述声母覆盖率,P为音素集合中音节的数量,U为音节表中的总数量。
比如,文本信息“张三喜欢跑步”的拼音为“zhang1san1xi3huan1pao3bu4”。该文本信息中出现了六个音节,分别为“zhang1”、“san1”、“xi3”、“huan1”、“pao3”“bu4”。假设音节表中具有3990个互不相同的音节,则音节覆盖率=6/3990=0.1504%。
如图9所示,在一实施例中,步骤S140所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,具体包括子步骤S141a、S141b和 S141c。
S141a、判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值。
具体的,声母覆盖率阀值和韵母覆盖率阀值可以根据实际应用场景设计为任意适宜的数值,比如声母覆盖率阀值设计为50%、韵母覆盖率阀值设计为30%。
S141b、若所述声母覆盖率大于所述声母覆盖率阀值,且所述韵母覆盖率大于所述韵母覆盖率阀值,判定所述语音信息声纹验证通过。
示例性的,假设声母覆盖率阀值为50%,韵母覆盖率阀值为30%。经计算,声母覆盖率为55%,韵母覆盖率为32%,该声母覆盖率55%大于声母覆盖率阀值50%,且韵母覆盖率32%大于韵母覆盖率阀值30%,此时判定用户输入的语音信息声纹验证通过,生成声纹验证通过的验证结果。
S141c、若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值,判定所述语音信息声纹验证未通过。
具体的,若声母覆盖率不大于声母覆盖率阀值以及所述韵母覆盖率不大于所述韵母覆盖率阀值至少有一个满足条件,判定上述语音信息声纹验证未通过,生成声纹验证未通过的验证结果。
示例性的,假设声母覆盖率阀值为50%,韵母覆盖率阀值为30%。经计算,声母覆盖率为48%,韵母覆盖率为32%,由于声母覆盖率48%小于声母覆盖率阀值50%,因而判定用户输入的语音信息声纹验证未通过,生成声纹验证未通过的验证结果。
在另一实施例中,步骤S140所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,包括:根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
具体的,对所述语音信息进行声纹验证的过程,如图10所示,即步骤S140包括子步骤S142a、S142b和S142c。
S142a、判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值,判断所述音节覆盖率是否大于音节覆盖率阀值。
具体的,声母覆盖率阀值、韵母覆盖率阀值和音节覆盖率阀值可以根据实际应用场景设计为任意适宜的数值,比如声母覆盖率阀值设计为50%、韵母覆盖率阀值设计为30%、音节覆盖率阀值设计为0.100%。
S142b、若所述声母覆盖率大于所述声母覆盖率阀值,所述韵母覆盖率大于所述韵母覆盖率阀值,且所述音节覆盖率大于所述音节覆盖率阀值,判定所述语音信息声纹验证通过。
示例性的,假设声母覆盖率阀值为50%,韵母覆盖率阀值为30%,音节覆盖率阀值为0.100%。经计算,声母覆盖率为55%,韵母覆盖率为32%,音节覆盖率为0.152%,该声母覆盖率55%大于声母覆盖率阀值50%,韵母覆盖率32%大于韵母覆盖率阀值30%,且音节覆盖率为0.152%大于韵母覆盖率阀值0.100%,此时判定用户输入的语音信息声纹验证通过,生成声纹验证通过的验证结果。
S142c、若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值;或,所述音节覆盖率不大于音节覆盖率阀值,判定所述语音信息声纹验证未通过。
具体的,若声母覆盖率不大于声母覆盖率阀值、所述韵母覆盖率不大于所述韵母覆盖率阀值和所述音节覆盖率不大于音节覆盖率阀值至少有一个满足条件,判定上述语音信息声纹验证未通过,生成声纹验证未通过的验证结果。
示例性的,假设声母覆盖率阀值为50%,韵母覆盖率阀值为30%,音节覆盖率阀值为0.100%。经计算,声母覆盖率为48%,韵母覆盖率为32%,音节覆盖率为0.152%,由于该声母覆盖率48%小于声母覆盖率阀值50%,因而判定用户输入的语音信息声纹验证未通过,生成声纹验证未通过的验证结果。
如图11所示,在一实施例中,步骤S140所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之后,还包括:
S150、若所述验证结果表示所述语音信息声纹验证通过,接收并存储所述语音信息。
在声纹注册应用场景中,即在对所输入的语音信息进行接收与存储之前,先对该语音信息进行声纹验证,若验证结果表示所述语音信息声纹验证通过,再接收并存储该语音信息为注册语音信息,用于供语音识别验证。如此通过声纹验证的语音信息能够更加完整地体现出用户的语音特征,为后续的声纹注册语音识别提供了重要参考,提高了声纹注册的安全性。
当用户输入验证语音信息时,将验证语音信息和注册语音信息输入预先训练好的语音识别模型,以输出语音识别结果。
其中,预先训练好的语音识别模型可以是采用大量的语音-文本样本数据对初始神经网络进行训练获得。初始神经网络可以是各种神经网络,例如,卷积神经网络、循环神经网络、长短期记忆神经网络等。
S160、若所述验证结果表示所述语音信息声纹验证未通过,生成提示信息,以提示用户重新输入语音信息。
示例性的,该提示信息可以为“声纹验证失败,请重新输入语音信息”,用户看到该提示信息后,重新输入语音信息,直至声母覆盖率和韵母覆盖率符合要求为止。
上述声纹验证方法,通过将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果,从而能够知悉该语音信息是否具有用户发音的声纹特征,是否能够涵盖该用户语音的大部分特征,进而找出具有涵盖用户大部分语音特征、用户语音特征完整度高的语音信息,为声纹身份验证提供了重要参考,确保该语音信息符合声纹注册的深层次需求。
请参阅图12,图12是本申请的实施例还提供一种声纹验证装置的示意性框图,该声纹验证装置用于执行前述任一项声纹验证方法。其中,该声纹验证装置可以配置于服务器或终端中。
其中,服务器可以为独立的服务器,也可以为服务器集群。该终端可以是手机、平板电脑、笔记本电脑、台式电脑、个人数字助理和穿戴式设备等电子设备。
如图12所示,声纹验证装置200包括:文本转化单元210、音素获取单元220、覆盖率计算单元230和声纹验证单元240。
文本转化单元210,用于将语音信息进行文本转化,以得到对应的文本信息。
音素获取单元220,用于根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母。
覆盖率计算单元230,用于根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述 音素集合的韵母覆盖率。
声纹验证单元240,用于根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
如图13所示,在一实施例中,音素获取单元220包括分词处理子单元221、拼音转化子单元222和音素获取子单元223。
分词处理子单元221,用于对所述文本信息进行分词处理,以得到多个词串。
拼音转化子单元222,用于对各所述词串进行拼音转换,以得到各所述词串对应的拼音串。
音素获取子单元223,用于将各所述拼音串输入预设的音素模型,以得到音素集合。
如图14所示,在一实施例中,音素获取子单元223包括声母输出模块223a、韵母输出模块223c和集合构造模块223c。
声母输出模块223a,用于将所述拼音串输入所述声母子模型,以输出所述词串中各字对应的声母。
韵母输出模块223c,用于将所述拼音串输入所述韵母子模型,以输出所述词串中各字对应的韵母和各所述韵母对应的声调。
集合构造模块223c,用于根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合。
如图15所示,覆盖率计算单元230包括数量计算子单元231、声母计算子单元232和韵母计算子单元233。
数量计算子单元231,用于计算所述音素集合中声母的数量和韵母的数量。
声母计算子单元232,用于基于声母覆盖率公式,根据所述声母的数量和所述声母表,计算所述声母覆盖率。
韵母计算子单元233,用于基于韵母覆盖率公式,根据所述韵母的数量和所述韵母表,计算所述韵母覆盖率。
如图12所示,在一实施例中,声纹验证装置200还包括音节计算单元201,用于根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率。
在该实施中,声纹验证单元240,用于根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
如图16所示,在一实施例中,声纹验证单元240包括覆盖率判断子单元241、第一判定子单元242和第二判定子单元243。
覆盖率判断子单元241,用于判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值;
第一判定子单元242,用于若所述声母覆盖率大于所述声母覆盖率阀值,且所述韵母覆盖率大于所述韵母覆盖率阀值,判定所述语音信息声纹验证通过;
第二判定子单元243,用于若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值,判定所述语音信息声纹验证未通过。
如图12所示,在一实施例中,声纹验证装置200还包括:信息存储单元250和信息生成单元260。
信息存储单元250,用于若所述验证结果表示所述语音信息声纹验证通过,接收并存储所述语音信息;
信息生成单元260,用于若所述验证结果表示所述语音信息声纹验证未通过,生成提示信息,以提示用户重新输入语音信息。
需要说明的是,所属领域的技术人员可以清楚地了解到,为了描述的方便 和简洁,上述描述的声纹验证装置和各单元的具体工作过程,可以参考前述声纹验证方法实施例中的对应过程,在此不再赘述。
上述的声纹验证装置可以实现为一种计算机程序的形式,该计算机程序可以在如图17所示的计算机设备上运行。
请参阅图17,图17是本申请实施例提供的一种计算机设备的示意性框图。该计算机设备可以是服务器或终端。
参阅图17,该计算机设备包括通过系统总线连接的处理器、存储器和网络接口,其中,存储器可以包括非易失性存储介质和内存储器。
非易失性存储介质可存储操作系统和计算机程序。该计算机程序包括程序指令,该程序指令被执行时,可使得处理器执行一种声纹验证方法。
处理器用于提供计算和控制能力,支撑整个计算机设备的运行。
内存储器为非易失性存储介质中的计算机程序的运行提供环境,该计算机程序被处理器执行时,可使得处理器执行一种声纹验证方法。
该网络接口用于进行网络通信,如发送分配的任务等。本领域技术人员可以理解,图17中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定,具体的计算机设备可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
应当理解的是,处理器可以是中央处理单元(Central Processing Unit,CPU),该处理器还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。其中,通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
其中,所述处理器用于运行存储在存储器中的计算机程序,以实现如下步骤:
将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
本申请的实施例中还提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序中包括程序指令,所述处理器执行所述程序指令,实现本申请实施例提供的任一项声纹验证方法。
其中,所述计算机可读存储介质可以是前述实施例所述的计算机设备的内部存储单元,例如所述计算机设备的硬盘或内存。所述计算机可读存储介质也可以是所述计算机设备的外部存储设备,例如所述计算机设备上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到各种等效的修改或替换,这些修改或替换都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以权利要求的保护范围为准。
Claims (20)
- 一种声纹验证方法,包括:将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求1所述的声纹验证方法,其中,所述根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母,包括:对所述文本信息进行分词处理,以得到多个词串;对各所述词串进行拼音转换,以得到各所述词串对应的拼音串;将各所述拼音串输入预设的音素模型,以得到音素集合。
- 根据权利要求2所述的声纹验证方法,其中,所述音素模型包括声母子模型和韵母子模型;所述将各所述拼音串输入预设的音素模型,以得到音素集合,包括:将所述拼音串输入所述声母子模型,以输出所述词串中各字对应的声母;将所述拼音串输入所述韵母子模型,以输出所述词串中各字对应的韵母和各所述韵母对应的声调;根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合。
- 根据权利要求1所述的声纹验证方法,其中,所述根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率,包括:计算所述音素集合中声母的数量和韵母的数量;基于声母覆盖率公式,根据所述声母的数量和所述声母表,计算所述声母覆盖率;基于韵母覆盖率公式,根据所述韵母的数量和所述韵母表,计算所述韵母覆盖率。
- 根据权利要求1所述的声纹验证方法,其中,所述音素集合还包括所述文本信息中每个字所对应的音节;所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之前,还包括:根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率;所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果,包括:根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求1所述的声纹验证方法,其中,所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,包括:判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值;若所述声母覆盖率大于所述声母覆盖率阀值,且所述韵母覆盖率大于所述韵母覆盖率阀值,判定所述语音信息声纹验证通过;若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值,判定所述语音信息声纹验证未通过。
- 根据权利要求1所述的声纹验证方法,其中,所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之后,还包括:若所述验证结果表示所述语音信息声纹验证通过,接收并存储所述语音信息;若所述验证结果表示所述语音信息声纹验证未通过,生成提示信息,以提示用户重新输入语音信息。
- 一种声纹验证装置,包括:文本转化单元,用于将语音信息进行文本转化,以得到对应的文本信息;音素获取单元,用于根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;覆盖率计算单元,用于根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;声纹验证单元,用于根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 一种计算机设备,所述计算机设备包括存储器和处理器;所述存储器用于存储计算机程序;所述处理器,用于执行所述计算机程序并在执行所述计算机程序时,实现如下步骤:将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求9所述的计算机设备,其中,所述处理器在实现所述根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母时,具体实现:对所述文本信息进行分词处理,以得到多个词串;对各所述词串进行拼音转换,以得到各所述词串对应的拼音串;将各所述拼音串输入预设的音素模型,以得到音素集合。
- 根据权利要求10所述的计算机设备,其中,所述音素模型包括声母子模型和韵母子模型;所述处理器在实现所述将各所述拼音串输入预设的音素模型,以得到音素集合时,具体实现:将所述拼音串输入所述声母子模型,以输出所述词串中各字对应的声母;将所述拼音串输入所述韵母子模型,以输出所述词串中各字对应的韵母和各所述韵母对应的声调;根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合。
- 根据权利要求9所述的计算机设备,其中,所述处理器在实现所述根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率时,具体实现:计算所述音素集合中声母的数量和韵母的数量;基于声母覆盖率公式,根据所述声母的数量和所述声母表,计算所述声母覆盖率;基于韵母覆盖率公式,根据所述韵母的数量和所述韵母表,计算所述韵母覆盖率。
- 根据权利要求9所述的计算机设备,其中,所述音素集合还包括所述文本信息中每个字所对应的音节;所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之前,还实现:根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率;所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果时,具体实现:根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求9所述的计算机设备,其中,所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证时,具体实现:判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值;若所述声母覆盖率大于所述声母覆盖率阀值,且所述韵母覆盖率大于所述韵母覆盖率阀值,判定所述语音信息声纹验证通过;若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值,判定所述语音信息声纹验证未通过。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序被处理器执行时使所述处理器实现如下步骤:将语音信息进行文本转化,以得到对应的文本信息;根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母;根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率;根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求15所述的计算机可读存储介质,其中,所述处理器在实现所述根据预设的音素模型,获取所述文本信息对应的音素集合,所述音素集合包括所述文本信息中每个字所对应的声母和韵母时,具体实现:对所述文本信息进行分词处理,以得到多个词串;对各所述词串进行拼音转换,以得到各所述词串对应的拼音串;将各所述拼音串输入预设的音素模型,以得到音素集合。
- 根据权利要求16所述的计算机可读存储介质,其中,所述音素模型包括声母子模型和韵母子模型;所述处理器在实现所述将各所述拼音串输入预设的音素模型,以得到音素集合时,具体实现:将所述拼音串输入所述声母子模型,以输出所述词串中各字对应的声母;将所述拼音串输入所述韵母子模型,以输出所述词串中各字对应的韵母和各所述韵母对应的声调;根据所述词串中各字对应的声母、韵母以及各所述韵母对应的声调,构建音素集合。
- 根据权利要求15所述的计算机可读存储介质,其中,所述处理器在实现所述根据声母表和所述音素集合中的各声母,计算所述音素集合的声母覆盖率;根据韵母表和所述音素集合中的各韵母,计算所述音素集合的韵母覆盖率时,具体实现:计算所述音素集合中声母的数量和韵母的数量;基于声母覆盖率公式,根据所述声母的数量和所述声母表,计算所述声母覆盖率;基于韵母覆盖率公式,根据所述韵母的数量和所述韵母表,计算所述韵母覆盖率。
- 根据权利要求15所述的计算机可读存储介质,其中,所述音素集合还包括所述文本信息中每个字所对应的音节;所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果之前,还实现:根据音节表和所述音素集合中的各音节,计算所述音素集合的音节覆盖率;所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证,以生成验证结果时,具体实现:根据所述声母覆盖率、所述韵母覆盖率和所述音节覆盖率,对所述语音信息进行声纹验证,以生成验证结果。
- 根据权利要求15所述的计算机可读存储介质,其中,所述处理器在实现所述根据所述声母覆盖率和所述韵母覆盖率,对所述语音信息进行声纹验证时,具体实现:判断所述声母覆盖率是否大于声母覆盖率阀值,所述韵母覆盖率是否大于韵母覆盖率阀值;若所述声母覆盖率大于所述声母覆盖率阀值,且所述韵母覆盖率大于所述韵母覆盖率阀值,判定所述语音信息声纹验证通过;若所述声母覆盖率不大于所述声母覆盖率阀值;或,所述韵母覆盖率不大于所述韵母覆盖率阀值,判定所述语音信息声纹验证未通过。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910522762.0 | 2019-06-17 | ||
| CN201910522762.0A CN110335608B (zh) | 2019-06-17 | 2019-06-17 | 声纹验证方法、装置、设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020252935A1 true WO2020252935A1 (zh) | 2020-12-24 |
Family
ID=68142005
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/103843 Ceased WO2020252935A1 (zh) | 2019-06-17 | 2019-08-30 | 声纹验证方法、装置、设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110335608B (zh) |
| WO (1) | WO2020252935A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116312513A (zh) * | 2023-02-13 | 2023-06-23 | 陕西省君凯电子科技有限公司 | 一种智能语音控制系统 |
| CN120877732A (zh) * | 2025-09-22 | 2025-10-31 | 珠海格力电器股份有限公司 | 电器设备控制方法、装置、电器设备、介质及产品 |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110880327B (zh) * | 2019-10-29 | 2024-07-09 | 平安科技(深圳)有限公司 | 一种音频信号处理方法及装置 |
| CN110970035B (zh) * | 2019-12-06 | 2022-10-11 | 广州国音智能科技有限公司 | 单机语音识别方法、装置及计算机可读存储介质 |
| CN111666469B (zh) * | 2020-05-13 | 2023-06-16 | 广州国音智能科技有限公司 | 语句库构建方法、装置、设备和存储介质 |
| CN112669820B (zh) * | 2020-12-16 | 2023-08-04 | 平安科技(深圳)有限公司 | 基于语音识别的考试作弊识别方法、装置及计算机设备 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150073796A1 (en) * | 2013-09-12 | 2015-03-12 | Electronics And Telecommunications Research Institute | Apparatus and method of generating language model for speech recognition |
| CN107016994A (zh) * | 2016-01-27 | 2017-08-04 | 阿里巴巴集团控股有限公司 | 语音识别的方法及装置 |
| CN108989341A (zh) * | 2018-08-21 | 2018-12-11 | 平安科技(深圳)有限公司 | 语音自主注册方法、装置、计算机设备及存储介质 |
| CN109473108A (zh) * | 2018-12-15 | 2019-03-15 | 深圳壹账通智能科技有限公司 | 基于声纹识别的身份验证方法、装置、设备及存储介质 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102779508B (zh) * | 2012-03-31 | 2016-11-09 | 科大讯飞股份有限公司 | 语音库生成设备及其方法、语音合成系统及其方法 |
| CN106057206B (zh) * | 2016-06-01 | 2019-05-03 | 腾讯科技(深圳)有限公司 | 声纹模型训练方法、声纹识别方法及装置 |
| CN109036377A (zh) * | 2018-07-26 | 2018-12-18 | 中国银联股份有限公司 | 一种语音合成方法及装置 |
-
2019
- 2019-06-17 CN CN201910522762.0A patent/CN110335608B/zh active Active
- 2019-08-30 WO PCT/CN2019/103843 patent/WO2020252935A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150073796A1 (en) * | 2013-09-12 | 2015-03-12 | Electronics And Telecommunications Research Institute | Apparatus and method of generating language model for speech recognition |
| CN107016994A (zh) * | 2016-01-27 | 2017-08-04 | 阿里巴巴集团控股有限公司 | 语音识别的方法及装置 |
| CN108989341A (zh) * | 2018-08-21 | 2018-12-11 | 平安科技(深圳)有限公司 | 语音自主注册方法、装置、计算机设备及存储介质 |
| CN109473108A (zh) * | 2018-12-15 | 2019-03-15 | 深圳壹账通智能科技有限公司 | 基于声纹识别的身份验证方法、装置、设备及存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| WANG, LINLIN: "Research on time-varying robustness in speaker recognition", CHINA DOCTORAL DISSERTATIONS FULL-TEXT DATABASE, INFORMATION SCIENCE & TECHNOLOGY, no. 7, 15 July 2014 (2014-07-15), ISSN: 1674-022X, DOI: 20200221133541X * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116312513A (zh) * | 2023-02-13 | 2023-06-23 | 陕西省君凯电子科技有限公司 | 一种智能语音控制系统 |
| CN120877732A (zh) * | 2025-09-22 | 2025-10-31 | 珠海格力电器股份有限公司 | 电器设备控制方法、装置、电器设备、介质及产品 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110335608B (zh) | 2023-11-28 |
| CN110335608A (zh) | 2019-10-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN113811946B (zh) | 数字序列的端到端自动语音识别 | |
| US12536989B2 (en) | Language-agnostic multilingual modeling using effective script normalization | |
| WO2020252935A1 (zh) | 声纹验证方法、装置、设备及存储介质 | |
| US11043213B2 (en) | System and method for detection and correction of incorrectly pronounced words | |
| JP5901001B1 (ja) | 音響言語モデルトレーニングのための方法およびデバイス | |
| JP5932869B2 (ja) | N−gram言語モデルの教師無し学習方法、学習装置、および学習プログラム | |
| CN114830148A (zh) | 可控制有基准的文本生成 | |
| WO2022121185A1 (zh) | 模型训练方法、方言识别方法、装置、服务器及存储介质 | |
| US9858923B2 (en) | Dynamic adaptation of language models and semantic tracking for automatic speech recognition | |
| WO2022121251A1 (zh) | 文本处理模型训练方法、装置、计算机设备和存储介质 | |
| US20140350934A1 (en) | Systems and Methods for Voice Identification | |
| WO2017127296A1 (en) | Analyzing textual data | |
| CN111727442A (zh) | 使用质量分数来训练序列生成神经网络 | |
| CN113178192A (zh) | 语音识别模型的训练方法、装置、设备及存储介质 | |
| JP7664330B2 (ja) | テキストエコー消去 | |
| WO2022086640A1 (en) | Fast emit low-latency streaming asr with sequence-level emission regularization | |
| CN119088912B (zh) | 基于大语言模型的复杂场景对话方法及系统 | |
| JP2017045054A (ja) | 言語モデル改良装置及び方法、音声認識装置及び方法 | |
| CN114372139A (zh) | 数据处理方法、摘要展示方法、装置、设备及存储介质 | |
| WO2022022049A1 (zh) | 文本长难句的压缩方法、装置、计算机设备及存储介质 | |
| US12272351B2 (en) | Automatic out of vocabulary word detection in speech recognition | |
| CN115691503A (zh) | 语音识别方法、装置、电子设备和存储介质 | |
| CN113889145A (zh) | 语音核验方法、装置、电子设备及介质 | |
| CN113096667A (zh) | 一种错别字识别检测方法和系统 | |
| JP2023007014A (ja) | 応答システム、応答方法、および応答プログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19934104 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19934104 Country of ref document: EP Kind code of ref document: A1 |