WO2020151317A1 - 语音验证方法、装置、计算机设备及存储介质 - Google Patents
语音验证方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2020151317A1 WO2020151317A1 PCT/CN2019/117613 CN2019117613W WO2020151317A1 WO 2020151317 A1 WO2020151317 A1 WO 2020151317A1 CN 2019117613 W CN2019117613 W CN 2019117613W WO 2020151317 A1 WO2020151317 A1 WO 2020151317A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- verification
- information
- preset
- voice information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/30—Authentication, i.e. establishing the identity or authorisation of security principals
- G06F21/31—User authentication
- G06F21/32—User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/04—Training, enrolment or model building
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/18—Artificial neural networks; Connectionist approaches
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/22—Interactive procedures; Man-machine interfaces
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
Definitions
- This application relates to the technical field of security verification, in particular to a voice verification method, device, computer equipment and storage medium.
- the traditional voice verification system directly calls the user client after receiving the verification request, and broadcasts the verification information to the user client by means of voice broadcast. After the user obtains the verification information, it returns to the client to fill in and verify the verification information.
- voice verification the user needs to listen to the verification information and record the verification information, and then return to the user client to fill in the verification information. This operation process is too cumbersome; at the same time, the verification information broadcast by voice generally only supports the use of numbers.
- the content of the information has certain limitations, so there is a certain risk of leakage; the traditional voice verification system has the disadvantages of cumbersome operation and high risk of leakage.
- a verification method based on voice recognition is derived.
- the user generates voice content according to the dynamic verification information, and the user's audio content is parsed through the voice recognition algorithm in the background and verified with dynamic verification. The information is compared to verify the accuracy.
- the main function of this technology is to use speech to recognize the semantic content of users, instead of the original manual input mode of verification information, and simplify the verification information steps.
- the effectiveness of the voice recognition verification technology is based on the authenticity of the user, and it is impossible to recognize whether the current voice content is sent by a real human or by a smart AI. After the voice content is deciphered, the smart AI will imitate a human voice. Verification of information voice, so the security of the verification cannot be guaranteed.
- the embodiments of the present application can provide a voice verification method, device, computer equipment, and storage medium that effectively guarantee the authenticity of the verified user and improve the security of the verification system.
- a technical solution adopted in the embodiment created by this application is to provide a voice verification method, which includes the following steps:
- the step of judging whether the voice content is a preset sound category includes the following steps: parsing the verified voice information to obtain characteristic data, wherein the characteristic data is time domain data and spectrum data obtained by processing the voice information; The characteristic data is input into a preset human voice judgment model, where the human voice judgment model is a neural network model that has been trained to convergence and is used to judge whether the voice information is a human voice according to the input characteristic data; The output result of the human voice judgment model determines whether the voice content is a preset voice category.
- an embodiment of the present application also provides a voice verification device, including:
- the obtaining module is used to obtain verification voice information, where the verification voice information is the voice content collected by the target terminal when the verification user reads the verification information aloud; the processing module is used to determine the voice content according to the verification voice information Whether it is a preset sound category, where the preset sound category is a voice category that characterizes that the voice content is a human voice; the execution module is used to determine that the voice content does not belong to the preset voice category, It is determined that the voice verification fails; the first parsing submodule is used to parse the verified voice information to obtain characteristic data, where the characteristic data is time domain data and spectrum data obtained by processing the voice information; the first input submodule uses Inputting the feature data into a preset human voice judgment model, where the human voice judgment model is a neural network model that has been trained to convergence and used to judge whether the voice information is a human voice according to the input feature data; The first processing sub-module is configured to determine whether the voice content is a preset sound category according to the output result of the human voice judgment
- an embodiment of the present application further provides a computer device including a memory and a processor.
- the memory stores computer-readable instructions.
- the processor executes the steps of the voice verification method described above.
- the embodiments of the present application also provide a storage medium storing computer readable instructions.
- the computer readable instructions are executed by one or more processors, the one or more processors execute the above Describe the steps of the voice verification method.
- the beneficial effect of the embodiments of the present application is that compared with the prior art, the technical solutions of the embodiments of the present application focus on mining the user’s biological voice features, which can distinguish the difference between the simulated human voice and the real voice based on This feature can effectively identify real users.
- malicious users such as machines, AI, crawlers, etc. can be effectively excluded, preventing such malicious users from attacking websites and platforms, ensuring the validity and authenticity of verified users, and improving the effectiveness of voice verification. safety.
- Figure 1 is a schematic diagram of the basic flow of a voice verification method according to an embodiment of this application.
- FIG. 2 is a schematic diagram of a flow of obtaining verification information according to an embodiment of the application
- FIG. 3 is a block diagram of the basic structure of a voice verification device according to an embodiment of the application.
- Fig. 4 is a block diagram of the basic structure of a computer device according to an embodiment of the application.
- terminal and “terminal device” used herein may be portable, transportable, installed in a vehicle, or suitable and/or configured to operate locally, and/or Distributed form, running on the earth and/or any other location in space.
- the “terminals” and “terminal devices” used here can also be communication terminals, internet terminals, music/video playback terminals, and other devices.
- FIG. 1 is a schematic diagram of the basic flow of the voice verification method in this embodiment.
- a voice verification method includes the following steps:
- verification voice information is voice content collected by the target terminal when the verification user reads the verification information aloud;
- the verified user After requesting verification, the verified user receives the verification request from the terminal, sends verification information to the terminal, triggers a prompt instruction to guide the user to perform voice verification, and collects the voice verification voice entered by the user.
- the verification information may be one or more words randomly generated, or a random word or a combination of one or more words searched from a preset verification information database.
- the terminal After receiving the verification information, the terminal will verify Information is displayed on the screen and a reminder is issued at the same time. The reminder can be through a specific voice broadcast or display a specific guiding sentence, such as "Please read the verification information on the screen".
- start the voice collection The end point of the collection is determined according to the size of the user's voice. For example, when there is no sound for more than a preset time (such as 1 second, but not limited to this), it is determined that the sound collection is over, and the collected sound is used as verification voice information.
- S1200 Determine whether the voice content is a preset sound category according to the verified voice information, where the preset voice category is a voice category that characterizes that the voice content is a human voice;
- the verification voice information is analyzed to obtain corresponding characteristic data.
- the characteristic data includes, but is not limited to, the time domain data and spectrum data of the verification voice.
- the human voice judgment model is trained to converge. It is a neural network model used to judge whether the voice information is human voice according to the input characteristic data, and judge the output of the model according to the human voice As a result, it is determined whether the verification voice information is a human voice.
- the preset sound category is human voice classification. When the sound category is human voice classification, the voice content belongs to human voice.
- human voice feature data is used as a positive sample
- non-human voice feature data such as voices, animal sounds, and noise synthesized by speech synthesis technology are used as negative samples
- the Inception-v3 neural network The 7x7 convolutional network is decomposed into two one-dimensional convolutions (1x7, 7x1), and the 3x3 convolutional network is also decomposed into two one-dimensional convolutions (1x3, 3x1) to train the Inception-v3 neural network model.
- the human voice judgment model can be set to only two categories, namely, human voice and non-human voice, or more than two categories, such as human voice, synthetic voice, animal voice, and noise, etc., but not limited to this, according to the actual situation Depending on the application scenario, the classification settings can be adjusted appropriately.
- the voice category is human voice
- determine that the voice content belongs to the preset voice category determines that the voice content belongs to the preset voice category
- determine that the voice content does not belong to the preset voice category determines that the voice content does not belong to the preset voice category.
- the current verification The user is an abnormal user and the voice verification fails.
- Step S1200 specifically includes the following steps:
- Step a Parse the verified voice information to obtain characteristic data, where the characteristic data are time domain data and frequency spectrum data obtained by processing the voice information;
- Analyze the obtained verification voice information into original time domain data perform anti-aliasing filtering, sampling, and A/D conversion on the original voice data for digitization, and then perform pre-emphasis to improve the high frequency part and filter out the unimportant To find out the beginning and end of the speech signal, and then perform windowing and framing.
- short-time Fourier transform the processed time-domain data is converted into a frequency signal.
- Mel spectrum transformation the frequency is converted into the linear relationship that human ears can perceive.
- DCT transformation is used to separate the DC signal component and the sinusoidal signal component, and the sound spectrum feature is extracted as the spectrum data, and the time domain data and The spectrum data is collectively used as characteristic data for verifying voice information.
- Step b Input the characteristic data into a preset human voice judgment model, where the human voice judgment model is a neural network that has been trained to convergence and is used to judge whether the voice information is a human voice according to the input characteristic data model;
- human voice feature data is used as a positive sample
- non-human voice feature data such as voices, animal sounds, and noise synthesized by speech synthesis technology are used as negative samples to train the neural network model.
- the neural network model used in this embodiment may be a CNN convolutional neural network model, a VGG convolutional neural network model, or an Inception-v3 neural network model, but is not limited to this.
- the 7x7 convolutional network of the Inception-v3 neural network is decomposed into two one-dimensional convolutions (1x7, 7x1), and the 3x3 convolutional network is also decomposed into two one-dimensional convolutions ( 1x3, 3x1), train the Inception-v3 neural network model.
- the human voice judgment model can be set to only two categories, namely, human voice and non-human voice, or more than two categories, such as human voice, synthetic voice, animal voice, and noise, etc., but not limited to this, according to the actual situation Depending on the application scenario, the classification settings can be adjusted appropriately.
- the feature data is input into the human voice judgment model, and then the output result of the human voice judgment model is obtained.
- Step c Determine whether the voice content is a preset sound category according to the output result of the human voice judgment model
- the preset sound category can be human voice classification.
- the human voice classification is used to characterize the voice content of the human voice. After obtaining the output result of the human voice judgment model, the voice content is determined according to the output result of the human voice judgment model Whether it belongs to the classification of human voice.
- the method of using the human voice judgment model to judge the verification voice can quickly and accurately determine whether the verification voice is a human voice.
- the verification voice of the verification user is obtained, it can be found in time, and when the verification is performed by an abnormal user, it can be verified according to the verification
- the classification result of the voice is intercepted.
- Step a includes the following steps:
- Step a1 Process the verified voice information according to a preset first processing rule to obtain time-domain data, where the first processing rule is to parse the voice information into time-domain data and enhance the high-frequency part therein.
- Voice information processing rules
- Analyze the obtained verification voice information into original time domain data perform anti-aliasing filtering, sampling, and A/D conversion on the original voice data for digitization, and then perform pre-emphasis to improve the high frequency part and filter out the unimportant It also eliminates the effects caused by the vocal cords and lips during the vocalization process to compensate for the high-frequency part of the voice signal suppressed by the pronunciation system, and highlight the high-frequency formant.
- Step a2 Process the time-domain data according to a preset second processing rule to obtain a sound spectrum, where the second processing rule is a data processing rule for converting time-domain data into spectral data according to Fourier transform ;
- the Fourier transform requires the input signal to be stable.
- the voice signal is not stable on the macroscopic level and stable on the microscopic level. It has short-term stability (the voice signal can be considered to be approximately unchanged within 10-30ms).
- the speech signal can be divided into some short segments for processing. Each short segment is called a frame. Since the subsequent operation needs to be windowed, when the frame is divided, the intercepted frame and the frame overlap each other partly, and then intercept The frame is multiplied by the preset window function, so that the original speech signal without periodicity shows part of the characteristics of the periodic function, and then the frame signal is Fourier transformed to obtain the corresponding frequency spectrum.
- the frequency Converting the linear relationship that human ears can perceive, through Mel cepstrum analysis, using DCT transformation to separate the DC signal component and the sinusoidal signal component, and extract the sound spectrum characteristics as spectrum data.
- Step a3 Define the time domain data and the frequency spectrum data as the characteristic data
- the time data and spectrum data obtained by analyzing the verification voice information are used together as the characteristic data of the verification voice information.
- the method of analyzing and processing the verification voice to obtain time-domain data and spectrum data can effectively eliminate the influence of environmental noise and other irrelevant sounds on the verification voice, and simultaneously characterize the verification voice from multiple angles, so that the characteristic data It can reflect the verification voice more truthfully, and the subsequent human voice judgment is more accurate.
- step S1100 the following steps are further included:
- the target terminal When the target terminal needs to perform voice verification, it sends a verification request to the server, and the server side obtains the verification request sent by the terminal.
- a verification database is set in the server.
- the verification database contains a large number of preset texts (for example, 1000).
- the text can be a vocabulary or a random combination of characters.
- a verification request from the target terminal is obtained, a random search is found in the verification database
- the text is used as the verification information for this voice verification.
- multiple words or words can be randomly searched in the verification database and randomly combined to generate verification information, so that the verification information has higher randomness.
- the verification information is sent to the target terminal according to the obtained verification request.
- the verification information is displayed on the screen, and the reminder instruction is triggered at the same time, and the reminder can be sent out. Announce through a specific voice or display a specific guide sentence, such as "Please read the verification information on the screen.”
- the verification information may be preprocessed to obtain the verification information picture, such as obfuscation, but not limited to this, the verification information picture after the preprocessing is displayed to the verification user to guide him/her to perform voice verification.
- step S1200 the following steps are further included:
- Step d When it is judged that the voice content belongs to a preset sound category, verify the voice information according to a preset verification rule, wherein the verification rule is to determine whether the content of the verified voice information is consistent with the verification Whether the similarity of information is greater than the preset similarity threshold data comparison rule;
- the preliminary verification is passed and the voice content is verified.
- Input the verification voice information into the natural language analysis model identify the content in it, output text information corresponding to the voice content, and use the obtained text information as the verification text to compare with the verification information of this voice verification to obtain the comparison
- the obtained similarity is judged whether the similarity is greater than the preset similarity threshold.
- the verification rule is met; when the similarity is not greater than the preset threshold, the verification rule is not met.
- Step e When the verified voice information meets the verification rule, it is determined that the voice verification is passed;
- Step f When the verified voice information does not meet the verification rule, it is determined that the voice verification fails;
- malicious users can prevent malicious users from arbitrarily obtaining permissions and causing damage to the platform or website.
- voice verification can also effectively reduce the possibility of most crawlers or intelligent AI bypassing verification. To improve the authenticity of users.
- Step d specifically includes the following steps:
- Step d1 generating a verification text according to the verification voice information, wherein the verification text is text information corresponding to the content of the verification voice information obtained after content recognition of the verification voice information;
- the voice information is input into the voice recognition model, and the verification text is determined according to the output result of the voice recognition model.
- the verification text is text information corresponding to the content in the voice information, that is, the voice information is converted into text information.
- the speech recognition model used may be an existing one, and a model that generates corresponding text information by recognizing content in the speech information, such as a natural speech analysis model or a neural network model that has been trained to convergence, is not limited here.
- Step d2 Determine text similarity according to the verification text, wherein the text similarity is similarity information between the verification text and the verification information;
- the similarity of the verification text is compared with the verification information to obtain the corresponding text similarity.
- the verification text is converted into Unicode characters or GBK ⁇ GB2312 characters, and compared with the characters of the verification information to determine the Hamming distance.
- the text similarity is determined by the ratio of the Hamming distance to the total number of characters in the verification information.
- each vocabulary or individual Chinese character in the text can be sorted and compared with the vocabulary or Chinese character in the corresponding position in the verification information for the Hamming distance between characters.
- the obtained Hamming distance is greater than zero, it is determined The corresponding vocabulary or Chinese characters do not correspond, the number of vocabularies or Chinese characters that do not correspond between the verification text and the verification information is counted, and the ratio is calculated with the total word volume of the verification information, and the ratio is used as the text similarity.
- Step d3 verify whether the text similarity is greater than the preset similarity threshold
- a similarity threshold is preset in the system to determine whether the similarity between the verification text and the verification information meets the verification rules.
- the value of the similarity threshold can be adjusted according to the actual situation. For example, when a more accurate similarity determination method is selected, you can Increase the value of the similarity threshold. When a rough method of determining the similarity is selected, the value of the similarity threshold can be reduced.
- the comparison result of text similarity and similarity threshold is used to determine whether the voice information meets the verification rule. When the text similarity is greater than the similarity threshold, it is determined that the voice information meets the verification rule and the verification is passed; when the text similarity is less than or equal to the similarity threshold , It is determined that the voice information does not meet the verification rules and the verification fails.
- Step d1 specifically includes the following steps:
- Step d11 Input the verification voice information into a preset voice recognition model, where the voice recognition model is a natural language analysis model that converts the input voice information to obtain text corresponding to the content of the voice information;
- the voice recognition model Enter the voice information into the voice recognition model. First, segment the voice information. The segmentation can be based on pauses in the speech process, or according to the syllable of the speech, the voice information is segmented to obtain the segmented voice, and then The segmented speech is input into a speech recognition model for word segmentation, and fragmented words or syllables are extracted.
- the speech recognition model can be an existing natural language analysis model that converts the input speech information into text.
- Step d12 Determine the verification text according to the output result of the speech recognition model
- the words or syllables output by the speech recognition model are spliced according to the sequence of the segments, and homophones are replaced and adjusted according to the semantics of the entire sentence to obtain a complete sentence as text information.
- Homophone adjustments can be based on preset word collocation relationships, or similarity matching with preset example sentences, and replacements based on words in similar sentences obtained by matching.
- the voice model By using the voice model to extract the content of the voice information and convert it into text, the corresponding text content can be accurately obtained, which is more convenient when compared with the verification information, and the accuracy of the voice verification is determined.
- an embodiment of the present application also provides a voice verification device. Please refer to Figure 3 for details.
- Figure 3 is a block diagram of the basic structure of the voice verification device in this embodiment.
- the voice verification device includes: an acquisition module 2100, a processing module 2200, and an execution module 2300.
- the obtaining module is used to obtain verification voice information, wherein the verification voice information is the voice content collected by the target terminal when the verification user reads the verification information aloud;
- the processing module is used to determine the voice content according to the verification voice information Whether it is a preset sound category, where the preset sound category is a voice category that characterizes that the voice content is human voice;
- the execution module is used to determine when it is determined that the voice content does not belong to the preset sound category Voice verification failed.
- the technical solution of the embodiment of the present application focuses on mining the user's biological voice characteristics, which can distinguish the difference between the simulated human voice and the real human voice based on the characteristics, and can effectively identify real users based on this feature .
- malicious users such as machines, AI, crawlers, etc. can be effectively eliminated, preventing such malicious users from attacking websites and platforms, ensuring the validity and authenticity of verified users, and improving the security of voice verification Sex.
- the voice verification device further includes: a first parsing submodule, a first input submodule, and a first processing submodule.
- the first parsing submodule is used to parse the verified voice information to obtain characteristic data, where the characteristic data is time domain data and spectrum data obtained by processing the voice information;
- the first input submodule is used to combine the characteristic data Input into a preset human voice judgment model, where the human voice judgment model is trained to convergence, and is used to judge whether the voice information is a human voice based on the input feature data;
- the first processing sub-module is used Determining whether the voice content is a preset voice category according to the output result of the human voice judgment model.
- the voice verification device further includes: a second processing submodule, a third processing submodule, and a first execution submodule.
- the second processing sub-module is configured to process the verified voice information according to a preset first processing rule to obtain time-domain data, where the first processing rule is to parse the voice information into time-domain data and improve Among them, the high-frequency part of the voice information processing rules;
- the third processing sub-module is used to process the time domain data according to a preset second processing rule to obtain the sound spectrum, wherein the second processing rule is based on Fu
- the inner leaf transform converts time domain data into data processing rules for spectrum data;
- the first execution submodule is used to define the time domain data and the spectrum data as the characteristic data.
- the voice verification device further includes: a first acquiring submodule, a first searching submodule, and a first sending submodule.
- the first obtaining submodule is used for obtaining the verification request of the target terminal;
- the first searching submodule is used for randomly searching a text as the verification information in a preset verification database according to the verification request;
- the first sending submodule is used for After sending the verification information to the target terminal, a preset reminder instruction is triggered to guide the verification user to perform voice verification according to the verification information.
- the voice verification device further includes: a second execution submodule, a third execution submodule, and a fourth execution submodule.
- the second execution sub-module is configured to verify the voice information according to a preset verification rule when it is judged that the voice content belongs to a preset sound category, wherein the verification rule is judging the verification voice information Whether the similarity between the content of the verification information and the verification information is greater than the preset similarity threshold; the third execution submodule is used to determine that the voice verification is passed when the verification voice information meets the verification rule; fourth The execution sub-module is used to determine that the voice verification fails when the verified voice information does not meet the verification rule.
- the voice verification device further includes: a fourth processing submodule, a fifth processing submodule, and a first verification submodule.
- the fourth processing sub-module is configured to generate a verification text according to the verification voice information, wherein the verification text is a text corresponding to the content of the verification voice information obtained after content recognition of the verification voice information Information; the fifth processing submodule is used to determine the text similarity according to the verification text, wherein the text similarity is the similarity information between the verification text and the verification information; the first verification submodule is used for It is verified whether the text similarity is greater than the preset similarity threshold.
- the voice verification device further includes: a second input submodule and a sixth processing submodule.
- the second input sub-module is used for inputting the verification voice information into a preset voice recognition model, wherein the voice recognition model is converted according to the input voice information to obtain text corresponding to the content of the voice information Natural language analysis model; the sixth processing sub-module is used to determine the verification text according to the output result of the speech recognition model.
- FIG. 4 is a block diagram of the basic structure of the computer device in this embodiment.
- the computer device includes a processor, a nonvolatile storage medium, a memory, and a network interface connected through a system bus.
- the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions.
- the database may store control information sequences.
- the processor can implement a A voice verification method.
- the processor of the computer equipment is used to provide calculation and control capabilities, and supports the operation of the entire computer equipment.
- a computer readable instruction may be stored in the memory of the computer device, and when the computer readable instruction is executed by the processor, the processor may execute a voice verification method.
- the network interface of the computer device is used to connect and communicate with the terminal.
- the structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
- the specific computer device may include More or fewer parts than shown in the figure, or some parts are combined, or have a different arrangement of parts.
- the processor is used to execute the specific functions of the acquisition module 2100, the processing module 2200, and the execution module 2300 in FIG. 3, and the memory stores the program codes and various data required to execute the above modules.
- the network interface is used for data transmission between user terminals or servers.
- the memory in this embodiment stores the program codes and data required to execute all the sub-modules in the voice verification device, and the server can call the program codes and data of the server to execute the functions of all the sub-modules.
- the present application also provides a storage medium storing computer-readable instructions.
- the computer-readable instructions are executed by one or more processors, the one or more processors execute the voice verification method described in any of the above embodiments.
- the storage medium may be a non-volatile readable storage medium.
- the computer program can be stored in a computer readable storage medium. When executed, it may include the procedures of the above-mentioned method embodiments.
- the aforementioned storage medium can be a magnetic disk, an optical disk, a read-only storage memory (Read-Only Memory, ROM) and other non-volatile storage media, or random access memory (Random Access Memory, RAM), etc.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Computer Security & Cryptography (AREA)
- Theoretical Computer Science (AREA)
- Computer Hardware Design (AREA)
- Software Systems (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Telephonic Communication Services (AREA)
Abstract
一种语音验证方法、装置、计算机设备及存储介质,包括下述步骤:获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容(S1100);根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类(S1200);当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败(S1300)。通过对验证语音是否为真实人声进行校验,可以有效排除机器、AI、爬虫等恶意用户,防止此类恶意用户对网站、平台的攻击,保证验证用户有效性和真实性,提升语音验证的安全性。
Description
本申请要求于2019年1月24日提交中国专利局、申请号为201910068827.9、发明名称为“语音验证方法、装置、计算机设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在申请中。
技术领域
本申请涉及安全验证技术领域,特别是涉及一种语音验证方法、装置、计算机设备及存储介质。
背景技术
传统的语音验证系统接收到验证请求后直接呼叫用户客户端,通过语音播报的方式给用户客户端进行播报验证信息,用户获取验证信息后返回客户端进行验证信息的填写并校验。在语音验证的操作过程中,用户需要听取验证信息并记录验证信息信息,然后返回用户客户端进行验证信息的填写,本操作过程过于繁琐;同时,语音播报的验证信息一般只支持使用数字,验证信息的内容具备一定的局限性,因此,存在的一定的泄密风险;传统的语音验证系统存在着操作繁琐、泄密风险高等缺点。
在此基础上衍生出了基于语音识别的验证方式,现有的基于语音识别的验证技术中,用户根据动态验证信息对照生成语音内容,后台通过语音识别算法解析出用户音频内容,并与动态验证信息作对比验证准确性。这一技术的主要功能在于利用语音识别用户语义内容,代替原始的验证信息手动输入模式,简化验证信息步骤。但是该语音识别验证技术的有效性建立在用户真实性的前提下,无法识别出当前的语音内容是否为真实的人类发出还是由智能AI发出,存在语音内容被破译后,由智能AI模仿人类发出验证信息语音,因此无法保证该验证的安全性。
发明内容
本申请实施例能够提供一种有效保证验证用户真实性、提升验证系统安全性的语音验证方法、装置、计算机设备及存储介质。
为解决上述技术问题,本申请创造的实施例采用的一个技术方案是:提供一种语音验证方法,包括以下步骤:
获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类;当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;其中,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤,包括以下步骤:解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
为解决上述技术问题,本申请实施例还提供一种语音验证装置,包括:
获取模块,用于获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;处理模块,用于根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类;执行模块,用于当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;第一解析子模块,用于解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;第一输入子模块,用于将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;第一处理子模块,用于根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
为解决上述技术问题,本申请实施例还提供一种计算机设备,包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被所述处理器执行时,使得所述处理器执行上述所述语音验证方法的步骤。
为解决上述技术问题,本申请实施例还提供一种存储有计算机可读指令的存储介质,所述计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行上述所述语音验证方法的步骤。
本申请实施例的有益效果是:与现有技术相比,本申请实施例的技术方案侧重于挖掘用户的生物学语音特征,此特征可以区分机器声模拟人声和真实人声的差别,基于该特征能够实现有效的鉴别真实用户。通过对验证语音是否为真实人声进行校验,可以有效排除机器、AI、爬虫等恶意用户,防止此类恶意用户对网站、平台的攻击,保证验证用户有效性和真实性,提升语音验证的安全性。
附图说明
图1为本申请实施例语音验证方法的基本流程示意图;
图2为本申请实施例获取验证信息的流程示意图;
图3为本申请实施例语音验证装置的基本结构框图;
图4为本申请实施例计算机设备基本结构框图。
具体实施方式
为了使本技术领域的人员更好地理解本申请方案,下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述。
在本申请的说明书和权利要求书及上述附图中的描述的一些流程中,包含了按照特定顺序出现的多个操作,但是应该清楚了解,这些操作可以不按照其在本文中出现的顺序来执行或并行执行,操作的序号如101、102等,仅仅是用于区分开各个不同的操作,序号本身不代表任何的执行顺序。另外,这些流程可以包括更多或更少的操作,并且这些操作可以按顺序执行或并行执行。需要说明的是,本文中的“第一”、“第二”等描述,是用于区分不同的消息、设备、模块等,不代表先后顺序,也不限定“第一”和“第二”是不同的类型。
本技术领域技术人员可以理解,这里所使用的“终端”、“终端设备”可以是便携式、可运输、安装在交通工具中的,或者适合于和/或配置为在本地运行,和/或以分布形式,运行在地球和/或空间的任何其他位置运行。这里所使用的“终端”、“终端设备”还可以是通信终端、上网终端、音乐/视频播放终端等设备。
具体地请参阅图1,图1为本实施例语音验证方法的基本流程示意图。
如图1所示,一种语音验证方法,包括以下步骤:
S1100、获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;
验证用户在请求验证之后,接收到终端的验证请求,将验证信息发送到终端处,并触发提示指令引导用户进行语音验证,采集用户录入的语验证语音。具体地,验证信息可以是随机生成的一个或多个词汇,或者是从预设的验证信息库中查找得到的随机词汇或者一个或多个文字进行组合,终端在接收到验证信息之后,将验证信息显示在屏幕中,同时发出提醒,提醒的方式可以是通过特定的语音播报或者显示特定的引导句式,例如“请朗读屏幕中的验证信息”,在引导验证用户进行验证之后,启动声音采集,采集的结束点根据用户声音的大小进行判断,例如当超过预设的时间(例如1秒,但不限于此)没有声音时,判断声音采集结束,将采集到的声音作为验证语音信息。
S1200、根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类;
将验证语音信息进行解析,得到对应的特征数据,特征数据包括但不限于验证语音的时域数据和频谱数据。将特征数据输入到预设的人声判断模型中,人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型,根据人声判断模型的输出结果确定验证语音信息是否为人类声音。预设的声音类别为人声分类,当声音类别为人声分类时,即语音内容属于人声。
本实施例使用的人声判断模型在训练时,将人声特征数据作为正样本,语音合成技术合成的声音、动物声和杂音等非人声特征数据作为负样本,将Inception-v3神经网络的7x7卷积网络分解成两个一维的卷积(1x7,7x1),3x3卷积网络也分解成两个一维的卷积(1x3,3x1),训练Inception-v3神经网络模型。人声判断模型可以仅设置两种分类,即属于人声和不属于人声,或者可以设置超过两种的分类,例如人声、合成声、动物声和杂音等,但不限于此,根据实际应用场景的不同,分类的设置可以适当进行调整。
S1300、当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;
根据人声判断模型的输出结果确定语音内容的分类,当声音类别为人类声音时,确定语音内容属于预设的声音类别,当声音类别不属于人类声音时,确定语音内容不属于预设的声音类别。当判断语音内容不属于预设的声音类别,即语音内容的声音类别不是人类声音时,此时可能是由智能AI或者爬虫等破译了验证信息之后利用模拟语音等方式企图绕过验证,当前验证用户为非正常用户,语音验证失败。
步骤S1200具体包括以下步骤:
步骤a、解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;
将获取到的验证语音信息解析成原始的时域数据,对原始的声音数据进行反混叠滤波、采样、A/D转换进行数字化,之后进行预加重,提升高频部分,滤掉其中不重要的信息以及背景噪音,并进行语音信号的端点检测,从而找出语音信号的始末,然后进行加窗分帧,通过短时傅里叶变换,将处理之后的时域数据转换为频哉信号,通过梅尔频谱变换,将频率转换成人耳能感知的线性关系,通过梅尔倒谱分析,采用DCT变换将直流信号分量和正弦信号分量分离,提取声音频谱特征作为频谱数据,将时域数据和频谱数据共同作为验证语音信息的特征数据。
步骤b、将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;
本实施例使用的人声判断模型在训练时,将人声特征数据作为正样本,语音合成技术合成的声音、动物声和杂音等非人声特征数据作为负样本,对神经网络模型进行训练。本实施例使用的神经网络模型可以是CNN卷积神经网络模型、VGG卷积神经网络模型或者Inception-v3神经网络模型,但不限于此。以Inception-v3神经网络为例,将Inception-v3神经网络的7x7卷积网络分解成两个一维的卷积(1x7,7x1),3x3卷积网络也分解成两个一维的卷积(1x3,3x1),训练Inception-v3神经网络模型。人声判断模型可以仅设置两种分类,即属于人声和不属于人声,或者可以设置超过两种的分类,例如人声、合成声、动物声和杂音等,但不限于此,根据实际应用场景的不同,分类的设置可以适当进行调整。
在确定了验证语音信息的特征数据之后,将特征数据输入到人声判断模型当中,然后获取人声判断模型的输出结果。
步骤c、根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别;
预设的声音类别可以是人声分类,人声分类即用于表征语音内容属于人类发声的声音类别,在获取到人声判断模型的输出结果之后,根据人声判断模型的输出结果确定语音内容是否属于人声分类。
利用人声判断模型对验证语音进行判断的方法,可以快速并准确地判断验证语音是否属于人声,当获取到验证用户的验证语音存在异常时可以及时发现,在非正常用户进行验证时根据验证语音的分类结果进行拦截。
步骤a包括以下步骤:
步骤a1、根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;
将获取到的验证语音信息解析成原始的时域数据,对原始的声音数据进行反混叠滤波、采样、A/D转换进行数字化,之后进行预加重,提升高频部分,滤掉其中不重要的信息以及背景噪音,同时消除发声过程中声带和嘴唇造成的效应,来补偿语音信号受到发音系统所压抑的高频部分,并且突显高频的共振峰。
步骤a2、根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;
进行语音信号的端点检测,找出语音信号的始末,然后进行加窗分帧。傅里叶变换要求输入的信号的平稳的,语音信号在宏观上是不平稳的,在微观上是平稳的,具有短时平稳性(10-30ms内可以认为语音信号近似不变),这个就可以把语音信号分为一些短段来进行处理,每一个短段称为一帧,由于后续操作需要加窗,则在分帧的时候,截取的帧与帧之间相互重叠一部分,然后将截取的帧与预设的窗函数相乘,使原本没有周期性的语音信号呈现出周期函数的部分特征,然后对帧信号进行傅里叶变换,得到对应的频谱,通过梅尔频谱变换,将频率转换成人耳能感知的线性关系,通过梅尔倒谱分析,采用DCT变换将直流信号分量和正弦信号分量分离,提取声音频谱特征作为频谱数据。
步骤a3、定义所述时域数据和所述频谱数据为所述特征数据;
将对验证语音信息进行解析得到的时哉数据和频谱数据共同作为验证语音信息的特征数据。
通过对验证语音进行解析并处理得到时域数据和频谱数据的方法,可以有效地消除环境杂等不相关声音对验证语音的影响,并且同时从多个角度去表征验证语音的特征,使特征数据可以更加真实地反映验证语音,例后续的人声判断更加准确。
如图2所示,步骤S1100之前还包括以下步骤:
S1010、获取目标终端的验证请求;
目标终端需要进行语音验证时,向服务器发送验证请求,服务器端获取终端发送的验证请求。
S1020、根据所述验证请求在预设的验证数据库随机查找一个文本作为所述验证信息;
服务器中设置有验证数据库,验证数据库中包含有预设的大量文本(例如1000个),文本可以是词汇或者随机的文字组合,在获取到目标终端的验证请求时,在验证数据库中随机查找一个文本作为本次语音验证的验证信息。在一些实施方式中,可以在验证数据库中随机查找多个文字或词汇进行随机组合生成验证信息,以使验证信息具备更高的随机性。
S1030、将所述验证信息发送至目标终端,触发预设的提醒指令,以引导验证用户根据所述验证信息进行语音验证;
当查找得到验证信息后,根据获取到的验证请求将验证信息发送到目标终端,终端在接收到验证信息之后,将验证信息显示在屏幕中,同时触发提醒指令,发出提醒,提醒的方式可以是通过特定的语音播报或者显示特定的引导句式,例如“请朗读屏幕中的验证信息”。在一些实施方式中,显示验证信息之前可以对验证信息进行预处理得到验证信息图片,例如模糊化,但不限于此,将预处理之后的验证信息图片展示给验证用户,引导其进行语音验证。
步骤S1200之后还包括下述步骤:
步骤d、当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证,其中,所述验证规则为判断所述验证语音信息的内容与所述验证信息的相似度是否大于预设的相似度阈值的数据对比规则;
判断语音内容属于预设的声音类别时,初步验证通过,对语音内容进行验证。将验证语音信息输入到自然语言解析模型中,识别其中的内容,输出与语音内容相对应的文本信息,将获取到的文本信息作为验证文本,与本次语音验证的验证信息进行对比,获取对比得到的相似度,判断相似度是否大于预设的相似度阈值,当相似度大于预设的阈值时,即符合验证规则,当相似度不大于预设的阈值时,不符合验证规则。
步骤e、当所述验证语音信息符合所述验证规则时,确定语音验证通过;
当提取得到的验证文本与验证信息的相似度大于预设的相似度阈值时,确定验证语音信息符合验证规则,语音验证通过。
步骤f、当所述验证语音信息不符合所述验证规则时,确定语音验证失败;
当提取得到的验证文本与验证信息的相似度小于或等于预设的相似度阈值时,确定验证语音信息不符合验证规则,语音验证失败。
通过建立验证规则,利用验证规则对用户进行验证的方式,防止恶意用户随意获得权限而对平台或网站造成破坏,利用语音验证的方式也可以有效的减少大部分爬虫或者智能AI绕过验证的可能性,提高用户的真实性。
步骤d具体包括下述步骤:
步骤d1、根据所述验证语音信息生成验证文本,其中,所述验证文本为对所述验证语音信息进行内容识别后得到的与所述验证语音信息的内容相对应的文本信息;
将语音信息输入到语音识别模型中,根据语音识别模型的输出结果确定验证文本,验证文本为与语音信息中的内容相对应的文本信息,即把语音信息转换为文本信息,本实施例中所使用的语音识别模型可以是现有的,通过识别语音信息中的内容生成对应的文本信息的模型,例如自然语音解析模型或者已经训练至收敛的神经网络模型,在此不作限定。
步骤d2、根据所述验证文本确定文本相似度,其中,所述文本相似度为所述验证文本与所述验证信息之间的相似度信息;
将验证文本与验证信息进行相似度对比,得到对应的文本相似度,具体地,将验证文本转化为Unicode字符或GBK\GB2312字符,并与验证信息的字符进行对比,判断其中的汉明距离,以汉明距离与验证信息的字符总数量的比值确定文本相似度。在一些实施方式中,可以将文本中的每一个词汇或单独的汉字按排序与验证信息中对应位置的词汇或汉字进行字符间的汉明距离对比,当得到的汉明距离大于零时,确定对应的词汇或汉字不对应,统计验证文本与验证信息间不对应的词汇数或汉字数,与验证信息的总字量计算得到比值,以该比值作为文本相似度。
由于汉字中存在大量的同音或近似音词汇或汉字,因此可以进行模糊对比,将获取得到的验证文本转化为拼音字符,与验证信息的拼音字符通过前述的多种方法中的一种得到文相似度。
步骤d3、验证所述文本相似度是否大于所述预设的相似度阈值;
系统中预设有相似度阈值,用于判断验证文本和验证信息的相似度是否符合验证规则,相似度阈值的取值可以根据实际情况进行调整,例如选用比较精确的相似度确定方法时,可以提高相似度阈值的取值,当选用比较粗略的相似度确定方法时,可以降低相似度阈值的取值。以文本相似度与相似度阈值的对比结果确定语音信息是否符合验证规则,当文本相似度大于相似度阈值时,确定语音信息符合验证规则,验证通过;当文本相似度小于或等于相似度阈值时,确定语音信息不符合验证规则,验证失败。
步骤d1具体包括下述步骤:
步骤d11、将所述验证语音信息输入到预设的语音识别模型中,其中,所述语音识别模型为根据输入的语音信息转换得到与语音信息的内容相对应的文本的自然语言解析模型;
将语音信息输入到语音识别模型中,首先根据语音信息进行分段,分段的依据可以是讲话过程中的停顿,或者按照讲话的音节,将语音信息进行分段后得到分段语音,再将分段语音输入到语音识别模型中进行分词提取,提取得到零散的词语或音节,语音识别模型可以是现有的,将输入的语音信息转换得为文本的自然语言解析模型。
步骤d12、根据所述语音识别模型的输出结果确定所述验证文本;
将语音识别模型输出的词语或音节根据分段的先后顺序进行拼接,并且根据整句的语义进行同音词的替换调整,获得完整的句子作为文本信息。同音词调整的依据可以是预设的词语搭配关系,或者与预设的例句进行相似度匹配,根据匹配得到的相近句子中的词语进行替换。
通过利用语音模型提取语音信息中的内容并转化为文本,可以准确得获得对应的文本内容,在与验证信息进行对比时更加便捷,确定语音验证的准确性。
为解决上述技术问题,本申请实施例还提供一种语音验证装置。具体请参阅图3,图3为本实施语音验证装置的基本结构框图。
如图3所示,语音验证装置,包括:获取模块2100、处理模块2200和执行模块2300。其中,获取模块用于获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;处理模块用于根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类;执行模块用于当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败。与现有技术相比,本申请实施例的技术方案侧重于挖掘用户的生物学语音特征,此特征可以区分机器声模拟人声和真实人声的差别,基于该特征能够实现有效的鉴别真实用户。通过对验证语音是否为真实人声校验,可以有效排除机器、AI、爬虫等恶意用户,防止此类恶意用户对网站、平台的攻击,保证验证用户有效性和真实性,提升语音验证的安全性。
在一些实施方式中,语音验证装置还包括:第一解析子模块、第一输入子模块、第一处理子模块。其中第一解析子模块用于解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;第一输入子模块用于将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;第一处理子模块用于根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。在一些实施方式中,语音验证装置还包括:第二处理子模块、第三处理子模块、第一执行子模块。其中,第二处理子模块用于根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;第三处理子模块用于根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;第一执行子模块用于定义所述时域数据和所述频谱数据为所述特征数据。在一些实施方式中,语音验证装置还包括:第一获取子模块、第一查找子模块、第一发送子模块。其中,第一获取子模块用于获取目标终端的验证请求;第一查找子模块用于根据所述验证请求在预设的验证数据库随机查找一个文本作为所述验证信息;第一发送子模块用于将所述验证信息发送至目标终端,触发预设的提醒指令,以引导验证用户根据所述验证信息进行语音验证。在一些实施方式中,语音验证装置还包括:第二执行子模块、第三执行子模块、第四执行子模块。其中,第二执行子模块用于当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证,其中,所述验证规则为判断所述验证语音信息的内容与所述验证信息的相似度是否大于预设的相似度阈值的数据对比规则;第三执行子模块用于当所述验证语音信息符合所述验证规则时,确定语音验证通过;第四执行子模块用于当所述验证语音信息不符合所述验证规则时,确定语音验证失败。在一些实施方式中,语音验证装置还包括:第四处理子模块、第五处理子模块、第一验证子模块。其中,第四处理子模块用于根据所述验证语音信息生成验证文本,其中,所述验证文本为对所述验证语音信息进行内容识别后得到的与所述验证语音信息的内容相对应的文本信息;第五处理子模块用于根据所述验证文本确定文本相似度,其中,所述文本相似度为所述验证文本与所述验证信息之间的相似度信息;第一验证子模块用于验证所述文本相似度是否大于所述预设的相似度阈值。在一些实施方式中,语音验证装置还包括:第二输入子模块、第六处理子模块。其中,第二输入子模块用于将所述验证语音信息输入到预设的语音识别模型中,其中,所述语音识别模型为根据输入的语音信息转换得到与语音信息的内容相对应的文本的自然语言解析模型;第六处理子模块用于根据所述语音识别模型的输出结果确定所述验证文本。
为解决上述技术问题,本申请实施例还提供一种计算机设备。具体请参阅图4,图4为本实施例计算机设备基本结构框图。
如图4所示,计算机设备的内部结构示意图。如图4所示,该计算机设备包括通过系统总线连接的处理器、非易失性存储介质、存储器和网络接口。其中,该计算机设备的非易失性存储介质存储有操作系统、数据库和计算机可读指令,数据库中可存储有控件信息序列,该计算机可读指令被处理器执行时,可使得处理器实现一种语音验证方法。该计算机设备的处理器用于提供计算和控制能力,支撑整个计算机设备的运行。该计算机设备的存储器中可存储有计算机可读指令,该计算机可读指令被处理器执行时,可使得处理器执行一种语音验证方法。该计算机设备的网络接口用于与终端连接通信。本领域技术人员可以理解,图中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定,具体的计算机设备可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
本实施方式中处理器用于执行图3中获取模块2100、处理模块2200和执行模块2300的具体功能,存储器存储有执行上述模块所需的程序代码和各类数据。网络接口用于向用户终端或服务器之间的数据传输。本实施方式中的存储器存储有语音验证装置中执行所有子模块所需的程序代码及数据,服务器能够调用服务器的程序代码及数据执行所有子模块的功能。
本申请还提供一种存储有计算机可读指令的存储介质,所述计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行上述任一实施例所述语音验证方法的步骤。所述存储介质可以为非易失性可读存储介质。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,该计算机程序可存储于一计算机可读取存储介质中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,前述的存储介质可为磁碟、光盘、只读存储记忆体(Read-Only
Memory,ROM)等非易失性存储介质,或随机存储记忆体(Random Access Memory,RAM)等。
以上所述实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对本申请专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。
Claims (20)
- 一种语音验证方法,其特征在于,包括以下步骤:获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为用于表征语音内容为人类声音的声音分类;当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;其中,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤,包括以下步骤:解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
- 如权利要求1所述的语音验证方法,其特征在于,所述解析所述验证语音信息得到特征数据的步骤,包括以下步骤:根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;定义所述时域数据和所述频谱数据为所述特征数据。
- 如权利要求1所述的语音验证方法,其特征在于,所述获取验证语音信息的步骤之前,包括以下步骤:获取目标终端的验证请求;根据所述验证请求在预设的验证数据库随机查找一个文本作为所述验证信息;将所述验证信息发送至目标终端,触发预设的提醒指令,以引导验证用户根据所述验证信息进行语音验证。
- 如权利要求1所述的语音验证方法,其特征在于,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤之后,包括下述步骤:当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证,其中,所述验证规则为判断所述验证语音信息的内容与所述验证信息的相似度是否大于预设的相似度阈值的数据对比规则;当所述验证语音信息符合所述验证规则时,确定语音验证通过;当所述验证语音信息不符合所述验证规则时,确定语音验证失败。
- 如权利要求4所述的语音验证方法,其特征在于,所述当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证的步骤,包括下述步骤:根据所述验证语音信息生成验证文本,其中,所述验证文本为对所述验证语音信息进行内容识别后得到的与所述验证语音信息的内容相对应的文本信息;根据所述验证文本确定文本相似度,其中,所述文本相似度为所述验证文本与所述验证信息之间的相似度信息;验证所述文本相似度是否大于所述预设的相似度阈值。
- 如权利要求5所述的语音验证方法,其特征在于,所述根据所述验证语音信息生成验证文本的步骤,包括下述步骤:将所述验证语音信息输入到预设的语音识别模型中,其中,所述语音识别模型为根据输入的语音信息转换得到与语音信息的内容相对应的文本的自然语言解析模型;根据所述语音识别模型的输出结果确定所述验证文本。
- 一种语音验证装置,其特征在于,包括:获取模块,用于获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;处理模块,用于根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为表征语音内容为人类声音的声音分类;执行模块,用于当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;第一解析子模块,用于解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;第一输入子模块,用于将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;第一处理子模块,用于根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
- 如权利要求7所述的语音验证装置,其特征在于,所述语音验证装置还包括:第二处理子模块用于根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;第三处理子模块用于根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;第一执行子模块用于定义所述时域数据和所述频谱数据为所述特征数据。
- 如权利要求7所述的语音验证装置,其特征在于,所述语音验证装置还包括:第一获取子模块,用于获取目标终端的验证请求;第一查找子模块,用于根据所述验证请求在预设的验证数据库随机查找一个文本作为所述验证信息;第一发送子模块,用于将所述验证信息发送至目标终端,触发预设的提醒指令,以引导验证用户根据所述验证信息进行语音验证。
- 如权利要求7所述的语音验证装置,其特征在于,所述语音验证装置还包括:第二执行子模块,用于当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证,其中,所述验证规则为判断所述验证语音信息的内容与所述验证信息的相似度是否大于预设的相似度阈值的数据对比规则;第三执行子模块,用于当所述验证语音信息符合所述验证规则时,确定语音验证通过;第四执行子模块,用于当所述验证语音信息不符合所述验证规则时,确定语音验证失败。
- 如权利要求10所述的语音验证装置,其特征在于,所述语音验证装置还包括:第四处理子模块,用于根据所述验证语音信息生成验证文本,其中,所述验证文本为对所述验证语音信息进行内容识别后得到的与所述验证语音信息的内容相对应的文本信息;第五处理子模块,用于根据所述验证文本确定文本相似度,其中,所述文本相似度为所述验证文本与所述验证信息之间的相似度信息;第一验证子模块,用于验证所述文本相似度是否大于所述预设的相似度阈值。
- 如权利要求11所述的语音验证装置,其特征在于,所述语音验证装置还包括:第二输入子模块,用于将所述验证语音信息输入到预设的语音识别模型中,其中,所述语音识别模型为根据输入的语音信息转换得到与语音信息的内容相对应的文本的自然语言解析模型;第六处理子模块,用于根据所述语音识别模型的输出结果确定所述验证文本。
- 一种计算机设备,其特征在于,包括:处理器;用于存储被处理器执行的计算机可读指令的存储器;其中,所述处理器被配置为执行如下步骤:获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为用于表征语音内容为人类声音的声音分类;当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;其中,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤,包括以下步骤:解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
- 如权利要求13所述的计算机设备,其特征在于,所述解析所述验证语音信息得到特征数据的步骤,包括以下步骤:根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;定义所述时域数据和所述频谱数据为所述特征数据。
- 如权利要求13所述的计算机设备,其特征在于,所述获取验证语音信息的步骤之前,包括以下步骤:获取目标终端的验证请求;根据所述验证请求在预设的验证数据库随机查找一个文本作为所述验证信息;将所述验证信息发送至目标终端,触发预设的提醒指令,以引导验证用户根据所述验证信息进行语音验证。
- 如权利要求13所述的计算机设备,其特征在于,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤之后,包括下述步骤:当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证,其中,所述验证规则为判断所述验证语音信息的内容与所述验证信息的相似度是否大于预设的相似度阈值的数据对比规则;当所述验证语音信息符合所述验证规则时,确定语音验证通过;当所述验证语音信息不符合所述验证规则时,确定语音验证失败。
- 如权利要求16所述的计算机设备,其特征在于,所述当判断所述语音内容属于预设的声音类别时,根据预设的验证规则对所述语音信息进行验证的步骤,包括下述步骤:根据所述验证语音信息生成验证文本,其中,所述验证文本为对所述验证语音信息进行内容识别后得到的与所述验证语音信息的内容相对应的文本信息;根据所述验证文本确定文本相似度,其中,所述文本相似度为所述验证文本与所述验证信息之间的相似度信息;验证所述文本相似度是否大于所述预设的相似度阈值。
- 如权利要求17所述的计算机设备,其特征在于,所述根据所述验证语音信息生成验证文本的步骤,包括下述步骤:将所述验证语音信息输入到预设的语音识别模型中,其中,所述语音识别模型为根据输入的语音信息转换得到与语音信息的内容相对应的文本的自然语言解析模型;根据所述语音识别模型的输出结果确定所述验证文本。
- 一种非临时性计算机可读存储介质,其特征在于,当所述存储介质中的计算机可读指令由移动终端的处理器执行时,使得移动终端能够执行以下步骤:获取验证语音信息,其中,所述验证语音信息为验证用户在朗读验证信息时,目标终端采集到的语音内容;根据所述验证语音信息判断所述语音内容是否为预设的声音类别,其中,所述预设的声音类别为用于表征语音内容为人类声音的声音分类;当判断所述语音内容不属于所述预设的声音类别时,确定语音验证失败;其中,所述根据所述验证语音信息判断所述语音内容是否为预设的声音类别的步骤,包括以下步骤:解析所述验证语音信息得到特征数据,其中,所述特征数据为将语音信息处理得到的时域数据和频谱数据;将所述特征数据输入到预设的人声判断模型中,其中,所述人声判断模型为已训练至收敛的,用于根据输入的特征数据判断语音信息是否为人声的神经网络模型;根据所述人声判断模型的输出结果确定所述语音内容是否为预设的声音类别。
- 如权利要求19所述的非临时性计算机可读存储介质,其特征在于,所述解析所述验证语音信息得到特征数据的步骤,包括以下步骤:根据预设的第一处理规则对所述验证语音信息进行处理,得到时域数据,其中,所述第一处理规则为将语音信息解析为时域数据并提升其中的高频部分的语音信息处理规则;根据预设的第二处理规则对所述时域数据进行处理,得到声音频谱,其中,所述第二处理规则为根据傅里叶变换将时域数据转换为频谱数据的数据处理规则;定义所述时域数据和所述频谱数据为所述特征数据。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910068827.9A CN109801638B (zh) | 2019-01-24 | 2019-01-24 | 语音验证方法、装置、计算机设备及存储介质 |
| CN201910068827.9 | 2019-01-24 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020151317A1 true WO2020151317A1 (zh) | 2020-07-30 |
Family
ID=66560320
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/117613 Ceased WO2020151317A1 (zh) | 2019-01-24 | 2019-11-12 | 语音验证方法、装置、计算机设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109801638B (zh) |
| WO (1) | WO2020151317A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116030820A (zh) * | 2022-11-28 | 2023-04-28 | 浙江大学 | 音频验证方法及装置、音频取证方法及装置 |
| CN117854185A (zh) * | 2024-01-16 | 2024-04-09 | 北京摇光智能科技有限公司 | 智能锁访客出入验证方法、装置及计算机设备 |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109801638B (zh) * | 2019-01-24 | 2023-10-13 | 平安科技(深圳)有限公司 | 语音验证方法、装置、计算机设备及存储介质 |
| CN110727934A (zh) * | 2019-10-22 | 2020-01-24 | 成都知道创宇信息技术有限公司 | 一种反爬虫方法及装置 |
| CN110931020B (zh) * | 2019-12-11 | 2022-05-24 | 北京声智科技有限公司 | 一种语音检测方法及装置 |
| CN112185417B (zh) * | 2020-10-21 | 2024-05-10 | 平安科技(深圳)有限公司 | 人工合成语音检测方法、装置、计算机设备及存储介质 |
| CN113516154A (zh) * | 2021-04-09 | 2021-10-19 | 北京小米移动软件有限公司 | 识别媒体文件中人声配音类型的方法、装置及存储介质 |
| CN112948788B (zh) * | 2021-04-13 | 2024-05-31 | 杭州网易智企科技有限公司 | 语音验证方法、装置、计算设备以及介质 |
| CN114822557B (zh) * | 2022-04-01 | 2025-04-04 | 北京中庆现代技术股份有限公司 | 课堂中不同声音的区分方法、装置、设备以及存储介质 |
| CN114822556B (zh) * | 2022-04-01 | 2025-11-04 | 北京中庆现代技术股份有限公司 | 教师声音和非教师声音的区分方法、装置、设备以及介质 |
| CN117201879B (zh) * | 2023-11-06 | 2024-04-09 | 深圳市微浦技术有限公司 | 机顶盒显示方法、装置、设备及存储介质 |
| CN119446143A (zh) * | 2024-10-31 | 2025-02-14 | 北京百度网讯科技有限公司 | 应用控制方法、装置、电子设备及存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102820033A (zh) * | 2012-08-17 | 2012-12-12 | 南京大学 | 一种声纹识别方法 |
| CN106954136A (zh) * | 2017-05-16 | 2017-07-14 | 成都泰声科技有限公司 | 一种集成麦克风接收阵列的超声定向发射参量阵 |
| US20170248955A1 (en) * | 2016-02-26 | 2017-08-31 | Ford Global Technologies, Llc | Collision avoidance using auditory data |
| CN109801638A (zh) * | 2019-01-24 | 2019-05-24 | 平安科技(深圳)有限公司 | 语音验证方法、装置、计算机设备及存储介质 |
Family Cites Families (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102402985A (zh) * | 2010-09-14 | 2012-04-04 | 盛乐信息技术(上海)有限公司 | 提高声纹识别安全性的声纹认证系统及其实现方法 |
| US9390245B2 (en) * | 2012-08-02 | 2016-07-12 | Microsoft Technology Licensing, Llc | Using the ability to speak as a human interactive proof |
| EP3078026B1 (en) * | 2013-12-06 | 2022-11-16 | Tata Consultancy Services Limited | System and method to provide classification of noise data of human crowd |
| CN104660413A (zh) * | 2015-01-28 | 2015-05-27 | 中国科学院数据与通信保护研究教育中心 | 一种声纹口令认证方法和装置 |
| CN107404381A (zh) * | 2016-05-19 | 2017-11-28 | 阿里巴巴集团控股有限公司 | 一种身份认证方法和装置 |
| US10692502B2 (en) * | 2017-03-03 | 2020-06-23 | Pindrop Security, Inc. | Method and apparatus for detecting spoofing conditions |
| CN108877813A (zh) * | 2017-05-12 | 2018-11-23 | 阿里巴巴集团控股有限公司 | 人机识别的方法、装置和系统 |
| CN109218269A (zh) * | 2017-07-05 | 2019-01-15 | 阿里巴巴集团控股有限公司 | 身份认证的方法、装置、设备及数据处理方法 |
| CN108198561A (zh) * | 2017-12-13 | 2018-06-22 | 宁波大学 | 一种基于卷积神经网络的翻录语音检测方法 |
| CN108039176B (zh) * | 2018-01-11 | 2021-06-18 | 广州势必可赢网络科技有限公司 | 一种防录音攻击的声纹认证方法、装置及门禁系统 |
| CN108281158A (zh) * | 2018-01-12 | 2018-07-13 | 平安科技(深圳)有限公司 | 基于深度学习的语音活体检测方法、服务器及存储介质 |
| CN108711436B (zh) * | 2018-05-17 | 2020-06-09 | 哈尔滨工业大学 | 基于高频和瓶颈特征的说话人验证系统重放攻击检测方法 |
| CN109065030B (zh) * | 2018-08-01 | 2020-06-30 | 上海大学 | 基于卷积神经网络的环境声音识别方法及系统 |
| CN109147799A (zh) * | 2018-10-18 | 2019-01-04 | 广州势必可赢网络科技有限公司 | 一种语音识别的方法、装置、设备及计算机存储介质 |
-
2019
- 2019-01-24 CN CN201910068827.9A patent/CN109801638B/zh active Active
- 2019-11-12 WO PCT/CN2019/117613 patent/WO2020151317A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102820033A (zh) * | 2012-08-17 | 2012-12-12 | 南京大学 | 一种声纹识别方法 |
| US20170248955A1 (en) * | 2016-02-26 | 2017-08-31 | Ford Global Technologies, Llc | Collision avoidance using auditory data |
| CN106954136A (zh) * | 2017-05-16 | 2017-07-14 | 成都泰声科技有限公司 | 一种集成麦克风接收阵列的超声定向发射参量阵 |
| CN109801638A (zh) * | 2019-01-24 | 2019-05-24 | 平安科技(深圳)有限公司 | 语音验证方法、装置、计算机设备及存储介质 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116030820A (zh) * | 2022-11-28 | 2023-04-28 | 浙江大学 | 音频验证方法及装置、音频取证方法及装置 |
| CN117854185A (zh) * | 2024-01-16 | 2024-04-09 | 北京摇光智能科技有限公司 | 智能锁访客出入验证方法、装置及计算机设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109801638B (zh) | 2023-10-13 |
| CN109801638A (zh) | 2019-05-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020151317A1 (zh) | 语音验证方法、装置、计算机设备及存储介质 | |
| CN110853615B (zh) | 一种数据处理方法、装置及存储介质 | |
| WO2020139058A1 (en) | Cross-device voiceprint recognition | |
| CN115641860A (zh) | 模型的训练方法、语音转换方法和装置、设备及存储介质 | |
| KR20240037592A (ko) | 인공지능 활용 상담 보조 시스템 및 방법 | |
| Płaza et al. | Call transcription methodology for contact center systems | |
| CN114996506A (zh) | 语料生成方法、装置、电子设备和计算机可读存储介质 | |
| Kopparapu | Non-linguistic analysis of call center conversations | |
| CN107886951A (zh) | 一种语音检测方法、装置及设备 | |
| CN112087726B (zh) | 彩铃识别的方法及系统、电子设备及存储介质 | |
| CN112231440A (zh) | 一种基于人工智能的语音搜索方法 | |
| WO2021251539A1 (ko) | 인공신경망을 이용한 대화형 메시지 구현 방법 및 그 장치 | |
| CN114125506B (zh) | 语音审核方法及装置 | |
| CN105957517A (zh) | 基于开源api的语音数据结构化转换方法及其系统 | |
| CN114707515A (zh) | 话术判别方法、装置、电子设备及存储介质 | |
| CN113158052B (zh) | 聊天内容推荐方法、装置、计算机设备及存储介质 | |
| WO2022154217A1 (ko) | 음성 장애 환자를 위한 음성 자가 훈련 방법 및 사용자 단말 장치 | |
| Burkhardt et al. | Masking speech contents by random splicing: is emotional expression preserved? | |
| CN110298150A (zh) | 一种基于语音识别的身份验证方法及系统 | |
| CN117174092A (zh) | 基于声纹识别与多模态分析的移动语料转写方法及装置 | |
| CN111916106B (zh) | 一种提高英语教学中发音质量的方法 | |
| CN114999444A (zh) | 语音合成模型的训练方法、装置、电子设备及存储介质 | |
| KR102755648B1 (ko) | 인공지능 기반의 다국어 통역 시스템 및 서비스 방법 | |
| CN118588112B (zh) | 一种针对非言语信号的交流状态分析方法、设备及介质 | |
| JP7422702B2 (ja) | バリアフリースマート音声システムとその制御方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19911634 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19911634 Country of ref document: EP Kind code of ref document: A1 |