WO2019196238A1 - 一种语音识别方法、终端设备及计算机可读存储介质 - Google Patents
一种语音识别方法、终端设备及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2019196238A1 WO2019196238A1 PCT/CN2018/096263 CN2018096263W WO2019196238A1 WO 2019196238 A1 WO2019196238 A1 WO 2019196238A1 CN 2018096263 W CN2018096263 W CN 2018096263W WO 2019196238 A1 WO2019196238 A1 WO 2019196238A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- text
- information
- voice
- type
- adjusted
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
Definitions
- the present application belongs to the field of information processing technologies, and in particular, to a voice recognition method, a terminal device, and a computer readable storage medium.
- the existing intelligent voice robot can perform business processing or information transmission according to the user's voice, in the process of recognizing the user's voice, if there are syllables in the voice content that are easily confused, for example, the number "1" and The letter “E” is easy to cause inaccurate recognition results.
- the embodiment of the present application provides a voice recognition method, a terminal device, and a computer readable storage medium, so as to solve the problem that the recognition result is inaccurate in the existing voice recognition technology.
- a first aspect of the embodiments of the present application provides a voice recognition method, including:
- a voice recognition method, a terminal device, and a computer readable storage medium provided by embodiments of the present application have the following beneficial effects:
- the received voice response information returned by the incoming call terminal according to the voice query information is used to divide the text to be adjusted from the voice content text corresponding to the voice response information.
- the reference text is determined from the preset database, and finally the text to be adjusted is adjusted according to the reference text to obtain the target information, thereby improving the accuracy of the voice recognition.
- FIG. 1 is a flowchart of implementing a voice recognition method according to an embodiment of the present application
- FIG. 2 is a flowchart of an implementation of a voice recognition method according to another embodiment of the present application.
- FIG. 3 is a flowchart of a specific implementation of a voice recognition method S12 according to another embodiment of the present application.
- FIG. 4 is a flowchart of a specific implementation of a voice recognition method S13 according to another embodiment of the present application.
- FIG. 5 is a flowchart of a specific implementation of a voice recognition method S14 according to another embodiment of the present application.
- FIG. 6 is a structural block diagram of a terminal device according to an embodiment of the present application.
- FIG. 7 is a schematic diagram of a terminal device according to another embodiment of the present application.
- the received voice response information returned by the incoming call terminal according to the voice query information is used to divide the text to be adjusted from the voice content text corresponding to the voice response information.
- the reference text is determined from the preset database, and finally the text to be adjusted is adjusted according to the reference text to obtain the target information, and the recognition result existing in the existing voice recognition technology is solved. Inaccurate question.
- the execution body of the voice recognition method is a server device.
- the server device includes, but is not limited to, a computer, or may be other network devices or communication devices having data processing capabilities, and the like.
- FIG. 1 is a flowchart of an implementation of a voice recognition method provided by an embodiment of the present application, which is described in detail as follows:
- the voice query information is voice content pre-recorded into the server, and is used for voice inquiry to the user corresponding to the incoming call terminal, wherein the content of the voice query can be customized by the operator according to requirements.
- the voice response information is voice information returned by the user to the server through the incoming call terminal after receiving the voice query information.
- the incoming call terminal may be a mobile terminal or a non-mobile terminal, such as a mobile phone, a tablet computer, or a fixed telephone.
- the server sends voice inquiry information to the incoming call terminal, and then receives the voice response information returned by the user through the incoming call terminal.
- the user sends an instruction to the incoming call terminal to receive the voice query information, and then the server sends the voice query message to the incoming terminal according to the instruction, and the received user returns through the incoming call terminal.
- Voice response information is the call link between the incoming call terminal and the server.
- the preset operation of when the voice inquiry information is sent to the incoming terminal may be detected, which may include, but is not limited to, the following scenarios.
- Scenario 1 When it is detected that a call link is established between the server and the incoming terminal, the triggering operation of the server to send voice inquiry information to the incoming terminal is triggered.
- the user sends a call request to the server through the terminal, and the server establishes a call link with the terminal according to the call request, and triggers an operation of sending the voice query information to the caller terminal, so as to send the voice query information to the terminal.
- Scenario 2 After the call link is established between the incoming call terminal and the server, the user sends an instruction to the incoming call terminal to receive the voice query information, and according to the instruction, the trigger server sends the voice query information to the incoming call terminal.
- the user triggers a request sending instruction on the terminal, so that the terminal sends an instruction to the server to receive the voice query information, and then triggers the server to receive the voice query information according to the request, and sends an incoming call.
- the terminal sends an operation of the voice inquiry information to implement sending the voice inquiry information to the terminal.
- the received call terminal can be made into a voice response file according to the voice response information returned by the voice inquiry information by voice recording, so as to facilitate optimization and recognition.
- step S12 the voice content text corresponding to the voice response information is obtained by text conversion of the voice response information.
- the text to be adjusted is part or all of the text in the voice content text corresponding to the voice response information.
- the voice content text corresponding to the voice response information may be divided by calling a pre-configured text segmentation policy to be adjusted, and then the text to be adjusted is divided.
- the text segmentation policy to be adjusted may include multiple policies.
- the corresponding text segmentation strategy to be adjusted may be formulated according to the character types included in the voice content text corresponding to the voice response information.
- S13 Determine a reference text from a preset database based on a phone number of the incoming terminal and a content type of the text to be adjusted.
- the data in the preset database is used to describe the correspondence between the phone number, the content type, and the reference text.
- the content type is the content type of the text to be adjusted, including: a single character type or a mixed character type, wherein the single character type refers to the content of the text to be adjusted is composed of the same character, and the mixed character type refers to the content of the text to be adjusted. Consists of at least two characters.
- the content of the text to be adjusted is used to describe the user name, that is, the content type of the text to be adjusted is text, and belongs to a single character type.
- the content of the text to be adjusted is used to describe the license plate number, that is, the content type of the text to be adjusted includes letters and numbers, or includes characters, letters, and numbers, and belongs to a mixed character type.
- the preset database is a database for storing user information
- different user information may be determined from a preset database according to different phone numbers
- the reference text is determined from the user information according to the content type.
- the content type of the reference text is the same as the content type of the text to be adjusted.
- the user information stored in the preset database can be obtained through a phone number search, wherein the user information includes all information related to the user, for example, an ID number, an address, a license plate number, and the like.
- step S14 the reference text is the text type and the content type of the text to be adjusted, as an index, and the obtained text is searched from the preset database.
- the content type of the reference text is the same as the content type of the text to be adjusted, that is, the character type constituting the reference text is the same as the character type constituting the text to be adjusted.
- the text to be adjusted is the license plate number
- the content of the text to be adjusted is “Beijing A12345”
- the content type of the text to be adjusted is a mixed character type, that is, the character types constituting the text to be adjusted include characters, letters, and numbers. Determining the user information from the preset database based on the telephone number of the incoming call terminal, and determining the reference text from the user information according to the content type of the text to be adjusted. Since the text to be adjusted is the license plate number, the content type of the text to be adjusted In order to mix character types, the reference text determined from the user information should also be the license plate number in the user information, and the license plate number should also include text, letters and numbers.
- the text to be adjusted is adjusted according to the reference text, and may be separately compared and adjusted based on the character types included in the reference text, and the target information is obtained.
- the content of the text to be adjusted is “Beijing AE2345”, the content of the reference text is “Jin A12345”, and the text to be adjusted is adjusted according to the reference text, and the obtained target information should be “Jin A12345”.
- the voice recognition method receives the voice response information returned by the incoming call terminal according to the voice query information, and receives the voice response information from the voice response information when the voice call inquiry information is sent to the incoming call terminal.
- the text to be adjusted is divided into the text content corresponding to the information, and the reference text is determined from the preset database based on the telephone number of the incoming call terminal and the content type of the text to be adjusted, and finally the text to be adjusted is adjusted according to the reference text to obtain the target information.
- FIG. 2 is a flowchart of an implementation of a voice recognition method according to another embodiment of the present application. As shown in FIG. 2, with reference to the embodiment shown in FIG. 1, the voice recognition method provided in this embodiment further includes S21, S201, and S22, which are specifically described as follows:
- the method before the text to be adjusted is divided from the voice content text corresponding to the voice response information, the method further includes:
- S21 Acquire an identifier of the voice query information, where the identifier is used to distinguish a character type included in the text corresponding to the voice response information.
- S22 Determine, according to the character type, a target state network from a preset form, where data in the preset form is used to describe a correspondence between the character type and the target state network; the target state The network is configured to perform text conversion on the voice response information to obtain a voice content text corresponding to the voice information.
- the voice query information is voice content pre-recorded into the server
- the voice content may be determined, and then the text included in the text corresponding to the voice response information returned by the electronic terminal according to the voice query information may be predicted. Character type.
- the content of the voice inquiry information is used to request the user to input the identity card number through the caller terminal, and thus the voice response information sent by the user through the caller terminal can be determined, and the corresponding voice content text is necessarily an ID card number composed of numbers. Therefore, it can be determined that the character type included in the voice content text corresponding to the voice response information is a number.
- the content of the voice inquiry information is used to request the user to input the license plate number through the incoming call terminal, thereby determining the voice response information sent by the user through the incoming call terminal, and the corresponding voice content text is necessarily composed of characters, letters, and numbers.
- the license plate number so it can be determined that the character types included in the voice content text corresponding to the voice response information are characters, letters, and numbers.
- different voice identification information may be configured with different identifiers, thereby distinguishing the text corresponding to the voice response information.
- the audio framing function can be called, for example, by calling a moving window function to framing a voice file to obtain multi-frame speech, and then performing acoustic feature extraction on each frame of speech. Processing, converting each frame waveform in the voice information into a multi-dimensional vector, thereby obtaining a matrix composed of a plurality of multi-dimensional vectors, wherein each multi-dimensional vector contains content information of the corresponding speech frame, and in the matrix, several frames of speech Corresponding to one state, every three states are combined into one phoneme, and several phonemes are combined into one word.
- the recognition of the speech content can be realized according to the relationship between the state, the phoneme, and the words.
- the character type includes at least one of a text type, a letter type, and a digital type, and different character types correspond to different target state networks.
- step S201 along with step S21 may be further included.
- step S21 and step S201 are performed in partial order.
- S201 Create a state network corresponding to each of the character types, where the state network is configured to reflect that the voice response information corresponding to the character type is converted into an optimal path of the voice content text.
- step S201 the state network is developed into a phoneme network by a word-level network, and the phoneme network is expanded into a state network.
- the cumulative transition probability includes: an observation probability, a transition probability, and a language probability.
- observation probability refers to the probability of each frame of speech and each state.
- transition probability refers to the probability that each state is transferred to itself or to the next state.
- the language probability is obtained by the law of language statistics. The probability. Both the observation probability and the transition probability can be obtained by inputting a preset acoustic model. The language probability can be obtained by inputting a preset language model.
- the language model is trained using a large amount of text, and can utilize the statistics of a certain language itself. Regularity helps to improve recognition accuracy.
- Different state networks are created according to different character types, so that when the text information is converted, the character type can be distinguished according to the pre-configured identifier, and the target state network is determined from the preset form based on the character type, that is, the selection and the The target state network corresponding to the character type included in the text corresponding to the voice information is text-converted to the voice information, and the voice content text corresponding to the voice information is obtained.
- the identifier of the voice query information is obtained to distinguish the character type included in the text corresponding to the voice response information, and the state network corresponding to each character type is created, and the target state is determined from the preset form based on the character type.
- the network that is, the voice response information corresponding to the character type is determined to be converted into the best path of the voice content text, thereby improving the conversion efficiency of the voice information into text information.
- FIG. 3 is a flowchart showing a specific implementation of a voice recognition method S12 according to another embodiment of the present application. As shown in FIG. 3, based on the foregoing embodiments, in the speech recognition method provided in this embodiment, S12 includes S121, S122, and S123, and the details are as follows:
- S121 Identify the number of character types included in the text information, where the number of character types is greater than or equal to 1.
- the number of character types included in the text information is used to reflect that the character type of the text information belongs to a single character type or a mixed character type.
- the character type indicating the text information belongs to a single character type; when the number of character types included in the text information is greater than 1, the character type indicating the text information belongs to a mixed character type.
- the text of the voice content is divided according to the position of the preset keyword in the text of the voice content, thereby obtaining Adjust the text.
- the text of the voice content is “Futian District, Shenzhen City, Guangdong province”, and the preset keywords are “province”, “city” and “zone”, and the text of the voice content is based on the position of the preset keyword in the text of the voice content.
- the division is carried out, and the texts to be adjusted are obtained as "Guangdong City", “Shenzhen City” and "Futian District”.
- composition words can be configured for different preset keywords.
- the name of the province with the longest name is “Heilongjiang Province”
- the default keyword is “province”
- the corresponding text constitutes 3 words.
- the name of the city with the longest name is “Hohhot City”
- the default keyword is “City”
- the corresponding text constitutes 4 words.
- the text message is “My address is Futian District, Shenzhen, Guangdong City”, and the text information is divided according to the preset key characters to obtain the texts to be adjusted as “Guangdong City”, “Shenzhen City” and “Futian District”.
- the text message is “My license plate number is Beijing AE2345”, and the content of the character information in the text information is separately divided, and the obtained text to be adjusted includes “My license plate number is Beijing”, “AE”, and “ 2345”.
- the complicated information is avoided when the constituent elements of the text information are relatively simple.
- the way to divide the text to be adjusted makes the data processing process more reasonable.
- FIG. 4 is a flowchart showing a specific implementation of a voice recognition method S13 according to another embodiment of the present application.
- the content type of the text to be adjusted includes any one of a text type, a letter type, and a numeric type.
- S13 includes S131 and S132, and the details are as follows:
- S131 Obtain target user information from a preset database according to the phone number.
- S132 Determine, from the target user information, information that matches a content type of the text to be adjusted as the reference text.
- the target user information there is a correspondence between the target user information and the phone number, and the corresponding target user information can be found from the preset database by using the phone number as an index, wherein the target user information may include multiple types of information of the target user. For example, an ID number, a license plate number, or an address.
- the target user information includes multiple types of information of the user
- the reference text cannot be directly determined therefrom.
- the matching information is determined from the target user information, and the information is used as a reference. text.
- the target user information is obtained from the preset database according to the phone number, and according to the content type of the adjusted text, the information matching the target user information is determined, and the target user information can be avoided. All the information is filtered one by one, which improves the speed of determining the reference text.
- FIG. 5 is a flowchart showing a specific implementation of a voice recognition method S14 according to another embodiment of the present application.
- S14 includes S141 and S142 in a voice recognition method according to the foregoing embodiments. The details are as follows:
- S141 Identify target content different from the reference text from the to-be-adjusted text.
- the target content is information in the text to be adjusted that is different from the reference text content.
- the target content different from the reference text is determined from the text to be adjusted.
- the reference text is used as the text to be compared with the text to be adjusted, when the voice content text corresponding to the voice response information does not include the target user information, it is not necessary to use the reference text to adjust the text to be adjusted.
- the target content in the text to be adjusted is part of the text to be adjusted, when the target user information is not included in the voice content text corresponding to the voice response information, the text content of the voice content is prevented from being adjusted, thereby preventing voice conversion.
- adjustment disorder or text conversion disorder occurs.
- the voice recognition method receives the voice response information returned by the incoming call terminal according to the voice query information, and receives the voice response information from the voice response information when the voice call inquiry information is sent to the incoming call terminal.
- the text to be adjusted is divided into the text content corresponding to the information, and the reference text is determined from the preset database based on the telephone number of the incoming call terminal and the content type of the text to be adjusted, and finally the text to be adjusted is adjusted according to the reference text to obtain the target information.
- the target state network is determined from the preset form based on the character type, that is, the voice response information corresponding to the character type is determined to be converted into the best path of the voice content text, thereby improving The conversion efficiency of voice information into text information.
- FIG. 6 is a structural block diagram of a terminal device according to an embodiment of the present application, where each unit included in the terminal device is used to execute each step in the embodiment corresponding to FIG. 2.
- each unit included in the terminal device is used to execute each step in the embodiment corresponding to FIG. 2.
- only the parts related to the present embodiment are shown.
- the terminal device includes: a receiving unit 31, a dividing unit 32, a first determining unit 33, and an adjusting unit 34. specifically:
- the receiving unit 31 is configured to: if the preset operation of sending the voice inquiry information to the incoming terminal is detected, receive the voice response information returned by the incoming terminal according to the voice inquiry information.
- the dividing unit 32 is configured to divide the text to be adjusted from the voice content text corresponding to the voice response information.
- the first determining unit 33 is configured to determine reference text from a preset database based on a phone number of the incoming terminal and a content type of the text to be adjusted, where data in the preset database is used to describe the phone number Corresponding relationship between the content type and the reference text.
- the adjusting unit 34 is configured to adjust the text to be adjusted according to the reference text to obtain target information.
- the character type includes at least one of a text type, a letter type, and a digital type.
- the terminal device further includes: an obtaining unit 301, a creating unit 302, and a second determining unit 303. specifically:
- the obtaining unit 301 is configured to acquire an identifier of the voice query information, where the identifier is used to distinguish a character type included in the text corresponding to the voice response information.
- the creating unit 302 is configured to create a state network corresponding to each of the character types, and the state network is configured to reflect that the voice response information corresponding to the character type is converted into an optimal path of the voice content text.
- the second determining unit 303 is configured to determine, according to the character type, a target state network from a preset form, where data in the preset form is used to describe a correspondence between the character type and the target state network.
- the target state network is configured to perform text conversion on the voice response information to obtain a voice content text corresponding to the voice information.
- the dividing unit 32 is specifically configured to identify the number of character types included in the text information, where the number of the character types is greater than or equal to 1; If the number is equal to 1, the text information is divided according to the preset key characters to obtain the text to be adjusted; if the number of the character types is greater than 1, the content of the text information in the text type is separately Divided to obtain the text to be adjusted.
- the content type of the text to be adjusted includes any one of a text type, a letter type, and a numeric type.
- the first determining unit 33 is specifically configured to: obtain target user information from a preset database according to the phone number; and determine, from the target user information, information that matches a content type of the text to be adjusted. As the reference text.
- the adjusting unit 34 is specifically configured to: identify, from the text to be adjusted, target content different from the reference text; if the target content is part of the text to be adjusted The content is replaced by the partial content according to the reference text to obtain target information.
- the solution of the embodiment of the present application receives the voice response information returned by the incoming call terminal according to the voice query information, and the voice content text corresponding to the voice response information, by detecting the preset operation of sending the voice query information to the incoming terminal. Dividing the text to be adjusted, based on the phone number of the incoming call terminal and the content type of the text to be adjusted, determining the reference text from the preset database, and finally adjusting the text to be adjusted according to the reference text to obtain the target information, thereby improving the accuracy of the speech recognition. degree.
- the target state network is determined from the preset form based on the character type, that is, the voice response information corresponding to the character type is determined to be converted into the best path of the voice content text, thereby improving The conversion efficiency of voice information into text information.
- FIG. 7 is a schematic diagram of a terminal device according to another embodiment of the present application.
- the terminal device 7 of this embodiment includes a processor 70, a memory 71, and computer readable instructions 72 stored in the memory 71 and operable on the processor 70, such as a voice recognition program. .
- the processor 70 executes the computer readable instructions 72, the steps in the embodiments of the various speech recognition methods described above are implemented, such as all of the steps shown in FIG. 2.
- the processor 70 when executing the computer readable instructions 72, implements the functions of the various units in the various apparatus embodiments described above, such as the functions of the modules 61 through 67 shown in FIG.
- the computer readable instructions 72 may be partitioned into one or more units, the one or more units being stored in the memory 71 and executed by the processor 70 to complete the application.
- the one or more units may be a series of computer readable instruction instructions segments capable of performing a particular function for describing the execution of the computer readable instructions 72 in the terminal device 7.
- the computer readable instructions 72 may be divided into a receiving unit, a dividing unit, a first determining unit, and an adjusting unit, each of which has a specific function as described above.
- the terminal device 7 may be a computing device such as a desktop computer, a notebook, a palmtop computer, and a cloud server.
- the terminal device may include, but is not limited to, a processor 70 and a memory 71. It will be understood by those skilled in the art that FIG. 7 is only an example of the terminal device 7, and does not constitute a limitation of the terminal device 7, and may include more or less components than those illustrated, or combine some components or different components.
- the terminal device may further include an input/output device, a network access device, a bus, and the like.
- the so-called processor 70 can be a central processing unit (Central Processing Unit, CPU), can also be other general purpose processors, digital signal processors (DSP), application specific integrated circuits (Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
- the general purpose processor may be a microprocessor or the processor or any conventional processor or the like.
- the memory 71 may be an internal storage unit of the terminal device 7, such as a hard disk or a memory of the terminal device 7.
- the memory 71 may also be an external storage device of the terminal device 7, for example, a plug-in hard disk provided on the terminal device 7, a smart memory card (SMC), and a secure digital (SD). Card, flash card, etc. Further, the memory 71 may also include both an internal storage unit of the terminal device 7 and an external storage device.
- the memory 71 is configured to store the computer readable instructions and other programs and data required by the terminal device.
- the memory 71 can also be used to temporarily store data that has been output or is about to be output.
- each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
- the integrated modules/units if implemented in the form of software functional units and sold or used as separate products, may be stored in a computer readable storage medium.
- the present application implements all or part of the processes in the foregoing embodiments, and may also be implemented by computer readable instructions, which may be stored in a computer readable storage medium.
- the computer readable instructions when executed by a processor, may implement the steps of the various method embodiments described above.
- the computer readable instructions comprise computer readable instruction code, which may be in the form of source code, an object code form, an executable file or some intermediate form or the like.
- the computer readable medium can include any entity or device capable of carrying the computer readable instruction code, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, a computer memory, a read only memory (ROM, Read-Only) Memory), random access memory (RAM, Random) Access Memory), electrical carrier signals, telecommunications signals, and software distribution media.
- ROM Read Only memory
- RAM Random Access Memory
- electrical carrier signals telecommunications signals
- telecommunications signals and software distribution media. It should be noted that the content contained in the computer readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in a jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, computer readable media Does not include electrical carrier signals and telecommunication signals.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Telephonic Communication Services (AREA)
Abstract
本申请适用于信息处理技术领域,提供了一种语音识别方法、终端设备及计算机可读存储介质,其中,一种语音识别方法,通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,提高了语音识别的准确程度。
Description
本申请申明享有2017年04月09日递交的申请号为201810309686.0、名称为“一种语音识别方法、终端设备及计算机可读存储介质”中国专利申请的优先权,该中国专利申请的整体内容以参考的方式结合在本申请中。
本申请属于信息处理技术领域,尤其涉及一种语音识别方法、终端设备及计算机可读存储介质。
随着人工成本越来越高,为了降低客服业务部门的人力成本,许多电话客服业务都采用智能语音机器人为用户来电进行服务。
虽然现有的智能语音机器人能够根据用户的语音进行业务办理或者信息发送,但是在对用户的语音进行识别的过程中,如果语音内容中存在容易被混淆的音节,例如,数字的“1”和字母的“E”,则容易造成识别结果不准确的现象。
有鉴于此,本申请实施例提供了一种语音识别方法、终端设备及计算机可读存储介质,以解决现有的语音识别技术中存在识别结果不准确的问题。
本申请实施例的第一方面提供了一种语音识别方法,包括:
若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息;
从所述语音响应信息对应的语音内容文本中划分出待调整文本;
基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系;
根据所述参考文本对所述待调整文本进行调整,得到目标信息。
实施本申请实施例提供的一种语音识别方法、终端设备及计算机可读存储介质具有以下有益效果:
本申请实施例通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,提高了语音识别的准确程度。
图1是本申请实施例提供的一种语音识别方法的实现流程图;
图2是本申请另一实施例提供的一种语音识别方法的实现流程图;
图3是本申请另一实施例提供的一种语音识别方法S12具体实现流程图;
图4是本申请另一实施例提供的一种语音识别方法S13具体实现流程图;
图5是本申请另一实施例提供的一种语音识别方法S14具体实现流程图;
图6是本申请实施例提供的一种终端设备的结构框图;
图7是本申请另一实施例提供的一种终端设备的示意图。
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
本申请实施例通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,解决了现有的语音识别技术中存在的识别结果不准确的问题。
在本申请的所有实施例中,语音识别方法的执行主体为服务器设备。该服务器设备包括但不限于:计算机,或者可以是具有数据处理能力的其他网络设备或通信设备等。图1示出了本申请实施例提供的语音识别方法的实现流程图,详述如下:
S11:若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息。
在步骤S11中,语音询问信息为预先录制到服务器中的语音内容,用于向来电终端对应的用户进行语音询问,其中语音询问的内容可以由运营商根据需求进行定制。语音响应信息为用户在接听到语音询问信息后,通过来电终端向服务器返回的语音信息。
在本实施例中,来电终端可以为移动终端或者非移动终端,如手机、平板电脑或者固定电话等。当来电终端与服务器之间建立了通话链路后,服务器向来电终端发送语音询问信息,再接收用户通过来电终端返回的语音响应信息。或者,当来电终端与服务器之间建立了通话链路后,用户向来电终端发送请求接收语音询问信息的指令,再由服务器根据该指令向来电终端发送语音询问信息,接收用户通过来电终端返回的语音响应信息。
至于何时会检测到向来电终端发送语音询问信息的预设操作,可以包括但不仅限于以下场景。
场景1:当检测到服务器与来电终端之间建立通话链路时,则触发服务器向来电终端发送语音询问信息的操作。
例如,用户通过终端向服务器发送通话请求,服务器根据该通话请求与终端建立通话链路,并触发向来电终端发送语音询问信息的操作,实现向终端发送语音询问信息。
场景2:在来电终端与服务器之间建立了通话链路后,用户向来电终端发送请求接收语音询问信息的指令,则根据该指令,触发服务器向来电终端发送语音询问信息的操作。
例如,终端与服务器之间建立通话链路后,用户在终端上触发请求发送指令,使终端向服务器发送请求接收语音询问信息的指令,进而触发服务器根据该请求接收语音询问信息的指令,向来电终端发送语音询问信息的操作,实现向终端发送语音询问信息。
可以理解的是,在实际应用中,可以通过语音录制的方式,将接收到的来电终端根据语音询问信息返回的语音响应信息,制作成语音响应文件,便于对其进行优化和识别。
S12:从所述语音响应信息对应的语音内容文本中划分出待调整文本。
在步骤S12中,语音响应信息对应的语音内容文本,是通过对语音响应信息进行文字转换得到。待调整文本为语音响应信息对应的语音内容文本中的部分或全部文本。
在本实施例中,可以通过调用预先配置好的待调整文本划分策略,对语音响应信息对应的语音内容文本进行划分,进而从中划分出待调整文本。其中,待调整文本划分策略可以包括多种策略,在实际应用中,可以根据语音响应信息对应的语音内容文本中包含的字符类型,制定对应的待调整文本划分策略。
S13:基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本。
在步骤S13中,预设数据库中的数据用于描述电话号码、内容类型以及参考文本之间的对应关系。内容类型为待调整文本的内容类型,包括:单一字符类型或者混合字符类型,其中,单一字符类型指的是待调整文本的内容由同一种字符组成,混合字符类型指的是待调整文本的内容由至少两种字符组成。
例如,待调整文本的内容用于描述用户姓名,即该待调整文本的内容类型为文字,属于单一字符类型。
再例如,待调整文本的内容用于描述车牌号码,即该待调整文本的内容类型包括字母和数字,或者包括文字、字母以及数字,属于混合字符类型。
在本实施例中,预设数据库为用于存储用户信息的数据库,可以根据不同的电话号码,从预设数据库中确定出不同的用户信息,再根据内容类型从用户信息中确定出参考文本,其中,参考文本的内容类型与待调整文本的内容类型相同。
需要说明的是,存储在预设数据库中的用户信息,均可以通过电话号码搜索得到,其中,用户信息包括与用户相关的所有信息,例如,身份证号码、地址、车牌号码等。
S14:根据所述参考文本对所述待调整文本进行调整,得到目标信息。
在步骤S14中,参考文本是以电话号码与待调整文本的内容类型,作为索引,从预设数据库中查找得到的文本。
在本实施例中,参考文本的内容类型与待调整文本的内容类型相同,即组成参考文本的字符类型与组成待调整文本的字符类型相同。
以待调整文本为车牌号码为例,待调整文本的内容为“京A12345”,待调整文本的内容类型为混合字符类型,也即组成该待调整文本的字符类型包括文字、字母以及数字。基于来电终端的电话号码从预设数据库中确定出用户信息,再根据待调整文本的内容类型,从用户信息中确定出参考文本,由于待调整文本为车牌号码,则该待调整文本的内容类型为混合字符类型,因此,从用户信息中确定出的参考文本也应当为用户信息中的车牌号码,且该车牌号码中也应当包括文字、字母以及数字。
需要说明的是,根据参考文本对所述待调整文本进行调整,可以是基于参考文本中所包含的字符类型,不同字符类型进行分别进行比对和调整,进而得到目标信息。
以待调整文本的内容为“京AE2345”,参考文本的内容为“津A12345”,根据参考文本对待调整文本进行调整,得到的目标信息应当为“津A12345”。
以上可以看出,本申请实施例提供的一种语音识别方法,通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,提高了语音识别的准确程度。
图2示出了本申请另一实施例提供的一种语音识别方法的实现流程图。参见图2所示,相对于图1所述实施例,本实施例提供的一种语音识别方法中还包括S21、S201以及S22,具体详述如下:
进一步地,作为本申请另一实施例,从所述语音响应信息对应的语音内容文本中划分出待调整文本之前,还包括:
S21:获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型。
S22:基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
在本实施例中,由于语音询问信息为预先录制到服务器中的语音内容,因此,该语音内容可以确定,进而可以预测电终端根据该语音询问信息返回的语音响应信息对应的文本中所包含的字符类型。
例如,语音询问信息的内容为用于请求用户通过来电终端输入身份证号码,进而可以确定用户通过来电终端发送的语音响应信息,其对应的语音内容文本中必然是由数字组成的身份证号码,因此可以确定语音响应信息对应的语音内容文本中所包含的字符类型为数字。
再例如,语音询问信息的内容为用于请求用户通过来电终端输入车牌号码,进而可以确定用户通过来电终端发送的语音响应信息,其对应的语音内容文本中必然是由文字、字母以及数字组成的车牌号码,因此可以确定语音响应信息对应的语音内容文本中所包含的字符类型为文字、字母以及数字。
在本实施例中,通过预测来电终端根据语音询问信息返回的语音响应信息对应的文本中,包含的字符类型,可以对不同的语音询问信息配置不同的标识,进而区分语音响应信息对应的文本中所包含的字符类型。
在实际中,将语音信息转换为文本信息的过程中,可以通过调用音频分帧函数,例如,调用移动窗函数对语音文件进行分帧,得到多帧语音,再对每帧语音进行声学特征提取处理,将语音信息中的每一帧波形转化为一个多维向量,进而得到由多个多维向量组成的矩阵,其中,每个多维向量包含了对应语音帧的内容信息,在矩阵中,若干帧语音对应一个状态,每三个状态组合成一个音素,若干个音素组合成一个单词。当确定了语音中每帧语音对应的状态后,就能够根据状态、音素以及单词之间的关系,实现对语音内容的识别。
作为本实施例一种可能实现的方式,字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型,不同字符类型对应不同的目标状态网络。如图2所示,在步骤S22之前,还可以包括与步骤S21并列的步骤S201,在本实施例中,步骤S21与步骤S201执行部分先后。
S201:创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
在步骤S201中,状态网络是由单词级网络展开成音素网络,再将音素网络展开成状态网络。
在本实施例中,创建状态网络时,需要考虑不同字符类型对应的累积转换概率,其中,累积转换概率包括:观察概率、转移概率以及语言概率。
需要说明的是,观察概率指的是每帧语音和每个状态对应的概率,转移概率指的是每个状态转移到自身或转移到下个状态的概率,语言概率是通过语言统计规律得出的概率。观察概率和转移概率都可以通过输入预设的声学模型中得到,语言概率则可以通过输入预设的语言模型中得到,语言模型是使用大量的文本训练出来的,可以利用某门语言本身的统计规律来帮助提升识别正确率。
根据不同的字符类型创建不同的状态网络,使得在对语音信息进行文本转换时,可以根据预先配置好的标识区分字符类型,基于字符类型从预设表单中确定出目标状态网络,也即选择与语音信息对应的文本中所包含的字符类型相对应的目标状态网络,对语音信息进行文本转换,得到语音信息对应的语音内容文本。
本实施例通过获取语音询问信息的标识,以区分语音响应信息对应的文本中所包含的字符类型,通过创建与每种字符类型对应的状态网络,基于字符类型从预设表单中确定出目标状态网络,也即确定字符类型对应的语音响应信息被转换为语音内容文本的最佳路径,从而提高了语音信息转换成文本信息的转换效率。
图3示出了本申请另一实施例提供的一种语音识别方法S12的具体实现流程图。参见图3所示,基于上述各个实施例,本实施例提供的一种语音识别方法中S12包括S121、S122以及S123,具体详述如下:
S121:识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1。
S122:若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本。
S123:若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
在本实施例中,文本信息中包含的字符类型个数,用于反映文本信息的字符类型属于单一字符类型,或者混合字符类型。当文本信息中包含的字符类型个数等于1时,表示文本信息的字符类型属于单一字符类型;当文本信息中包含的字符类型个数大于1时,表示文本信息的字符类型属于混合字符类型。
当文本信息中包含的字符类型个数等于1时,通过识别语音内容文本中是否存在预设关键字,根据预设关键字在语音内容文本中的位置,对语音内容文本进行划分,进而得到待调整文本。
例如,语音内容文本为“广东省深圳市福田区”,预设关键字为“省”、“市”以及“区”,则根据预设关键字在语音内容文本中的位置,对语音内容文本进行划分,进而得到待调整文本为“广东省”、“深圳市”以及“福田区”。
需要说明的是,针对不同的预设关键字可以配置对应的文本组成字数。
例如,我国省份中,名字最长的省份名称的为“黑龙江省”,预设关键字为“省”,对应的文本组成字数为3。
再例如,我国城市中,名字最长的城市名称的为“呼和浩特市”,预设关键字为“市”,对应的文本组成字数为4。
当文本信息中包含的字符类型个数大于1时,将文本信息中字符类型不同的内容进行分别划分,以得到待调整文本。
例如,文本信息为“我的地址是广东省深圳市福田区”,根据预设的关键字符划分文本信息,以得到待调整文本为“广东省”、“深圳市”以及“福田区”。
再例如,文本信息为“我的车牌号码是京AE2345”,将文本信息中字符类型不同的内容进行分别划分,则得到的待调整文本包括“我的车牌号码是京”、“AE”以及“2345”。
通过确定语音信息对应的文本信息中包含的字符类型个数,进而根据字符类型个数的不同,确定不同的待调整文本的划分策略,避免在文本信息的构成元素较为单一时,采用较为复杂的方式进行待调整文本的划分,使得数据处理过程变得更加合理。
图4示出了本申请另一实施例提供的一种语音识别方法S13的具体实现流程图。
在本实施例中,待调整文本的内容类型包括文字类型、字母类型以及数字类型中的任一种字符类型。
参见图4所示,基于上述各个实施例,本实施例提供的一种语音识别方法中S13包括S131以及S132,具体详述如下:
S131:根据所述电话号码从预设数据库中获取目标用户信息。
S132:从所述目标用户信息中,确定出与所述待调整文本的内容类型相匹配的信息作为所述参考文本。
在本实施例中,目标用户信息与电话号码之间存在对应关系,以电话号码为索引可以从预设数据库中查找到对应的目标用户信息,其中,目标用户信息可以包括目标用户的多类信息,例如,身份证号码、车牌号码或者地址等。
需要说明的是,由于目标用户信息包含了用户的多类信息,因此当确定了目标用户信息后,并不能直接从中确定出参考文本。为了能够从目标用户信息中确定参考文本,通过识别测待调整文本的内容类型,再根据测待调整文本的内容类型,从目标用户信息中确定出与其相匹配的信息,并将该信息作为参考文本。
在本实施例中,根据电话号码从预设数据库中获取目标用户信息后,再根据测待调整文本的内容类型,从目标用户信息中确定出与其相匹配的信息,可以避免对目标用户信息中的所有信息进行一一筛选,提高了确定参考文本的速度。
图5示出了本申请另一实施例提供的一种语音识别方法S14的具体实现流程图。参见图5所示,基于上述各个实施例,本实施例提供的一种语音识别方法中S14包括S141以及S142,具体详述如下:
S141:从所述待调整文本中识别出与所述参考文本不同的目标内容。
S142:若所述目标内容为所述待调整文本的部分内容,则根据所述参考文本将所述部分内容进行替换,以得到目标信息。
在本实施例中,目标内容为待调整文本中,与参考文本内容不同的信息。通过比对待调整文本与参考文本,进而从待调整文本中确定出与参考文本不同的目标内容。
在实际中,虽然参考文本作为与待调整文本进行比对的文本,但是当语音响应信息对应的语音内容文本不包含目标用户信息时,则无需使用参考文本对待调整文本进行调整。
通过确定待调整文本中的目标内容是否为待调整文本的部分内容,能够在语音响应信息对应的语音内容文本中不包含目标用户信息时,避免对该语音内容文本对进行调整,进而防止语音转化过程中出现调整错乱或者文本转换错乱的现象。
以上可以看出,本申请实施例提供的一种语音识别方法,通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,提高了语音识别的准确程度。
通过创建与每种字符类型对应的状态网络,基于字符类型从预设表单中确定出目标状态网络,也即确定字符类型对应的语音响应信息被转换为语音内容文本的最佳路径,从而提高了语音信息转换成文本信息的转换效率。
图6示出了本申请实施例提供的一种终端设备的结构框图,该终端设备包括的各单元用于执行图2对应的实施例中的各步骤。具体请参阅图2与图2所对应的实施例中的相关描述。为了便于说明,仅示出了与本实施例相关的部分。
参见图6,所述终端设备包括:接收单元31、划分单元32、第一确定单元33以及调整单元34。具体地:
接收单元31用于,若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息。
划分单元32用于,从所述语音响应信息对应的语音内容文本中划分出待调整文本。
第一确定单元33用于,基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系。
调整单元34用于,根据所述参考文本对所述待调整文本进行调整,得到目标信息。
进一步地,作为本实施例一种可能实现的方式,字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型。终端设备还包括:获取单元301、创建单元302以及第二确定单元303。具体地:
获取单元301用于,获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型。
创建单元302用于,创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
第二确定单元303用于,基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
进一步地,作为本实施例一种可能实现的方式,划分单元32具体用于,识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1;若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本;若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
作为本实施例一种可能实现的方式,待调整文本的内容类型包括文字类型、字母类型以及数字类型中的任一种字符类型。
进一步地,第一确定单元33具体用于,根据所述电话号码从预设数据库中获取目标用户信息;从所述目标用户信息中,确定出与所述待调整文本的内容类型相匹配的信息作为所述参考文本。
作为本实施例一种可能实现的方式,调整单元34具体用于,从所述待调整文本中识别出与所述参考文本不同的目标内容;若所述目标内容为所述待调整文本的部分内容,则根据所述参考文本将所述部分内容进行替换,以得到目标信息。
以上可以看出,本申请实施例的方案通过在检测到向来电终端发送语音询问信息的预设操作时,接收来电终端根据语音询问信息返回的语音响应信息,从语音响应信息对应的语音内容文本中划分出待调整文本,基于来电终端的电话号码与待调整文本的内容类型,从预设数据库中确定参考文本,最后根据参考文本对待调整文本进行调整,得到目标信息,提高了语音识别的准确程度。
通过创建与每种字符类型对应的状态网络,基于字符类型从预设表单中确定出目标状态网络,也即确定字符类型对应的语音响应信息被转换为语音内容文本的最佳路径,从而提高了语音信息转换成文本信息的转换效率。
图7是本申请另一实施例提供的一种终端设备的示意图。如图7所示,该实施例的终端设备7包括:处理器70、存储器71以及存储在所述存储器71中并可在所述处理器70上运行的计算机可读指令72,例如语音识别程序。所述处理器70执行所述计算机可读指令72时实现上述各个语音识别方法实施例中的步骤,例如图2所示的所有步骤。或者,所述处理器70执行所述计算机可读指令72时实现上述各装置实施例中各单元的功能,例如图6所示模块61至67功能。
示例性的,所述计算机可读指令72可以被分割成一个或多个单元,所述一个或者多个单元被存储在所述存储器71中,并由所述处理器70执行,以完成本申请。所述一个或多个单元可以是能够完成特定功能的一系列计算机可读指令指令段,该指令段用于描述所述计算机可读指令72在所述终端设备7中的执行过程。例如,所述计算机可读指令72可以被分割成接收单元、划分单元、第一确定单元以及调整单元各单元具体功能如上所述。
所述终端设备7可以是桌上型计算机、笔记本、掌上电脑及云端服务器等计算设备。所述终端设备可包括,但不仅限于,处理器70、存储器71。本领域技术人员可以理解,图7仅仅是终端设备7的示例,并不构成对终端设备7的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件,例如所述终端设备还可以包括输入输出设备、网络接入设备、总线等。
所称处理器70可以是中央处理单元(Central
Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application
Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
所述存储器71可以是所述终端设备7的内部存储单元,例如终端设备7的硬盘或内存。所述存储器71也可以是所述终端设备7的外部存储设备,例如所述终端设备7上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital, SD)卡,闪存卡(Flash Card)等。进一步地,所述存储器71还可以既包括所述终端设备7的内部存储单元也包括外部存储设备。所述存储器71用于存储所述计算机可读指令以及所述终端设备所需的其他程序和数据。所述存储器71还可以用于暂时地存储已经输出或者将要输出的数据。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的模块/单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请实现上述实施例方法中的全部或部分流程,也可以通过计算机可读指令来指令相关的硬件来完成,所述的计算机可读指令可存储于一计算机可读存储介质中,该计算机可读指令在被处理器执行时,可实现上述各个方法实施例的步骤。。其中,所述计算机可读指令包括计算机可读指令代码,所述计算机可读指令代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机可读指令代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only
Memory)、随机存取存储器(RAM,Random
Access Memory)、电载波信号、电信信号以及软件分发介质等。需要说明的是,所述计算机可读介质包含的内容可以根据司法管辖区内立法和专利实践的要求进行适当的增减,例如在某些司法管辖区,根据立法和专利实践,计算机可读介质不包括电载波信号和电信信号。
以上所述实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围,均应包含在本申请的保护范围之内。
Claims (20)
- 一种语音识别方法,其特征在于,包括:若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息;从所述语音响应信息对应的语音内容文本中划分出待调整文本;基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系;根据所述参考文本对所述待调整文本进行调整,得到目标信息。
- 根据权利要求1所述的语音识别方法,其特征在于,所述从所述语音响应信息对应的语音内容文本中划分出待调整文本之前,还包括:获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型;基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
- 根据权利要求2所述的语音识别方法,其特征在于,所述字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型;基于所述字符类型,从预设表单中确定出目标状态网络之前,还包括:创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
- 根据权利要求1所述的语音识别方法,其特征在于,所述从所述语音信息对应的文本信息中划分出待调整文本,包括:识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1;若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本;若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
- 根据权利要求1所述的语音识别方法,其特征在于,所述待调整文本的内容类型包括文字类型、字母类型以及数字类型中的任一种字符类型;所述基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,包括:根据所述电话号码从预设数据库中获取目标用户信息;从所述目标用户信息中,确定出与所述待调整文本的内容类型相匹配的信息作为所述参考文本。
- 根据权利要求1至5任一项所述的语音识别方法,其特征在于,所述根据所述参考文本对所述待调整文本进行调整,得到目标信息,包括:从所述待调整文本中识别出与所述参考文本不同的目标内容;若所述目标内容为所述待调整文本的部分内容,则根据所述参考文本将所述部分内容进行替换,以得到目标信息。
- 一种语音识别装置,其特征在于,包括:接收单元,用于若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息;划分单元,用于从所述语音响应信息对应的语音内容文本中划分出待调整文本;第一确定单元,用于基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系;调整单元,用于根据所述参考文本对所述待调整文本进行调整,得到目标信息。
- 如权利要求7所述的装置,其特征在于,所述语音识别装置还包括:获取单元,用于获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型;第二确定单元,用于基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
- 如权利要求8所述的装置,其特征在于,所述字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型;所述语音识别装置还包括:创建单元,用于创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
- 如权利要求7所述的装置,其特征在于,划分单元具体用于,识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1;若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本;若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
- 如权利要求7所述的装置,其特征在于,所述待调整文本的内容类型包括文字类型、字母类型以及数字类型中的任一种字符类型;第一确定单元具体用于,根据所述电话号码从预设数据库中获取目标用户信息;从所述目标用户信息中,确定出与所述待调整文本的内容类型相匹配的信息作为所述参考文本。
- 一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,其特征在于,所述处理器执行所述计算机可读指令时实现如下步骤:若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息;从所述语音响应信息对应的语音内容文本中划分出待调整文本;基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系;根据所述参考文本对所述待调整文本进行调整,得到目标信息。
- 如权利要求12所述的终端设备,其特征在于,所述从所述语音响应信息对应的语音内容文本中划分出待调整文本之前,还包括:获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型;基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
- 如权利要求13所述的终端设备,其特征在于,所述字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型;基于所述字符类型,从预设表单中确定出目标状态网络之前,还包括:创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
- 如权利要求12所述的终端设备,所述从所述语音信息对应的文本信息中划分出待调整文本,包括:识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1;若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本;若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
- 如权利要求12所述的终端设备,其特征在于,所述待调整文本的内容类型包括文字类型、字母类型以及数字类型中的任一种字符类型;所述基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,包括:根据所述电话号码从预设数据库中获取目标用户信息;从所述目标用户信息中,确定出与所述待调整文本的内容类型相匹配的信息作为所述参考文本。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,其特征在于,所述计算机可读指令被处理器执行时实现如下步骤:若检测到向来电终端发送语音询问信息的预设操作,则接收所述来电终端根据所述语音询问信息返回的语音响应信息;从所述语音响应信息对应的语音内容文本中划分出待调整文本;基于所述来电终端的电话号码与所述待调整文本的内容类型,从预设数据库中确定参考文本,所述预设数据库中的数据用于描述所述电话号码、所述内容类型以及所述参考文本之间的对应关系;根据所述参考文本对所述待调整文本进行调整,得到目标信息。
- 如权利要求17所述的计算机可读存储介质,其特征在于,所述从所述语音响应信息对应的语音内容文本中划分出待调整文本之前,还包括:获取所述语音询问信息的标识,所述标识用于区分所述语音响应信息对应的文本中所包含的字符类型;基于所述字符类型,从预设表单中确定出目标状态网络,所述预设表单中的数据用于描述所述字符类型与所述目标状态网络之间的对应关系;所述目标状态网络用于对所述语音响应信息进行文本转换,以得到所述语音信息对应的语音内容文本。
- 如权利要求18所述的计算机可读存储介质,其特征在于,所述字符类型包括:文字类型、字母类型以及数字类型中的至少一种字符类型;基于所述字符类型,从预设表单中确定出目标状态网络之前,还包括:创建与每种所述字符类型对应的状态网络,所述状态网络用于反映所述字符类型对应的语音响应信息被转换为所述语音内容文本的最佳路径。
- 根据权利要求17所述的计算机可读存储介质,其特征在于,所述从所述语音信息对应的文本信息中划分出待调整文本,包括:识别所述文本信息中包含的字符类型个数,所述字符类型个数大于或等于1;若所述字符类型个数等于1,则根据预设的关键字符划分所述文本信息,以得到所述待调整文本;若所述字符类型个数大于1,则将所述文本信息中字符类型不同的内容进行分别划分,以得到所述待调整文本。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810309686.0 | 2018-04-09 | ||
| CN201810309686.0A CN108682421B (zh) | 2018-04-09 | 2018-04-09 | 一种语音识别方法、终端设备及计算机可读存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019196238A1 true WO2019196238A1 (zh) | 2019-10-17 |
Family
ID=63800836
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/096263 Ceased WO2019196238A1 (zh) | 2018-04-09 | 2018-07-19 | 一种语音识别方法、终端设备及计算机可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN108682421B (zh) |
| WO (1) | WO2019196238A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111782172A (zh) * | 2020-06-24 | 2020-10-16 | 大众问问(北京)信息科技有限公司 | 一种信息展示方法和装置 |
| CN112541774A (zh) * | 2020-12-08 | 2021-03-23 | 四川众信佳科技发展有限公司 | Ai质检方法,装置,系统,电子设备及存储介质 |
| CN114944149A (zh) * | 2022-04-15 | 2022-08-26 | 科大讯飞股份有限公司 | 语音识别方法、语音识别设备及计算机可读存储介质 |
| CN115171695A (zh) * | 2022-06-29 | 2022-10-11 | 东莞爱源创科技有限公司 | 语音识别方法、装置、电子设备和计算机可读介质 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110010131B (zh) * | 2019-04-04 | 2022-01-04 | 深圳市语芯维电子有限公司 | 一种语音信息处理的方法和装置 |
| CN111143525A (zh) * | 2019-12-17 | 2020-05-12 | 广东广信通信服务有限公司 | 车辆信息获取方法、装置和智能移车系统 |
| CN111667835A (zh) * | 2020-06-01 | 2020-09-15 | 马上消费金融股份有限公司 | 语音识别方法、活体检测方法、模型训练方法及装置 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090319272A1 (en) * | 2008-06-18 | 2009-12-24 | International Business Machines Corporation | Method and system for voice ordering utilizing product information |
| US20100161315A1 (en) * | 2008-12-24 | 2010-06-24 | At&T Intellectual Property I, L.P. | Correlated call analysis |
| CN105810197A (zh) * | 2014-12-30 | 2016-07-27 | 联想(北京)有限公司 | 语音处理方法、语音处理装置和电子设备 |
| CN106331392A (zh) * | 2016-08-19 | 2017-01-11 | 美的集团股份有限公司 | 控制方法及控制装置 |
| CN107437416A (zh) * | 2017-05-23 | 2017-12-05 | 阿里巴巴集团控股有限公司 | 一种基于语音识别的咨询业务处理方法及装置 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2323693B (en) * | 1997-03-27 | 2001-09-26 | Forum Technology Ltd | Speech to text conversion |
| CN106340293B (zh) * | 2015-07-06 | 2019-11-29 | 无锡天脉聚源传媒科技有限公司 | 一种音频数据识别结果的调整方法及装置 |
| CN105895103B (zh) * | 2015-12-03 | 2020-01-17 | 乐融致新电子科技(天津)有限公司 | 一种语音识别方法及装置 |
| CN105869642B (zh) * | 2016-03-25 | 2019-09-20 | 海信集团有限公司 | 一种语音文本的纠错方法及装置 |
| CN106328145B (zh) * | 2016-08-19 | 2019-10-11 | 北京云知声信息技术有限公司 | 语音修正方法及装置 |
| CN107045496B (zh) * | 2017-04-19 | 2021-01-05 | 畅捷通信息技术股份有限公司 | 语音识别后文本的纠错方法及纠错装置 |
| CN107293296B (zh) * | 2017-06-28 | 2020-11-20 | 百度在线网络技术(北京)有限公司 | 语音识别结果纠正方法、装置、设备及存储介质 |
| CN107731229B (zh) * | 2017-09-29 | 2021-06-08 | 百度在线网络技术(北京)有限公司 | 用于识别语音的方法和装置 |
-
2018
- 2018-04-09 CN CN201810309686.0A patent/CN108682421B/zh active Active
- 2018-07-19 WO PCT/CN2018/096263 patent/WO2019196238A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090319272A1 (en) * | 2008-06-18 | 2009-12-24 | International Business Machines Corporation | Method and system for voice ordering utilizing product information |
| US20100161315A1 (en) * | 2008-12-24 | 2010-06-24 | At&T Intellectual Property I, L.P. | Correlated call analysis |
| CN105810197A (zh) * | 2014-12-30 | 2016-07-27 | 联想(北京)有限公司 | 语音处理方法、语音处理装置和电子设备 |
| CN106331392A (zh) * | 2016-08-19 | 2017-01-11 | 美的集团股份有限公司 | 控制方法及控制装置 |
| CN107437416A (zh) * | 2017-05-23 | 2017-12-05 | 阿里巴巴集团控股有限公司 | 一种基于语音识别的咨询业务处理方法及装置 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111782172A (zh) * | 2020-06-24 | 2020-10-16 | 大众问问(北京)信息科技有限公司 | 一种信息展示方法和装置 |
| CN111782172B (zh) * | 2020-06-24 | 2024-03-12 | 大众问问(北京)信息科技有限公司 | 一种信息展示方法和装置 |
| CN112541774A (zh) * | 2020-12-08 | 2021-03-23 | 四川众信佳科技发展有限公司 | Ai质检方法,装置,系统,电子设备及存储介质 |
| CN114944149A (zh) * | 2022-04-15 | 2022-08-26 | 科大讯飞股份有限公司 | 语音识别方法、语音识别设备及计算机可读存储介质 |
| CN115171695A (zh) * | 2022-06-29 | 2022-10-11 | 东莞爱源创科技有限公司 | 语音识别方法、装置、电子设备和计算机可读介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN108682421A (zh) | 2018-10-19 |
| CN108682421B (zh) | 2023-04-14 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019196238A1 (zh) | 一种语音识别方法、终端设备及计算机可读存储介质 | |
| CN111566638B (zh) | 向应用编程接口添加描述性元数据以供智能代理使用 | |
| CN110532107B (zh) | 接口调用方法、装置、计算机设备及存储介质 | |
| CN107809550A (zh) | 调整业务语音播放顺序的方法及设备 | |
| WO2023272616A1 (zh) | 一种文本理解方法、系统、终端设备和存储介质 | |
| US11734341B2 (en) | Information processing method, related device, and computer storage medium | |
| CN112786041B (zh) | 语音处理方法及相关设备 | |
| CN111177358B (zh) | 意图识别方法、服务器及存储介质 | |
| WO2021063089A1 (zh) | 规则匹配方法、规则匹配装置、存储介质及电子设备 | |
| CN114242047B (zh) | 一种语音处理方法、装置、电子设备及存储介质 | |
| WO2022257452A1 (zh) | 表情回复方法、装置、设备及存储介质 | |
| CN111368549A (zh) | 一种支持多种服务的自然语言处理方法、装置及系统 | |
| WO2020103447A1 (zh) | 视频信息链式存储方法、装置、计算机设备及存储介质 | |
| CN108520471A (zh) | 重叠社区发现方法、装置、设备及存储介质 | |
| CN112562654A (zh) | 一种音频分类方法及计算设备 | |
| CN110597765A (zh) | 一种大零售呼叫中心异构数据源数据处理方法及装置 | |
| CN115691474A (zh) | 一种说话人音频分离方法、终端设备及存储介质 | |
| US9264870B2 (en) | Mobile terminal, server and calling method based on cloud contact list | |
| CN110443291B (zh) | 一种模型训练方法、装置及设备 | |
| US20210209172A1 (en) | Name matching using enhanced name keys | |
| CN115470800A (zh) | 一种机器人对话方法及装置 | |
| WO2021244424A1 (zh) | 中心词提取方法、装置、设备及存储介质 | |
| CN117059096B (zh) | 车载语义结果的处理方法及装置 | |
| WO2021159668A1 (zh) | 机器人对话方法、装置、计算机设备和存储介质 | |
| WO2021103594A1 (zh) | 一种默契度检测方法、设备、服务器及可读存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18914712 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18914712 Country of ref document: EP Kind code of ref document: A1 |