WO2017219495A1 - 语音识别方法及系统 - Google Patents

语音识别方法及系统 Download PDF

Info

Publication number
WO2017219495A1
WO2017219495A1 PCT/CN2016/097466 CN2016097466W WO2017219495A1 WO 2017219495 A1 WO2017219495 A1 WO 2017219495A1 CN 2016097466 W CN2016097466 W CN 2016097466W WO 2017219495 A1 WO2017219495 A1 WO 2017219495A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
recognition result
user
speech recognition
module
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/097466
Other languages
English (en)
French (fr)
Inventor
林瑞华
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yulong Computer Telecommunication Scientific Shenzhen Co Ltd
Original Assignee
Yulong Computer Telecommunication Scientific Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yulong Computer Telecommunication Scientific Shenzhen Co Ltd filed Critical Yulong Computer Telecommunication Scientific Shenzhen Co Ltd
Publication of WO2017219495A1 publication Critical patent/WO2017219495A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/28Constructional details of speech recognition systems
    • G10L15/32Multiple recognisers used in sequence or in parallel; Score combination systems therefor, e.g. voting systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems

Definitions

  • the present invention relates to the field of speech recognition technologies, and in particular, to a speech recognition method and system based on geographic location assistance.
  • Mandarin is widely used as the language of communication.
  • people in different regions still have differences in dialects. Therefore, influenced by dialects, Mandarin in different localities have different characteristics, but people in the same region speak the speed of Mandarin. Voice and semantics are similar.
  • the existing voice input recognizes the output result by taking the user voice data and identifying the output result according to the voice recognition algorithm, and does not make reference to the geographic location and other data of the user.
  • the recognition rate for some speech recognition with local accent and partial local dialect is not high.
  • a voice recognition method is applied to an electronic device, and the method includes:
  • the first speech recognition method is a large vocabulary speech recognition method based on a preset model
  • the second speech recognition method is a speech recognition method based on an auxiliary speech data packet.
  • the voice recognition method based on the auxiliary voice data packet includes:
  • Receiving the voice information acquiring current geographic location information of the user
  • the method further comprises:
  • a plurality of geographically-based voice data packets are preset and stored in the electronic device or in a server connected to the electronic device.
  • the method before the corresponding auxiliary voice data packet is called according to the geographical location information, the method further includes:
  • a corresponding auxiliary voice data packet is jointly determined based on the voice type and the geographic location information.
  • the method further comprises:
  • Receiving the voice information acquiring current geographic location information and historical geographic location information of the user;
  • the invoked auxiliary voice data packet is determined based on the historical geographical location information and the current geographic location information.
  • the method further comprises:
  • preset rules Updating the preset rules in combination with the obtained user feedback information, where the preset rules include:
  • the display mode includes a displayed time or a displayed position.
  • the updating the pre-set rules comprises:
  • the weight value or the recognition score value corresponding to the voice recognition result is increased, and/or the weight value or the recognition score value corresponding to the voice recognition result not selected by the user is reduced.
  • a speech recognition system running in an electronic device, the system comprising:
  • An obtaining module configured to obtain voice information input by a user
  • a first identification module configured to identify the voice information to obtain a first voice recognition result
  • a second identification module configured to identify the voice information to obtain a second voice recognition result
  • a display module configured to display the first voice recognition result and the second voice recognition result according to a preset rule.
  • the first speech recognition module is a large vocabulary speech recognition module based on a preset model
  • the second speech recognition module is a speech recognition module based on an auxiliary speech data packet.
  • the second identification module includes:
  • Calling a submodule configured to invoke a corresponding auxiliary voice data packet according to the geographic location information
  • the second identification module is configured to obtain the second voice recognition result by identifying the voice information according to the auxiliary voice data packet.
  • the system further comprises:
  • a setting module configured to preset a plurality of geographical location-based voice data packets, and store the voice data packets in the electronic device or in a server connected to the electronic device.
  • the system further comprises a determining sub-module:
  • a corresponding auxiliary voice data packet is jointly determined based on the voice type and the geographic location information.
  • the acquiring module is further configured to: when receiving the voice information, acquire current geographic location information and historical geographic location information of the user; and
  • the calling sub-module is further configured to determine the invoked auxiliary voice data packet according to the historical geographical location information and the current geographic location information.
  • the system further comprises:
  • an update module configured to update the preset rule in combination with the obtained user feedback information, where the preset rule is set by the setting module, including:
  • the display mode includes a displayed time or a displayed position.
  • the updating module updating the preset rules includes:
  • the weight value or the recognition score value corresponding to the voice recognition result is increased, and/or the weight value or the recognition score value corresponding to the voice recognition result not selected by the user is reduced.
  • the voice recognition method and system of the present invention can establish multiple auxiliary voice data packets according to the characteristics of different areas of Mandarin, and call different auxiliary voice data packets for users in different geographical locations, which can be effective. Reduce the variety of speech recognition libraries and improve speech recognition rates.
  • FIG. 1 is a schematic diagram of a hardware architecture of a preferred embodiment of an electronic device for performing a speech recognition system of the present invention.
  • FIG. 2 is a flow chart of a preferred embodiment of the speech recognition method of the present invention.
  • FIG. 3 is a flow chart of a preferred embodiment of a voice recognition method based on an auxiliary voice data packet of the present invention Figure.
  • Figure 4 is a functional block diagram of a first embodiment of the speech recognition system of the present invention.
  • Figure 5 is a functional block diagram of a second embodiment of the speech recognition system of the present invention.
  • Speech recognition system 10 Storage unit 20 Display unit 30 Processing unit 40 Voice receiving unit 50 Acquisition module 100 First identification module 102 Second identification module 104 Calling a submodule 1040 Download submodule 1042 Determine submodule 1044 Display module 106 Setting module 108 Update module 110
  • the electronic device 1 is a schematic diagram of a hardware architecture of a preferred embodiment of an electronic device for performing a speech recognition system of the present invention. As shown in the hardware architecture diagram, the electronic device 1 includes a speech recognition system 10. The electronic device 1 further includes a storage unit 20, a display unit 30, a processing unit 40, and voice receiving Unit 50.
  • the speech recognition method of the present invention is implemented by the speech recognition system 10 in the electronic device 1.
  • the electronic device 1 includes an electronic device capable of automatically performing numerical calculation and/or information processing according to an instruction set or stored in advance, and the hardware includes but not limited to a microprocessor and an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), Field Programmable Gate Array (FPGA), Digital Signal Processor (DSP), embedded devices, etc.
  • the electronic device 1 may also include a user device.
  • the user equipment includes, but is not limited to, any electronic product that can interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device, such as a personal computer, a tablet computer, a smart phone, and a personal digital device.
  • the network where the user equipment is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), and the like.
  • VPN virtual private network
  • user equipment is only an example, and other existing or future user equipments, such as may be adapted to the present invention, are also included in the scope of the present invention and are incorporated herein by reference.
  • the voice recognition system 10 is configured to acquire voice information input by the user when the user inputs voice information, and utilize a large vocabulary voice recognition method based on a preset model (for example, based on a hidden Markov model) Large vocabulary speech recognition method) identifying the input speech information to obtain a first speech recognition result, and using a speech recognition method based on the auxiliary speech data packet (for example, calling the geographic location information according to the current geographic location information of the user) The corresponding auxiliary voice data packet is identified to obtain a second voice recognition result.
  • the speech recognition system 10 obtains an optimal recognition result by comparing the first speech recognition result and the second speech recognition result, which not only improves the speech recognition rate but also improves the user experience.
  • the storage unit 20 is configured to store software programs and data installed in the electronic device 1, such as the speech recognition system 10.
  • the storage unit 20 may be an internal storage unit of the electronic device 1, such as a hard disk or a memory of the electronic device 1.
  • the storage unit 20 may also be an external storage device of the electronic device 1, such as a plug-in hard disk, a smart media card (SMC), and a secure digital card (Secure Digital) on the electronic device 1. Card, SD), flash memory card and other storage units.
  • the storage unit 20 may also include an internal storage unit of the electronic device 1 or an external storage device.
  • auxiliary voice data packets and voice information corresponding to the plurality of auxiliary voice data packets are pre-stored in the storage unit 20.
  • the auxiliary voice data packet may be a voice data packet based on a geographical location, and correspondingly, the storage unit 20 stores voice information having the voice feature of the geographic location.
  • the geographical location is divided by the city.
  • it may be subdivided into an area below the prefecture, for example, divided by a county-level city or divided by a set area.
  • the geographical location-based voice data packet stored in the storage unit 20 further includes voice data packets based on dialect and geographic location in other embodiments. Voice packets based on accents and geography.
  • voice packets based on dialects and geographic locations may include: Cantonese_Hong Kong, Cantonese_Guangzhou, Minnanese_Quanzhou, Minnanese_Xiamen.
  • Voice packets based on accents and geography can include: accent_Fujian, accent_Guangzhou. It should be noted that the voice data package based on the accent and the geographical location includes, but is not limited to, the initials of the initials and the finals.
  • the display unit 30 is configured to display a graphical user interface (GUI), where the graphical user interface may include multiple application icons and/or multiple virtual buttons, and the application icons and The virtual button may be representative of various functions that the electronic device 1 can provide.
  • GUI graphical user interface
  • the voice input icon represents a function that the electronic device 1 can provide voice input
  • the geographic location selection list button represents that the electronic device 1 can provide a selection.
  • the functionality of the geographic location, as well as the text input box keys, represent the functionality by which the electronic device 1 can provide an input geographic location.
  • the display unit 30 may be, but not limited to, a display unit having a touch function such as a touch display screen. Therefore, in addition to viewing the application icon and/or the virtual button displayed by the electronic device 1 through the display unit 30, the user may input a function instruction through the display unit 30, for example, running the application icon corresponding to The instructions of the application, or the instructions that activate the virtual button to initiate the corresponding function.
  • the processing unit 40 is one or more central processors (Central) Processing unit, CPU), microprocessor or other digital processing chip.
  • the processing unit 40 is operative to execute software program code or operational data, such as to perform the speech recognition system 10.
  • the processing unit 40 receives the voice information input by the user, and acquires the current geographic location information of the user, and combines the large vocabulary speech recognition based on the preset model when performing voice recognition (for example, based on hidden marin Large vocabulary speech recognition method of Kraft model, or speech recognition method based on artificial neural network model) and speech recognition based on auxiliary speech data packet (for example, speech recognition based on geographic location-based auxiliary speech data packet) respectively output first Identifying the result and the second recognition result, dynamically adjusting the weight of the large vocabulary speech recognition based on the preset model and the speech recognition based on the auxiliary speech data packet according to the selection made by the user comparing the first recognition result and the second recognition result Improve the accuracy of speech recognition.
  • voice recognition for example, based on hidden marin Large vocabulary speech recognition method of
  • the processing unit 40 is communicatively coupled to the voice recognition system 10, the storage unit 20, the display unit 30, and the voice input unit 50.
  • the communication can be implemented via a Serial Peripheral Interface Bus (USB) or other communication path or protocol.
  • USB Serial Peripheral Interface Bus
  • the voice input unit 50 is configured to input voice information of the user.
  • the display unit 30 includes, but is not limited to, a microphone.
  • FIG. 2 is a flow chart of a preferred embodiment of the speech recognition method of the present invention.
  • the order of the steps in the flowchart may be changed according to different requirements, and some steps may be omitted.
  • S100 Acquire voice information input by a user.
  • the user can directly input voice through the voice receiving unit 50 of the electronic device 1, and the voice recognition system 10 acquires voice information according to the content of the voice input by the user.
  • the display unit 30 of the electronic device 1 provides a graphical user interface including a voice input icon, the voice recognition system 10 when the user clicks on the voice input icon
  • the voice information input by the user is acquired by the voice receiving unit 50.
  • S102 Identify the voice information by using a first voice recognition method to obtain a first recognition result, and identify the voice information by using a second voice recognition method to obtain a second recognition result.
  • the first voice recognition method identifies a large vocabulary voice recognition method based on a preset model
  • the second voice recognition method may be a voice recognition method based on the auxiliary voice data packet. That is, the voice recognition method based on the auxiliary voice data packet is used to assist the voice recognition based on the large vocabulary voice recognition method based on the preset model.
  • the voice recognition method based on the auxiliary voice data packet may be It is a speech recognition method based on a geographically established auxiliary voice packet.
  • the speech recognition system 10 may first perform the first speech recognition method to identify the speech information, and then perform the second speech recognition method to identify the second speech information.
  • the speech recognition system 10 may perform the first speech recognition method and the second speech recognition method in parallel to identify the speech information, respectively.
  • the voice information is identified by the large vocabulary speech recognition method based on the preset model
  • the voice information is recognized by using the voice recognition method based on the auxiliary voice data packet, that is, the voice recognition system 10 is first.
  • the thread runs the large vocabulary speech recognition method based on the preset model to identify the speech information, and in parallel, a second thread runs the speech recognition based on the auxiliary speech data packet to identify the speech information.
  • the large vocabulary speech recognition method based on the preset model refers to a speech recognition library established according to standard Mandarin, and any user can call the speech recognition library to identify according to standard Mandarin.
  • Large vocabulary speech recognition based on a preset model does not take into account the influence of dialects and geographic locations and/or accents and geographic locations.
  • the large vocabulary speech recognition method based on the preset model may adopt a speech recognition method in the prior art, and learn and train through a plurality of models established in advance to recognize the user's voice, and convert the voice information into text information.
  • auxiliary voice recognition method for convenience of description
  • the voice recognition method based on the auxiliary voice data packet considers the influence of dialect and geographical position and/or accent and geographical position, and needs to establish a geography based on training and learning in advance.
  • Location voice packets Please refer to FIG. 3 and the corresponding description regarding the geographical location based speech recognition method.
  • the preset rule may be that the voice recognition system 10 pre-allocates a first weight for the first voice recognition result, and pre-allocates a second weight for the second voice recognition result, according to the weight.
  • the size of the value determines how the speech recognition result corresponding to the weight value is displayed.
  • the sum of the first weight value and the second weight value may be a fixed number, for example, an integer of one.
  • the first weight value preset by the voice recognition system 10 is greater than the second weight value, that is, the voice recognition system 10 assigns a weight value to the first voice recognition method that is greater than the second voice recognition method. Weights.
  • the preset rule may also be that the voice recognition system 10 And setting a first recognition score for the first voice recognition result, setting a second recognition score for the second voice recognition result, and determining a display manner of the voice recognition result corresponding to the recognition score according to the size of the recognition score.
  • the first recognition score value preset by the voice recognition system 10 is greater than the second recognition score value.
  • the manner in which the speech recognition result is displayed includes, but is not limited to, the displayed time and/or the displayed position.
  • the rule set by the voice recognition system 10 is to assign a weight to the voice recognition result, and when the first weight value set in advance is greater than the preset second weight value, the display unit 30 of the electronic device 1 may be The first speech recognition result corresponding to the weight value is displayed in the first position, such as the upper half of the user interface provided by the display unit 30; when the preset first weight value is less than the preset second weight value And displaying the first voice recognition result corresponding to the small weight value in the second position, such as the lower half of the user interface provided by the display unit 30.
  • the first voice recognition result is displayed on the display unit 30 of the electronic device 1 after a preset time (for example, after 2 seconds)
  • a second speech recognition result is displayed on the display unit 30 of the electronic device 1.
  • the voice recognition method further includes: updating the preset rule in combination with the acquired user feedback information.
  • the user feedback information can be obtained according to a user's operation. For example, if the user selects the first speech recognition result, the user feedback information acquired by the speech recognition system 10 indicates that the optimal speech recognition result is obtained by using the first speech recognition method. If the user selects the second speech recognition result, the user feedback information acquired by the speech recognition system 10 indicates that the best speech recognition result is obtained by using the second speech recognition method.
  • the updating the preset rule may be adjusting a preset weight value or adjusting a preset recognition score value.
  • the voice recognition system 10 increases the weight value or the recognition score value corresponding to the voice recognition result according to the voice recognition result selected by the user, and/or the weight value or the identification corresponding to the voice recognition result that the user does not select.
  • the fractional value is reduced. For example, when the obtained user feedback information is that the first speech recognition result is selected, the first weight value or the first recognition score value corresponding to the first speech recognition result is increased, and/or the second speech recognition result is to be corresponding.
  • the second weight value or the second identification score value is decreased.
  • the obtained user feedback information is the second speech recognition result selected, it will correspond to The second weight value or the second recognition score value of the second voice recognition result becomes larger, and/or the first weight value or the first recognition score value corresponding to the first voice recognition result is decreased.
  • the above-mentioned weight value or the value of the point value may be increased or decreased according to a preset ratio or value.
  • FIG. 3 Please refer to FIG. 3 as a flowchart of a preferred embodiment of a speech recognition method based on an auxiliary voice packet.
  • the order of the steps in the flowchart may be changed according to different requirements, and some steps may be omitted.
  • S1020 When receiving the voice information of the user, obtain the current geographic location information of the user.
  • the voice recognition system 10 acquires the geographical location information of the electronic device 1 by using the positioning module and/or the network connection module built in the electronic device 1 .
  • the positioning module includes, but is not limited to, a Global Positioning System (GPS).
  • the network connection module includes, but is not limited to, The 3rd Generation Telecommunication (3G), General Packet Radio Service (GPRS), and Wireless Fidelity (Wi -Fi).
  • the geographical location information of the electronic device 1 is considered to be the geographical location information of the current location of the user.
  • the voice recognition system 10 can also determine the current geographic location information of the user by receiving an instruction set by the user and according to an instruction set by the user.
  • the electronic device 1 is provided with a location selection list, which includes the names of all cities in China.
  • the user selects the geographical location information corresponding to the user input voice information by triggering the location selection list.
  • the electronic device 1 is provided with a text input box, and the user inputs the current geographic location information in the corresponding interface by activating the text input box function.
  • the electronic device 1 calls a corresponding auxiliary voice data packet from the storage unit 20 according to the geographical location information.
  • the storage unit 20 pre-stores the auxiliary voice data packet and the voice information having the geographical location voice feature included in the auxiliary voice data packet.
  • the speech recognition system 10 invokes an auxiliary speech data packet that identifies the Guangdong speech feature.
  • the voice recognition system 10 is acquiring the current geographic location information of the user. And downloading the auxiliary voice data packet from a server communicatively connected to the electronic device 1.
  • the communication connection can be a wireless communication connection.
  • the auxiliary voice data packet is obtained and trained by the user in advance and deployed to the server, and the voice recognition system 10 may request the server to send an auxiliary voice data packet corresponding to the geographical location information through a network.
  • the voice recognition system 10 uses the second voice recognition method to identify the voice information to obtain the second voice recognition result.
  • the speech recognition system 10 calls the corresponding auxiliary speech data packet according to the geographical location information,
  • the S1022 may further include: determining a voice type of the user according to the voice information, and jointly determining a corresponding auxiliary voice data packet based on the voice type and the geographic location information.
  • the user's voice type is determined by the pronunciation and pitch of the user's language and can include dialects and accents.
  • the voice recognition system 10 invokes the "acoustic_Guangzhou" auxiliary voice data packet to identify the voice information.
  • the voice recognition system 10 can also acquire the voice type of the user by acquiring information input on the interface provided by the display unit 30 including the text input box.
  • the electronic device 1 acquires the current geographical location information of the user, and invokes the corresponding auxiliary voice data packet according to the current geographical location information to cause a low recognition rate.
  • the S1022 may further include: acquiring current geographic location information of the user and historical geographical location information, and determining the invoked auxiliary voice data packet according to the historical geographical location information and the current geographic location information.
  • the historical geographical location information refers to geographic location information of a user's frequent residence.
  • the electronic device 1 calls the auxiliary voice data packet identifying the Fujian voice feature to identify the voice information.
  • a voice recognition method disclosed in the embodiment of the present invention obtains a plurality of auxiliary voice data packets in advance through training and learning, and the auxiliary voice data packets are voice databases divided by geographic location.
  • the auxiliary voice data packet is further subdivided into an auxiliary voice data packet based on dialect and geographic location, and an auxiliary voice data packet based on accent and geographic location.
  • the voice information of the user is recognized by the large vocabulary voice recognition method based on the preset model
  • the voice information of the user is also recognized by the auxiliary voice data packet to assist the large vocabulary voice recognition method based on the preset model, which not only improves The user's speech recognition rate also improves the user experience.
  • the voice recognition system 10 includes an acquisition module 100, a first identification module 102, a second identification module 104, a display module 106, a setting module 108, and an update module 110.
  • a module referred to in the present invention refers to a series of computer program segments that can be executed by the processing unit 40 and that are capable of performing a fixed function, which are stored in the storage unit 20. In the present embodiment, the functions of the respective modules will be described in detail in the subsequent embodiments.
  • the obtaining module 100 is configured to acquire voice information input by a user.
  • the user can directly input voice through the voice receiving unit 50 of the electronic device 1, and the acquiring module 100 acquires voice information according to the content of the voice input by the user.
  • the display unit 30 of the electronic device 1 provides a graphical user interface, and the graphical user interface includes a voice input icon.
  • the obtaining module 100 passes The voice receiving unit 50 acquires voice information input by the user.
  • the first identification module 102 is configured to identify the voice information to obtain a first recognition result.
  • the second identification module 102 is configured to identify the voice information to obtain a second recognition result.
  • the first identification module 102 may be a large vocabulary speech recognition module based on a preset model
  • the second identification module 102 may be a speech recognition module based on an auxiliary voice data packet. That is, the speech recognition module based on the auxiliary voice data packet is used to assist the large vocabulary speech recognition module based on the preset model for speech recognition.
  • the voice recognition module based on the auxiliary voice data packet may be a voice recognition module based on a geographic location established auxiliary voice data packet.
  • the voice recognition system 10 may first perform the first voice recognition module 102 to identify the voice information, and then perform the second voice recognition module 102 to identify the second voice information.
  • the speech recognition system 10 may perform the first speech recognition module 102 and the second speech recognition module 102 to identify the speech, respectively. information.
  • the voice information is identified by the large vocabulary speech recognition module based on the preset model, the voice information is recognized by the voice recognition module based on the auxiliary voice data packet, that is, the voice recognition system 10 is operated by the first thread.
  • the first identification module 102 identifies the voice information, and in parallel, a second thread runs the second identification module 102 to identify the voice information.
  • the large vocabulary speech recognition module based on the preset model refers to a speech recognition library established according to standard Mandarin, and any user can call the speech recognition library to identify according to standard Mandarin.
  • Large vocabulary speech recognition based on a preset model does not take into account the influence of dialects and geographic locations and/or accents and geographic locations.
  • the large vocabulary speech recognition module based on the preset model is the same as in the prior art.
  • auxiliary voice recognition module The voice recognition module based on the auxiliary voice data packet (hereinafter referred to as "auxiliary voice recognition module" for convenience of description) considers the influence of dialect and geographical position and/or accent and geographical position, and needs to establish a geographical basis through training and learning in advance.
  • Location voice packets Refer to Figure 5 and the corresponding description for the location-based speech recognition module.
  • the display module 106 is configured to display the first voice recognition result and the second voice recognition result according to a preset rule.
  • the preset rules are preset by the setting module 108.
  • the setting module 108 may pre-allocate a first weight for the first voice recognition result, pre-allocate a second weight for the second voice recognition result, and determine a display of the voice recognition result corresponding to the weight value according to the magnitude of the weight value. the way.
  • the sum of the first weight value and the second weight value may be a fixed number, for example, an integer of one.
  • the first weight value preset by the setting module 108 is greater than the second weight value, that is, the weight value assigned by the setting module 108 to the first voice recognition method is greater than the weight value assigned to the second voice recognition method. .
  • the setting module 108 may further set a rule that the first recognition score is preset for the first voice recognition result, and the second recognition score is preset for the second voice recognition result, according to The size of the recognition score determines how the speech recognition result corresponding to the score is displayed.
  • the first identification score value preset by the setting module 108 is greater than the second identification score value.
  • the manner in which the speech recognition result is displayed includes, but is not limited to, the displayed time and/or the displayed position. However, it is not limited to the displayed time and the displayed position.
  • the rule set by the setting module 108 is to assign a weight to the speech recognition result, then
  • the display module 106 may display the first voice recognition result with the corresponding weight value on the display unit 30 of the electronic device 1 at the first a location, such as the upper half of the user interface provided by the display unit 30; when the preset first weight value is less than a preset second weight value, the display module 106 will identify the first voice with a small weight value The result is displayed in the second position, such as the lower half of the user interface provided by the display unit 30.
  • the display module 106 displays the first voice recognition result on the display unit 30 of the electronic device 1 after a preset time (eg The second speech recognition result is displayed on the display unit 30 of the electronic device 1 after 2 seconds.
  • the voice recognition system 10 further includes the update module 110, configured to update the preset rule in combination with the acquired user feedback information.
  • the user feedback information may be obtained according to an operation of the user. For example, if the user selects the first voice recognition result, the user feedback information acquired by the acquiring module 100 indicates that the best voice recognition result is obtained by using the first voice recognition method. If the user selects the second voice recognition result, the user feedback information acquired by the acquiring module 100 indicates that the best voice recognition result is obtained by using the second voice recognition method.
  • the updating the module 110 to update the preset rule may be to adjust a preset weight value or adjust a preset recognition score value.
  • the update module 110 increases the weight value or the recognition score value corresponding to the voice recognition result according to the voice recognition result selected by the user, and/or the weight value or the identification score corresponding to the voice recognition result that the user does not select. The value is reduced. For example, when the obtained user feedback information is that the first voice recognition result is selected, the update module 110 increases the first weight value or the first recognition score value corresponding to the first voice recognition result, and/or corresponds to The second weight value or the second recognition score value of the second voice recognition result is decreased.
  • the update module 110 increases the second weight value or the second recognition score value corresponding to the second voice recognition result, and/or corresponds to the first The first weight value or the first recognition score value of the speech recognition result is decreased.
  • the above-mentioned weight value or the value of the point value may be increased or decreased according to a preset ratio or value.
  • the second identification module 104 includes a calling submodule 1040, a download submodule 1042, and a determining submodule 1044.
  • a module referred to in the present invention refers to a series of computer program segments that can be executed by the processing unit 40 and that are capable of performing a fixed function, which are stored in the storage unit 20. In the present embodiment, the functions of the respective modules will be described in detail in the subsequent embodiments.
  • the acquiring module 100 is further configured to: when receiving the voice information of the user, obtain the current geographic location information of the user.
  • the acquiring module 100 acquires the geographical location information of the electronic device 1 by using the positioning module and/or the network connection module built in the electronic device 1 .
  • the positioning module includes, but is not limited to, a Global Positioning System (GPS).
  • the network connection module includes, but is not limited to, The 3rd Generation Telecommunication (3G), General Packet Radio Service (GPRS), and Wireless Fidelity (Wi -Fi).
  • the geographical location information of the electronic device 1 is considered to be the geographical location information of the current location of the user.
  • the obtaining module 100 may further determine the current geographic location information of the user by receiving an instruction set by the user and according to an instruction set by the user.
  • the electronic device 1 is provided with a location selection list, which includes the names of all cities in China.
  • the user selects the geographical location information corresponding to the user input voice information by triggering the location selection list.
  • the electronic device 1 is provided with a text input box, and the user inputs the current geographic location information in the corresponding interface by activating the text input box function.
  • the calling sub-module 1040 is configured to invoke a corresponding auxiliary voice data packet according to the geographic location information.
  • the calling sub-module 1040 calls the corresponding auxiliary voice data packet from the storage unit 20 according to the geographical location information.
  • the storage unit 20 pre-stores the auxiliary voice data packet and the voice information having the geographical location voice feature included in the auxiliary voice data packet.
  • the calling sub-module 1040 invokes an auxiliary voice data packet identifying the Guangdong voice feature.
  • the acquiring module 100 when acquiring the current geographic location information of the user, the download sub-module 102 is executed.
  • the download sub-module 1042 downloads the auxiliary voice data packet from a server communicatively coupled to the electronic device 1.
  • the communication connection can be a wireless communication connection.
  • the auxiliary voice data packet is obtained and trained by the user in advance and deployed to the server, and the download sub-module 1042 may request the server to send an auxiliary voice data packet corresponding to the geographical location information through a network.
  • the second identification module 104 is configured to identify the voice information according to the auxiliary voice data packet to obtain a second voice recognition result.
  • the second identification module 104 uses the second voice recognition method to identify the voice information to obtain the second voice recognition result.
  • the second identification module 104 may further include a determining submodule 1044 for using the speech according to the speech.
  • the information determines the type of voice of the user.
  • the calling sub-module 1040 jointly determines a corresponding auxiliary voice data packet based on the voice type and the geographic location information.
  • the user's voice type is determined by the pronunciation and pitch of the user's language and can include dialects and accents.
  • the calling sub-module 1040 calls the auxiliary voice data packet of “Acoustic_Guangzhou” to identify the voice information.
  • the obtaining module 100 may further acquire the voice type of the user by acquiring information input on the interface provided by the display unit 30 including the text input box.
  • the obtaining module 100 acquires the current geographic location information of the user, and the calling sub-module 1040 calls the corresponding auxiliary voice data packet according to the current geographic location information.
  • the acquiring module 100 is further configured to acquire current geographic location information and historical geographic location information of the user, and the calling sub-module 1040 determines the invoked auxiliary voice data packet according to the historical geographical location information and the current geographic location information. .
  • the historical geographical location information refers to geographic location information of a user's frequent residence.
  • the calling sub-module 1040 calls the auxiliary voice data packet identifying the Fujian voice feature to identify the voice information.
  • a voice recognition system disclosed in the embodiment of the present invention obtains a plurality of auxiliary voice data packets in advance through training and learning, and the auxiliary voice data packets are voice databases divided by geographic location.
  • the auxiliary voice data packet is further subdivided into an auxiliary voice data packet based on dialect and geographic location, and an auxiliary voice data packet based on accent and geographic location.
  • the voice information of the user is recognized by the large vocabulary voice recognition module based on the preset model
  • the voice information of the user is also recognized by the auxiliary voice data packet to assist the large vocabulary voice recognition method based on the preset model,
  • the user's speech recognition rate is improved, and the user experience is also improved.
  • modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of the embodiment.
  • each functional module in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the above integrated unit can be implemented in the form of hardware or in the form of hardware plus software function modules.
  • the above-described integrated unit implemented in the form of a software function module can be stored in a computer readable storage medium.
  • the software function modules described above are stored in a storage medium and include instructions for causing a computer unit (which may be a personal computer, a server, or a network unit, etc.) or a processor to perform the methods of the various embodiments of the present invention. Part of the steps.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Telephonic Communication Services (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种语音识别方法,应用于电子设备中,该方法包括:获取用户输入的语音信息(S100);利用第一语音识别方法识别所述语音信息得到第一语音识别结果,利用第二语音识别方法识别所述语音信息得到第二语音识别结果(S102),根据预先设置的规则显示所述第一语音识别结果及所述第二语音识别结果(S104)。一种语音识别系统(10),其具有获取模块(100)、第一识别模块(102)、第二识别模块(104)、显示模块(106)、设置模块(108)以及更新模块(110)。上述语音识别方法和系统可提高语音识别的准确率。

Description

语音识别方法及系统
本申请要求于2016年6月21日提交中国专利局,申请号为201610465607.6、发明名称为“语音识别方法及系统”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本发明涉及语音识别技术领域,尤其涉及一种基于地理位置辅助的语音识别方法及系统。
背景技术
近年来,语音识别技术取得显著进步,已经从实验室走向市场。在实际应用中,例如自动电话应答系统,通过识别用户的语音输入信息,自动完成和用户的交互。
目前普通话作为交流的语言已经基本普及,但不同地区的人还有方言的差别,因此受到方言的影响,各地人群的普通话都有不同的特征,但同一区域的人群,其说普通话的语速、语音、语义具有类似性。
由于不同地区的用户普通话并不完全标准,带有地区特点,现有的语音输入通过录取用户语音数据并根据语音识别算法识别输出结果,并未结合用户所处的地理位置等数据做辅助参考,对于某些带有地方口音、带有部分地方方言时的语音识别的识别率并不高。
发明内容
鉴于以上内容,有必要提供一种语音识别方法及系统,能根据用户当前的地理位置调用辅助语音数据包来识别用户语音,从而提高语音识别的准确率。
一种语音识别方法,应用于电子设备中,该方法包括:
获取用户输入的语音信息;
利用第一语音识别方法识别所述语音信息得到第一语音识别结果,利用第二语音识别方法识别所述语音信息得到第二语音识别结果;及
根据预先设置的规则显示所述第一语音识别结果及所述第二语音识别结 果。
根据本发明的一个优选实施例,所述第一语音识别方法是基于预设模型的大词汇量语音识别方法,所述第二语音识别方法是基于辅助语音数据包的语音识别方法。
根据本发明的一个优选实施例,所述基于辅助语音数据包的语音识别方法包括:
接收到所述语音信息时,获取该用户当前的地理位置信息;
根据所述地理位置信息调用对应的辅助语音数据包;及
根据所述辅助语音数据包识别所述语音信息得到所述第二语音识别结果。
根据本发明的一个优选实施例,所述方法还包括:
预先设置多个基于地理位置的语音数据包,并将所述语音数据包存储于所述电子设备中或者存储于与所述电子设备连接的服务器中。
根据本发明的一个优选实施例,在根据所述地理位置信息调用对应的辅助语音数据包之前,所述方法还包括:
根据所述语音信息确定该用户的语音类型,所述语音类型包括口音及方言;及
基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
根据本发明的一个优选实施例,所述方法还包括:
接收到所述语音信息时,获取该用户当前的地理位置信息及历史地理位置信息;及
根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
根据本发明的一个优选实施例,所述方法还包括:
结合获取的用户反馈信息更新所述预先设置的规则,所述预先设置的规则包括:
为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式;或
为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识 别结果的显示方式,
其中,所述显示方式包括显示的时间或显示的位置。
根据本发明的一个优选实施例,所述更新所述预先设置的规则包括:
根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分数值减小。
一种语音识别系统,运行于电子设备中,该系统包括:
获取模块,用于获取用户输入的语音信息;
第一识别模块,用于识别所述语音信息得到第一语音识别结果;
第二识别模块,用于识别所述语音信息得到第二语音识别结果;及
显示模块,用于根据预先设置的规则显示所述第一语音识别结果及所述第二语音识别结果。
根据本发明的一个优选实施例,所述第一语音识别模块是基于预设模型的大词汇量语音识别模块,所述第二语音识别模块是基于辅助语音数据包的语音识别模块。
根据本发明的一个优选实施例,
所述获取模块,还用户接收到所述语音信息时,获取该用户当前的地理位置信息;
所述第二识别模块包括:
调用子模块,用于根据所述地理位置信息调用对应的辅助语音数据包;及
该第二识别模块,用于根据所述辅助语音数据包识别所述语音信息得到所述第二语音识别结果。
根据本发明的一个优选实施例,所述系统还包括:
设置模块,用于预先设置多个基于地理位置的语音数据包,并将所述语音数据包存储于所述电子设备中或者存储于与所述电子设备连接的服务器中。
根据本发明的一个优选实施例,所述系统还包括确定子模块:
用于根据所述语音信息确定该用户的语音类型,所述语音类型包括口音及方言;及
基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
根据本发明的一个优选实施例,其特征在于,
所述获取模块,还用于接收到所述语音信息时,获取该用户当前的地理位置信息及历史地理位置信息;及
所述调用子模块,还用于根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
根据本发明的一个优选实施例,所述系统还包括:
更新模块,用于结合获取的用户反馈信息更新所述预先设置的规则,所述预先设置的规则是由所述设置模块设置的,包括:
为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式;或
为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识别结果的显示方式,
其中,所述显示方式包括显示的时间或显示的位置。
根据本发明的一个优选实施例,所述更新模块更新所述预先设置的规则包括:
根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分数值减小。
由以上技术方案可以看出,本发明的语音识别方法及系统能够根据不同区域普通话的特征,建立多个辅助语音数据包,对处于不同地理位置的用户调用不同的辅助语音数据包,可有效的减少语音识别库的种类,并提高语音识别率。
附图说明
图1是本发明用于执行一个语音识别系统的电子设备较佳实施方式的硬件架构示意图。
图2是本发明语音识别方法较佳实施例的流程图。
图3是本发明基于辅助语音数据包的语音识别方法的较佳实施例的流程 图。
图4是本发明语音识别系统第一实施例的功能模块图。
图5是本发明语音识别系统第二实施例的功能模块图。
【主要元件符号说明】
电子设备 1
语音识别系统 10
存储单元 20
显示单元 30
处理单元 40
语音接收单元 50
获取模块 100
第一识别模块 102
第二识别模块 104
调用子模块 1040
下载子模块 1042
确定子模块 1044
显示模块 106
设置模块 108
更新模块 110
具体实施方式
为了使本发明的目的、技术方案和优点更加清楚,下面结合附图和具体实施例对本发明进行详细描述。显然,所描述的实施例仅仅是本发明的一部分实施例,而不是全部的实施例。此外,应当理解,本文所描述的具体实施例,仅用以解释本发明,并不用于限定本发明。
如图1所示,是本发明用于执行一个语音识别系统的电子设备较佳实施例的硬件架构示意图。如该硬件架构示意图所示,电子设备1包括语音识别系统10。该电子设备1还包括存储单元20、显示单元30、处理单元40及语音接收 单元50。
优选地,本发明的语音识别方法通过所述电子设备1中的语音识别系统10来实现。
所述电子设备1包括一种能够按照事先设定或存储的指令,自动进行数值计算和/或信息处理的电子设备,其硬件包括但不限于微处理器、专用集成电路(Application Specific Integrated Circuit,ASIC)、可编程门阵列(Field Programmable Gate Array,FPGA)、数字处理器(Digital Signal Processor,DSP)、嵌入式设备等。所述电子设备1还可包括用户设备。所述用户设备包括但不限于任何一种可与用户通过键盘、鼠标、遥控器、触摸板或声控设备等方式进行人机交互的电子产品,例如,个人计算机、平板电脑、智能手机、个人数字助理(Personal Digital Assistant,PDA)、游戏机、交互式网络电视(Internet Protocol Television,IPTV)、智能式穿戴设备等。其中,所述用户设备所处的网络包括但不限于互联网、广域网、城域网、局域网、虚拟专用网络(Virtual Private Network,VPN)等。
需要说明的是,所述用户设备仅为举例,其他现有的或今后可能出现的用户设备如可适应于本发明,也应包含在本发明的保护范围以内,并以引用方式包含于此。
在一个实施例中,所述语音识别系统10用于当用户输入语音信息时,获取该用户输入的语音信息,利用基于预设模型的大词汇量语音识别方法(例如,基于隐马尔可夫模型的大词汇量语音识别方法)对所输入的语音信息进行识别得到第一语音识别结果,利用基于辅助语音数据包的语音识别方法(例如,根据该用户当前的地理位置信息调用与该地理位置信息相对应的辅助语音数据包)进行识别得到第二语音识别结果。所述语音识别系统10通过比较第一语音识别结果和第二语音识别结果得到一个最优识别结果,不仅提高了语音识别率,还提高了用户的体验。
在一个实施例中,所述存储单元20用于存储安裝于所述电子设备1中的软件程序及数据,例如所述语音识别系统10。该存储单元20可以是所述电子设备1的内部存储单元,例如所述电子设备1的硬盘或者内存。该存储单元20也可以是所述电子设备1的外部存储设备,例如所述电子设备1上的插接式硬盘、智能媒体卡(Smart Media Card,SMC)、安全数字卡(Secure Digital  Card,SD)、快闪存储器卡(flash card)等储存单元。进一步地,所述存储单元20还可以既包括所述电子设备1的内部存储单元,也可以包括外部存储设备。
在本实施例中,所述存储单元20中预先存储有多个辅助语音数据包及与该多个辅助语音数据包相对应的语音信息。所述辅助语音数据包可以是基于地理位置的语音数据包,对应地,所述存储单元20中存储的是具有该地理位置语音特征的语音信息。
在本实施例中,所述的地理位置是以地市为单位进行划分的。在其他实施例中,对于方言复杂的地理位置,还可细分到地市以下的区域,例如,以县级市为单位进行划分或者以设定的区域为单位进行划分。
由于在同一地理位置,所讲的普通话也会存在口音和方言的区别。或者即使不在同一地理位置,方言或者口音也有可能相同,因此,所述存储单元20中存储的基于地理位置的语音数据包在其他的一些实施例中进一步包括基于方言和地理位置的语音数据包及基于口音和地理位置的语音数据包。
例如,基于方言和地理位置的语音数据包可以包括:粤语_香港、粤语_广州、闽南语_泉州、闽南语_厦门。基于口音和地理位置的语音数据包可以包括:口音_福建、口音_广州。需要说明的是,基于口音和地理位置的语音数据包包括,但不限于,声母、韵母的吐字方式。
在一个实施例中,所述显示单元30用来显示图形用户界面(Graphic User Interface,GUI),该图形用户界面中可包括多个应用程序图标及/或多个虚拟按键,该应用程序图标及虚拟按键可以是代表所述电子设备1所能提供的各个功能,例如语音输入图标代表了所述电子设备1可提供语音输入的功能,地理位置选择列表按键代表了所述电子设备1可提供选择地理位置的功能,以及文本输入框按键代表了所述电子设备1可提供输入地理位置的功能。
所述显示单元30可以是,但不限于,触摸显示屏等具有触摸功能的显示单元。故用户除了可通过所述显示单元30观看所述电子设备1所显示的应用程序图标及/或虚拟按键外,也可通过所述显示单元30输入功能指令,例如,运行所述应用程序图标对应的应用程序的指令,或者激活虚拟按键启动相应的功能的指令。
在一个实施例中,所述处理单元40是一个或者多个中央处理器(Central  Processing unit,CPU)、微处理器或其他数字处理芯片等。该处理单元40用于执行软件程序代码或运算数据,例如执行所述的语音识别系统10。本实施例中,所述处理单元40接收用户输入的语音信息,同时获取该用户当前的地理位置信息,在进行语音识别时,结合基于预设模型的大词汇量语音识别(例如,基于隐马尔可夫模型的大词汇量语音识别方法,或者基于人工神经网络模型的语音识别方法)和基于辅助语音数据包的语音识别(例如,基于地理位置的辅助语音数据包的语音识别)分别输出第一识别结果和第二识别结果,根据用户比较第一识别结果和第二识别结果做出的选择,动态调整基于预设模型的大词汇量语音识别和基于辅助语音数据包的语音识别的权重,以提高语音识别的准确率。
所述处理单元40与所述语音识别系统10、存储单元20、显示单元30及语音输入单元50通讯连接。所述通讯可以通过串行外围设备接口总线(Universal Serial Bus,USB)或其他通信路径或协议来实现。
所述语音输入单元50用于录入用户的语音信息。所述显示单元30包括,但不限于,麦克风。
如图2所示,是本发明语音识别方法的较佳实施例的流程图。根据不同的需求,该流程图中步骤的顺序可以改变,某些步骤可以省略。
S100,获取用户输入的语音信息。
在本实施例中,用户可以直接通过所述电子设备1的语音接收单元50输入语音,所述语音识别系统10根据用户输入语音的内容获取语音信息。
在其他实施例中,所述电子设备1的显示单元30提供了一个图形用户界面,所述图形用户界面上包括一个语音输入图标,在用户点击所述语音输入图标时,所述语音识别系统10通过所述语音接收单元50获取用户输入的语音信息。
S102,利用第一语音识别方法识别所述语音信息得到第一识别结果,以及利用第二语音识别方法识别所述语音信息得到第二识别结果。
在本实施例中,所述第一语音识别方法识别可以是基于预设模型的大词汇量语音识别方法,所述第二语音识别方法可以是基于辅助语音数据包的语音识别方法。即利用基于辅助语音数据包的语音识别方法协助基于预设模型的大词汇量语音识别方法进行语音识别。所述基于辅助语音数据包的语音识别方法可 以是基于地理位置建立的辅助语音数据包的语音识别方法。在一些实施例中,所述语音识别系统10可以先执行所述第一语音识别方法识别所述语音信息,再执行所述第二语音识别方法识别所述第二语音信息。
在一些实施例中,为了提高识别效率,所述语音识别系统10可以并行执行所述第一语音识别方法与所述第二语音识别方法分别识别所述语音信息。利用所述基于预设模型的大词汇量语音识别方法识别所述语音信息时,同时利用所述基于辅助语音数据包的语音识别方法识别所述语音信息,即所述语音识别系统10以第一线程运行所述基于预设模型的大词汇量语音识别方法以识别所述语音信息,并行地一第二线程运行所述基于辅助语音数据包的语音识别以识别所述语音信息。
在本实施例中,所述基于预设模型的大词汇量语音识别方法是指按照标准普通话建立的语音识别库,任何用户均可以调用所述语音识别库,按照标准普通话进行识别。基于预设模型的大词汇量语音识别不考虑方言和地理位置及/或口音和地理位置的影响。所述基于预设模型的大词汇量语音识别方法可采用现有技术中的语音识别方法,通过预先建立的多个模型进行学习、训练以识别用户的语音,并将语音信息转换成文字信息。
所述基于辅助语音数据包的语音识别方法(为便于描述,下文简称为“辅助语音识别方法”)考虑方言和地理位置及/或口音和地理位置的影响,需要事先通过训练和学习建立基于地理位置的语音数据包。关于所述基于地理位置的语音识别方法请参阅图3及相应描述。
S104,根据预先设置的规则显示所述第一语音识别结果和第二语音识别结果。
本实施例中,所述预先设置的规则可以是,所述语音识别系统10为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式。所述第一权重值和所述第二权重值的总和可以为一固定数,例如,为整数1。优选地,所述语音识别系统10预先设置的第一权重值大于第二权重值,也就是说所述语音识别系统10为第一语音识别方法分配的权重值大于为第二语音识别方法分配的权重值。
在其他实施例中,所述预先设置的规则还可以是,所述语音识别系统10 为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识别结果的显示方式。优选地,所述语音识别系统10预先设置的第一识别分数值大于第二识别分数值。
所述语音识别结果的显示方式包括,但不限于:显示的时间及/或显示的位置。
例如,所述语音识别系统10预先设置的规则是为语音识别结果分配权重,则当预先设置的第一权重值大于预先设置的第二权重值时,可以在所述电子设备1的显示单元30上将对应权重值大的第一语音识别结果显示在第一位置,如所述显示单元30提供的用户界面的上半部分;当预先设置的第一权重值小于预先设置的第二权重值时,将对应权重值小的第一语音识别结果显示在第二位置,如所述显示单元30提供的用户界面的下半部分。
此外,当预先设置的第一权重值大于预先设置的第二权重值时,在所述电子设备1的显示单元30上显示第一语音识别结果,在预设时间之后(例如,2秒后)在所述电子设备1的显示单元30上显示第二语音识别结果。
在本实施例中,所述的语音识别方法进一步包括:结合获取的用户反馈信息更新所述预先设置的规则。
所述用户反馈信息可以根据用户的操作得到。例如,用户选取了第一语音识别结果,则所述语音识别系统10获取到的用户反馈信息表示最佳语音识别结果是利用第一语音识别方法得到的。若用户选取了第二语音识别结果,则所述语音识别系统10获取到的用户反馈信息表示最佳语音识别结果是利用第二语音识别方法得到的。
所述更新所述预先设置的规则可以是调整预先设置的权重值或者调整预先设置的识别分数值。
具体地,所述语音识别系统10根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分数值减小。例如,当获取的用户反馈信息是选取了第一语音识别结果,则将对应该第一语音识别结果的第一权重值或者第一识别分数值变大,及/或将对应第二语音识别结果的第二权重值或者第二识别分数值减小。当获取的用户反馈信息是选取了第二语音识别结果,则将对应该 第二语音识别结果的第二权重值或者第二识别分数值变大,及/或将对应第一语音识别结果的第一权重值或者第一识别分数值减小。
其中,上述的权重值或者分数值的变大或减小可根据预先设置的比例或者数值进行。
请一并参阅图3所示,为基于辅助语音数据包的语音识别方法的较佳实施例的流程图。根据不同的需求,该流程图中步骤的顺序可以改变,某些步骤可以省略。
S1020,接收到用户的语音信息时,获取该用户当前的地理位置信息。
在本实施例中,所述语音识别系统10通过所述电子设备1内置的定位模块及/或网络连接模块获取所述电子设备1当前所在的地理位置信息。所述定位模块包括,但不限于:全球定位系统(Global Positioning System,GPS)。所述所述网络连接模块包括,但不限于:第3代移动通信技术(The 3rd Generation Telecommunication,3G)、通用分组无线业务(General Packet Radio Service,GPRS)以及无线保真技术(wireless fidelity,Wi-Fi)。所述电子设备1当前所在的地理位置信息即被认为是该用户当前所在的地理位置信息。
在一些实施例中,所述语音识别系统10还可以通过接收用户设置的指令,并根据该用户设置的指令确定该用户当前的地理位置信息。
例如,所述电子设备1中设置有位置选择列表,该位置选择列表包括中国所有城市的名称。用户通过触发该位置选择列表,选择与用户输入语音信息相应的地理位置信息。
又如,所述电子设备1中设置有文本输入框,用户通过激活该文本输入框功能,在相应的界面中输入当前地理位置信息。
S1022,根据所述地理位置信息调用对应的辅助语音数据包。
在本实施例中,所述电子设备1根据所述地理位置信息从所述存储单元20中调用对应的辅助语音数据包。
所述存储单元20中预先存储有辅助语音数据包及该辅助语音数据包包括的具有地理位置语音特征的语音信息。
例如,所述地理位置信息是广东,则所述语音识别系统10调用识别广东语音特征的辅助语音数据包。
在一些实施例中,如果所述电子设备1的存储单元20中没有预先存储有对应所述地理位置信息的辅助语音数据包时,则所述语音识别系统10在获取用户当前的地理位置信息时,从与所述电子设备1通讯连接的服务器下载该辅助语音数据包。所述通讯连接可以是无线通讯连接。所述辅助语音数据包由用户事先进行训练和学习得到并布署于所述服务器,所述语音识别系统10可以通过网络请求所述服务器发送对应所述地理位置信息的辅助语音数据包。
S1024,根据所述辅助语音数据包识别所述语音信息得到第二语音识别结果。
在本实施例中,所述语音识别系统10利用所述第二语音识别方法识别所述语音信息得到所述第二语音识别结果。
进一步地,为了解决即使在同一地理位置也会存在方言或者口音的差别而造成的语音识别率不高的问题,所述语音识别系统10根据所述地理位置信息调用对应的辅助语音数据包之前,所述S1022还可以包括:根据所述语音信息确定该用户的语音类型,并基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
该用户的语音类型由用户语言的发音和音调决定,可以包括方言和口音。
例如,用户的当前的地理位置为广州,用户的语音类型可以是口音(例如,粤语),则所述语音识别系统10调用“口音_广州”的辅助语音数据包识别所述语音信息。在一些实施例中,所述语音识别系统10还可以通过获取所述显示单元30提供的包括有文本输入框的界面上输入的信息获取用户的语音类型。
更进一步地,为了避免用户临时去某地出差或者旅游时,所述电子设备1获取该用户当前的地理位置信息,并根据该当前的地理位置信息调用相应的辅助语音数据包造成识别率低时,所述S1022还可以包括:获取用户当前的地理位置信息以及历史地理位置信息,并根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
在本实施例中,所述历史地理位置信息是指用户的经常居住地的地理位置信息。
例如,用户当前的地理位置为广州,而用户的经常居住地在福建,则电子设备1调用识别福建语音特征的辅助语音数据包来识别所述语音信息。
综上所述,本发明实施例公开的一种语音识别方法,预先通过训练和学习得到多个辅助语音数据包,该辅助语音数据包是以地理位置为单位进行划分的语音数据库。同时基于用户的语音类型,辅助语音数据包进一步细分为基于方言和地理位置的辅助语音数据包,以及基于口音和地理位置的辅助语音数据包。利用基于预设模型的大词汇量语音识别方法识别用户的语音信息时,同时也利用该辅助语音数据包识别用户的语音信息从而协助所述基于预设模型的大词汇量语音识别方法,不仅提高了用户的语音识别率,也提高了用户体验。
如图4所示,是本发明语音识别系统的第一实施例的功能模块图。所述语音识别系统10包括获取模块100、第一识别模块102、第二识别模块104、显示模块106、设置模块108及更新模块110。本发明所称的模块是指一种能够被处理单元40所执行并且能够完成固定功能的一系列计算机程序段,其存储在存储单元20中。在本实施例中,关于各模块的功能将在后续的实施例中详述。
所述获取模块100,用于获取用户输入的语音信息。
在本实施例中,用户可以直接通过所述电子设备1的语音接收单元50输入语音,所述获取模块100根据用户输入语音的内容获取语音信息。
在其他实施例中,所述电子设备1的显示单元30提供了一个图形用户界面,所示图形用户界面上包括一个语音输入图标,在用户点击所述语音输入图标时,所述获取模块100通过所述语音接收单元50获取用户输入的语音信息。
所述第一识别模块102,用于识别所述语音信息得到第一识别结果。
所述第二识别模块102,用于识别所述语音信息得到第二识别结果。
在本实施例中,所述第一识别模块102可以是基于预设模型的大词汇量语音识别模块,所述第二识别模块102可以是基于辅助语音数据包的语音识别模块。即利用基于辅助语音数据包的语音识别模块协助基于预设模型的大词汇量语音识别模块进行语音识别。所述基于辅助语音数据包的语音识别模块可以是基于地理位置建立的辅助语音数据包的语音识别模块。在一些实施例中,所述语音识别系统10可以先执行所述第一语音识别模块102识别所述语音信息,再执行所述第二语音识别模块102识别所述第二语音信息。
在一些实施例中,为了提高识别效率,所述语音识别系统10可以并行执行所述第一语音识别模块102与所述第二语音识别模块102分别识别所述语音 信息。利用基于预设模型的大词汇量语音识别模块识别所述语音信息时,同时利用所述基于辅助语音数据包的语音识别模块识别所述语音信息,即所述语音识别系统10以第一线程运行所述第一识别模块102以识别所述语音信息,并行地一第二线程运行所述第二识别模块102以识别所述语音信息。
在本实施例中,基于预设模型的大词汇量语音识别模块是指按照标准普通话建立的语音识别库,任何用户均可以调用所述语音识别库,按照标准普通话进行识别。基于预设模型的大词汇量语音识别不考虑方言和地理位置及/或口音和地理位置的影响。所述基于预设模型的大词汇量语音识别模块与现有技术中的相同。
所述基于辅助语音数据包的语音识别模块(为便于描述,下文简称为“辅助语音识别模块”)考虑方言和地理位置及/或口音和地理位置的影响,需要事先通过训练和学习建立基于地理位置的语音数据包。关于所述基于地理位置的语音识别模块请参阅图5及相应描述。
所述显示模块106,用于根据预先设置的规则显示所述第一语音识别结果和第二语音识别结果。
本实施例中,所述预先设置的规则由所述设置模块108预先设置。所述设置模块108可以为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式。所述第一权重值和所述第二权重值的总和可以为一固定数,例如,为整数1。优选地,所述设置模块108预先设置的第一权重值大于第二权重值,也就是说所述设置模块108为第一语音识别方法分配的权重值大于为第二语音识别方法分配的权重值。
在其他实施例中,所述设置模块108预先设置的规则还可以是,为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识别结果的显示方式。优选地,所述设置模块108预先设置的第一识别分数值大于第二识别分数值。
所述语音识别结果的显示方式包括,但不限于:显示的时间及/或显示的位置。但不限于显示的时间和显示的位置。
例如,所述设置模块108预先设置的规则是为语音识别结果分配权重,则 当预先设置的第一权重值大于预先设置的第二权重值时,所述显示模块106可以在所述电子设备1的显示单元30上将对应权重值大的第一语音识别结果显示在第一位置,如所述显示单元30提供的用户界面的上半部分;当预先设置的第一权重值小于预先设置的第二权重值时,所述显示模块106将对应权重值小的第一语音识别结果显示在第二位置,如所述显示单元30提供的用户界面的下半部分。
此外,当预先设置的第一权重值大于预先设置的第二权重值时,所述显示模块106在所述电子设备1的显示单元30上显示第一语音识别结果,在预设时间之后(例如,2秒后)在所述电子设备1的显示单元30上显示第二语音识别结果。
在本实施例中,所述的语音识别系统10进一步包括所述更新模块110,用于结合获取的用户反馈信息更新所述预先设置的规则。
本实施例中,所述用户反馈信息可以根据用户的操作得到。例如,用户选取了第一语音识别结果,则所述获取模块100获取到的用户反馈信息表示最佳语音识别结果是利用第一语音识别方法得到的。若用户选取了第二语音识别结果,则所述获取模块100获取到的用户反馈信息表示最佳语音识别结果是利用第二语音识别方法得到的。
所述更新模块110更新所述预先设置的规则可以是调整预先设置的权重值或者调整预先设置的识别分数值。
具体地,所述更新模块110根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分数值减小。例如,当获取的用户反馈信息是选取了第一语音识别结果,则所述更新模块110将对应该第一语音识别结果的第一权重值或者第一识别分数值变大,及/或将对应第二语音识别结果的第二权重值或者第二识别分数值减小。当获取的用户反馈信息是选取了第二语音识别结果,则所述更新模块110将对应该第二语音识别结果的第二权重值或者第二识别分数值变大,及/或将对应第一语音识别结果的第一权重值或者第一识别分数值减小。
其中,上述的权重值或者分数值的变大或减小可根据预先设置的比例或者数值进行。
请一并参阅图5所示,是本发明语音识别系统的第二实施例的功能模块图。其中,所述第二识别模块104包括调用子模块1040、下载子模块1042及确定子模块1044。本发明所称的模块是指一种能够被处理单元40所执行并且能够完成固定功能的一系列计算机程序段,其存储在存储单元20中。在本实施例中,关于各模块的功能将在后续的实施例中详述。
所述获取模块100,还用于接收到用户的语音信息时,获取该用户当前的地理位置信息。
在本实施例中,所述获取模块100通过所述电子设备1内置的定位模块及/或网络连接模块获取所述电子设备1当前所在的地理位置信息。所述定位模块包括,但不限于:全球定位系统(Global Positioning System,GPS)。所述所述网络连接模块包括,但不限于:第3代移动通信技术(The 3rd Generation Telecommunication,3G)、通用分组无线业务(General Packet Radio Service,GPRS)以及无线保真技术(wireless fidelity,Wi-Fi)。所述电子设备1当前所在的地理位置信息即被认为是该用户当前所在的地理位置信息。
在一些实施例中,所述获取模块100还可以通过接收用户设置的指令,并根据该用户设置的指令确定该用户当前的地理位置信息。
例如,所述电子设备1中设置有位置选择列表,该位置选择列表包括中国所有城市的名称。用户通过触发该位置选择列表,选择与用户输入语音信息相应的地理位置信息。
又如,所述电子设备1中设置有文本输入框,用户通过激活该文本输入框功能,在相应的界面中输入当前地理位置信息。
所述调用子模块1040,用于根据所述地理位置信息调用对应的辅助语音数据包。
在本实施例中,所述调用子模块1040根据所述地理位置信息从所述存储单元20中调用对应的辅助语音数据包。
所述存储单元20中预先存储有辅助语音数据包及该辅助语音数据包包括的具有地理位置语音特征的语音信息。
例如,所述地理位置信息是广东,则所述调用子模块1040调用识别广东语音特征的辅助语音数据包。
在一些实施例中,如果所述电子设备1的存储单元20中没有预先存储有对应所述地理位置信息的辅助语音数据包时,则所述获取模块100在获取用户当前的地理位置信息时,执行所述下载子模块102。所述下载子模块1042从与所述电子设备1通讯连接的服务器下载该辅助语音数据包。所述通讯连接可以是无线通讯连接。所述辅助语音数据包由用户事先进行训练和学习得到并布署于所述服务器,下载子模块1042可以通过网络请求所述服务器发送对应所述地理位置信息的辅助语音数据包。
所述第二识别模块104,用于根据所述辅助语音数据包识别所述语音信息得到第二语音识别结果。
在本实施例中,第二识别模块104利用所述第二语音识别方法识别所述语音信息得到所述第二语音识别结果。
进一步地,为了解决即使在同一地理位置也会存在方言或者口音的差别而造成的语音识别率不高的问题,所述第二识别模块104还可以包括确定子模块1044:用于根据所述语音信息确定该用户的语音类型。所述调用子模块1040基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
该用户的语音类型由用户语言的发音和音调决定,可以包括方言和口音。
例如,用户的当前的地理位置为广州,用户的语音类型是口音(例如,粤语),则所述调用子模块1040调用“口音_广州”的辅助语音数据包识别所述语音信息。
在一些实施例中,所述获取模块100还可以通过获取所述显示单元30提供的包括有文本输入框的界面上输入的信息获取用户的语音类型。
更进一步地,为了避免用户临时去某地出差或者旅游时,所述获取模块100获取该用户当前的地理位置信息,所述调用子模块1040根据该当前的地理位置信息调用相应的辅助语音数据包造成识别率低时,所述获取模块100还用于获取用户当前的地理位置信息以及历史地理位置信息,所述调用子模块1040根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
在本实施例中,所述历史地理位置信息是指用户的经常居住地的地理位置信息。
例如,用户当前的地理位置为广州,而用户的经常居住地在福建,则所述 调用子模块1040调用识别福建语音特征的辅助语音数据包来识别所述语音信息。
综上所述,本发明实施例公开的一种语音识别系统,预先通过训练和学习得到多个辅助语音数据包,该辅助语音数据包是以地理位置为单位进行划分的语音数据库。同时基于用户的语音类型,辅助语音数据包进一步细分为基于方言和地理位置的辅助语音数据包,以及基于口音和地理位置的辅助语音数据包。利用基于预设模型的大词汇量语音识别模块识别用户的语音信息时,同时也利用用该辅助语音数据包识别用户的语音信息从而协助所述基于预设模型的大词汇量语音识别方法,不仅提高了用户的语音识别率,也提高了用户体验。
在本发明所提供的几个实施例中,应该理解到,所揭露的系统,装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述模块的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式。
所述作为分离部件说明的模块可以是或者也可以不是物理上分开的,作为模块显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。
另外,在本发明各个实施例中的各功能模块可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用硬件加软件功能模块的形式实现。
上述以软件功能模块的形式实现的集成的单元,可以存储在一个计算机可读取存储介质中。上述软件功能模块存储在一个存储介质中,包括若干指令用以使得一台计算机单元(可以是个人计算机,服务器,或者网络单元等)或处理器(processor)执行本发明各个实施例所述方法的部分步骤。
对于本领域技术人员而言,显然本发明不限于上述示范性实施例的细节,而且在不背离本发明的精神或基本特征的情况下,能够以其他的具体形式实现本发明。因此,无论从哪一点来看,均应将实施例看作是示范性的,而且是非限制性的,本发明的范围由所附权利要求而不是上述说明限定,因此旨在将落在权利要求的等同要件的含义和范围内的所有变化涵括在本发明内。不应将权 利要求中的任何附图标记视为限制所涉及的权利要求。此外,显然“包括”一词不排除其他单元或步骤,单数不排除复数。系统权利要求中陈述的多个单元或装置也可以由一个单元或装置通过软件或者硬件来实现。第一,第二等词语用来表示名称,而并不表示任何特定的顺序。
最后应说明的是,以上实施例仅用以说明本发明的技术方案而非限制,尽管参照较佳实施例对本发明进行了详细说明,本领域的普通技术人员应当理解,可以对本发明的技术方案进行修改或等同替换,而不脱离本发明技术方案的精神和范围。

Claims (16)

  1. 一种语音识别方法,应用于电子设备中,其特征在于,该方法包括:
    获取用户输入的语音信息;
    利用第一语音识别方法识别所述语音信息得到第一语音识别结果,利用第二语音识别方法识别所述语音信息得到第二语音识别结果;及
    根据预先设置的规则显示所述第一语音识别结果及所述第二语音识别结果。
  2. 如权利要求1所述的语音识别方法,其特征在于,所述第一语音识别方法是基于预设模型的大词汇量语音识别方法,所述第二语音识别方法是基于辅助语音数据包的语音识别方法。
  3. 如权利要求2所述的语音识别方法,其特征在于,所述基于辅助语音数据包的语音识别方法包括:
    接收到所述语音信息时,获取该用户当前的地理位置信息;
    根据所述地理位置信息调用对应的辅助语音数据包;及
    根据所述辅助语音数据包识别所述语音信息得到所述第二语音识别结果。
  4. 如权利要求3所述的语音识别方法,其特征在于,所述方法还包括:
    预先设置多个基于地理位置的语音数据包,并将所述语音数据包存储于所述电子设备中或者存储于与所述电子设备连接的服务器中。
  5. 如权利要求3所述的语音识别方法,其特征在于,在根据所述地理位置信息调用对应的辅助语音数据包之前,所述方法还包括:
    根据所述语音信息确定该用户的语音类型,所述语音类型包括口音及方言;及
    基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
  6. 如权利要求3至5任意一项所述的语音识别方法,其特征在于,所述方法还包括:
    接收到所述语音信息时,获取该用户当前的地理位置信息及历史地理位置信息;及
    根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
  7. 如权利要求6所述的语音识别方法,其特征在于,所述方法还包括:
    结合获取的用户反馈信息更新所述预先设置的规则,所述预先设置的规则包括:
    为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式;或
    为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识别结果的显示方式,
    其中,所述显示方式包括显示的时间或显示的位置。
  8. 如权利要求7所述的语音识别方法,其特征在于,所述更新所述预先设置的规则包括:
    根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分数值减小。
  9. 一种语音识别系统,运行于电子设备中,其特征在于,该系统包括:
    获取模块,用于获取用户输入的语音信息;
    第一识别模块,用于识别所述语音信息得到第一语音识别结果;
    第二识别模块,用于识别所述语音信息得到第二语音识别结果;及
    显示模块,用于根据预先设置的规则显示所述第一语音识别结果及所述第二语音识别结果。
  10. 如权利要求9所述的语音识别系统,其特征在于,所述第一语音识别模块是基于预设模型的大词汇量语音识别模块,所述第二语音识别模块是基于辅助语音数据包的语音识别模块。
  11. 如权利要求10所述的语音识别系统,其特征在于,
    所述获取模块,还用于接收到所述语音信息时,获取该用户当前的地理位置信息;
    所述第二识别模块包括:
    调用子模块,用于根据所述地理位置信息调用对应的辅助语音数据包;及
    该第二识别模块,用于根据所述辅助语音数据包识别所述语音信息得到所述第二语音识别结果。
  12. 如权利要求11所述的语音识别系统,其特征在于,所述系统还包括:
    设置模块,用于预先设置多个基于地理位置的语音数据包,并将所述语音数据包存储于所述电子设备中或者存储于与所述电子设备连接的服务器中。
  13. 如权利要求11所述的语音识别系统,其特征在于,所述系统还包括确定子模块:
    用于根据所述语音信息确定该用户的语音类型,所述语音类型包括口音及方言;及
    基于所述语音类型和所述地理位置信息共同确定对应的辅助语音数据包。
  14. 如权利要求11至13任意一项所述的语音识别系统,其特征在于,
    所述获取模块,还用于接收到所述语音信息时,获取该用户当前的地理位置信息及历史地理位置信息;及
    所述调用子模块,还用于根据历史地理位置信息和当前地理位置信息确定调用的辅助语音数据包。
  15. 如权利要求14所述的语音识别系统,其特征在于,所述系统还包括:
    更新模块,用于结合获取的用户反馈信息更新所述预先设置的规则,所述预先设置的规则是由所述设置模块设置的,包括:
    为所述第一语音识别结果预先分配第一权重,为所述第二语音识别结果预先分配第二权重,根据权重值的大小确定对应该权重值的语音识别结果的显示方式;或
    为所述第一语音识别结果预先设置第一识别分数,为所述第二语音识别结果预先设置第二识别分数,根据识别分数的大小确定对应该识别分数的语音识别结果的显示方式,
    其中,所述显示方式包括显示的时间或显示的位置。
  16. 如权利要求15所述的语音识别系统,其特征在于,所述更新模块更新所述预先设置的规则包括:
    根据用户选取的语音识别结果,将对应该语音识别结果的权重值或者识别分数值变大,及/或将用户没有选取的语音识别结果对应的权重值或者识别分 数值减小。
PCT/CN2016/097466 2016-06-21 2016-08-31 语音识别方法及系统 Ceased WO2017219495A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201610465607.6A CN106128462A (zh) 2016-06-21 2016-06-21 语音识别方法及系统
CN201610465607.6 2016-06-21

Publications (1)

Publication Number Publication Date
WO2017219495A1 true WO2017219495A1 (zh) 2017-12-28

Family

ID=57268724

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/097466 Ceased WO2017219495A1 (zh) 2016-06-21 2016-08-31 语音识别方法及系统

Country Status (2)

Country Link
CN (1) CN106128462A (zh)
WO (1) WO2017219495A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109360565A (zh) * 2018-12-11 2019-02-19 江苏电力信息技术有限公司 一种通过建立资源库提高语音识别精度的方法

Families Citing this family (24)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107316637A (zh) * 2017-05-31 2017-11-03 广东欧珀移动通信有限公司 语音识别方法及相关产品
CN107170454B (zh) * 2017-05-31 2022-04-05 Oppo广东移动通信有限公司 语音识别方法及相关产品
CN107274885B (zh) * 2017-05-31 2020-05-26 Oppo广东移动通信有限公司 语音识别方法及相关产品
WO2018228515A1 (en) 2017-06-15 2018-12-20 Beijing Didi Infinity Technology And Development Co., Ltd. Systems and methods for speech recognition
CN109145281B (zh) * 2017-06-15 2020-12-25 北京嘀嘀无限科技发展有限公司 语音识别方法、装置及存储介质
CN107767872A (zh) * 2017-10-13 2018-03-06 深圳市汉普电子技术开发有限公司 语音识别方法、终端设备及存储介质
CN107733762B (zh) * 2017-11-20 2020-07-24 宁波向往智能科技有限公司 一种智能家居的语音控制方法及装置、系统
CN107894882B (zh) * 2017-11-21 2021-02-09 南京硅基智能科技有限公司 一种移动终端的语音输入方法
CN108847222B (zh) * 2018-06-19 2020-09-08 Oppo广东移动通信有限公司 语音识别模型生成方法、装置、存储介质及电子设备
CN110875039B (zh) * 2018-08-30 2023-12-01 阿里巴巴集团控股有限公司 语音识别方法和设备
KR102863864B1 (ko) * 2018-09-20 2025-09-25 삼성전자주식회사 전자 장치, 및 이의 학습 데이터 제공 또는 획득 방법
CN108986823A (zh) * 2018-09-27 2018-12-11 深圳市易控迪智能家居科技有限公司 一种语音识别解码器及语音操作系统
CN109377990A (zh) * 2018-09-30 2019-02-22 联想(北京)有限公司 一种信息处理方法和电子设备
CN109273000B (zh) * 2018-10-11 2023-05-12 河南工学院 一种语音识别方法
CN109147762A (zh) * 2018-10-19 2019-01-04 广东小天才科技有限公司 一种语音识别方法及系统
CN111161718A (zh) * 2018-11-07 2020-05-15 珠海格力电器股份有限公司 语音识别方法、装置、设备、存储介质及空调
CN110610697B (zh) * 2019-09-12 2020-07-31 上海依图信息技术有限公司 一种语音识别方法及装置
CN110648668A (zh) * 2019-09-24 2020-01-03 上海依图信息技术有限公司 关键词检测装置和方法
CN110956955B (zh) * 2019-12-10 2022-08-05 思必驰科技股份有限公司 一种语音交互的方法和装置
CN111292727B (zh) * 2020-02-03 2023-03-24 北京声智科技有限公司 一种语音识别方法及电子设备
CN111339770B (zh) * 2020-02-18 2023-07-21 百度在线网络技术(北京)有限公司 用于输出信息的方法和装置
CN114495945B (zh) * 2020-11-12 2025-05-30 阿里巴巴集团控股有限公司 语音识别方法、装置、电子设备及计算机可读存储介质
CN114627874A (zh) 2021-06-15 2022-06-14 宿迁硅基智能科技有限公司 文本对齐方法、存储介质、电子装置
CN115148192A (zh) * 2022-06-30 2022-10-04 上海近则生物科技有限责任公司 基于方言语义提取的语音识别方法及装置

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5519809A (en) * 1992-10-27 1996-05-21 Technology International Incorporated System and method for displaying geographical information
CN101042867A (zh) * 2006-03-24 2007-09-26 株式会社东芝 语音识别设备和方法
CN103903611A (zh) * 2012-12-24 2014-07-02 联想(北京)有限公司 一种语音信息的识别方法和设备
CN104916283A (zh) * 2015-06-11 2015-09-16 百度在线网络技术(北京)有限公司 语音识别方法和装置
CN105225665A (zh) * 2015-10-15 2016-01-06 桂林电子科技大学 一种语音识别方法及语音识别装置

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8468012B2 (en) * 2010-05-26 2013-06-18 Google Inc. Acoustic model adaptation using geographic information
CN102074231A (zh) * 2010-12-30 2011-05-25 万音达有限公司 语音识别方法和语音识别系统
CN102638605A (zh) * 2011-02-14 2012-08-15 苏州巴米特信息科技有限公司 一种识别方言背景普通话的语音系统
CN103811000A (zh) * 2014-02-24 2014-05-21 中国移动(深圳)有限公司 语音识别系统及方法
CN104391673A (zh) * 2014-11-20 2015-03-04 百度在线网络技术(北京)有限公司 语音交互方法和装置

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5519809A (en) * 1992-10-27 1996-05-21 Technology International Incorporated System and method for displaying geographical information
CN101042867A (zh) * 2006-03-24 2007-09-26 株式会社东芝 语音识别设备和方法
CN103903611A (zh) * 2012-12-24 2014-07-02 联想(北京)有限公司 一种语音信息的识别方法和设备
CN104916283A (zh) * 2015-06-11 2015-09-16 百度在线网络技术(北京)有限公司 语音识别方法和装置
CN105225665A (zh) * 2015-10-15 2016-01-06 桂林电子科技大学 一种语音识别方法及语音识别装置

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109360565A (zh) * 2018-12-11 2019-02-19 江苏电力信息技术有限公司 一种通过建立资源库提高语音识别精度的方法

Also Published As

Publication number Publication date
CN106128462A (zh) 2016-11-16

Similar Documents

Publication Publication Date Title
WO2017219495A1 (zh) 语音识别方法及系统
US11114091B2 (en) Method and system for processing audio communications over a network
JP2025500981A (ja) Apiコール呼び出し及び音声応答の言語モデル予測
US10152965B2 (en) Learning personalized entity pronunciations
EP3032532B1 (en) Disambiguating heteronyms in speech synthesis
CN105895103B (zh) 一种语音识别方法及装置
US8909532B2 (en) Supporting multi-lingual user interaction with a multimodal application
US20160093298A1 (en) Caching apparatus for serving phonetic pronunciations
KR102096590B1 (ko) Gui 음성제어 장치 및 방법
CN111261144A (zh) 一种语音识别的方法、装置、终端以及存储介质
HK1225504A1 (zh) 在話音合成中消除同形異音詞的歧義
WO2013173504A1 (en) Systems and methods for interating third party services with a digital assistant
KR20170033722A (ko) 사용자의 발화 처리 장치 및 방법과, 음성 대화 관리 장치
CN112270925A (zh) 用于创建可定制对话系统引擎的平台
CN104718569A (zh) 改进语音发音
JP2021022928A (ja) 人工知能基盤の自動応答方法およびシステム
US11545140B2 (en) System and method for language-based service hailing
CN107170446A (zh) 语义处理服务器及用于语义处理的方法
JP5616390B2 (ja) 応答生成装置、応答生成方法および応答生成プログラム
CN104125548A (zh) 一种对通话语言进行翻译的方法、设备和系统
US10235133B2 (en) Tooltip surfacing with a screen reader
US20200194003A1 (en) Meeting minute output apparatus, and control program of meeting minute output apparatus
CN112820294B (zh) 语音识别方法、装置、存储介质及电子设备
CN112242143A (zh) 一种语音交互方法、装置、终端设备及存储介质
WO2018043137A1 (ja) 情報処理装置及び情報処理方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16906035

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16906035

Country of ref document: EP

Kind code of ref document: A1