WO2017012242A1 - 语音识别方法和装置 - Google Patents

语音识别方法和装置 Download PDF

Info

Publication number
WO2017012242A1
WO2017012242A1 PCT/CN2015/096596 CN2015096596W WO2017012242A1 WO 2017012242 A1 WO2017012242 A1 WO 2017012242A1 CN 2015096596 W CN2015096596 W CN 2015096596W WO 2017012242 A1 WO2017012242 A1 WO 2017012242A1
Authority
WO
WIPO (PCT)
Prior art keywords
recognition result
mute
result
voice information
recognition
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/096596
Other languages
English (en)
French (fr)
Inventor
谢延
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Baidu Online Network Technology Beijing Co Ltd
Original Assignee
Baidu Online Network Technology Beijing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Baidu Online Network Technology Beijing Co Ltd filed Critical Baidu Online Network Technology Beijing Co Ltd
Publication of WO2017012242A1 publication Critical patent/WO2017012242A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/20Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems

Definitions

  • the present invention relates to the field of speech recognition technologies, and in particular, to a speech recognition method and apparatus.
  • the speech recognition system mainly recognizes the speech by receiving the voice input by the user, thereby obtaining the speech recognition result.
  • the voice search product can not only identify the voice input by the user, but also send a search request to the search server according to the voice recognition result, and further obtain the search result.
  • the content may be many, and it is necessary to wait for a long time to obtain the recognition result after the user inputs the voice. If it is a voice search product, it needs to wait for the process of obtaining the recognition result, and then wait for the process of obtaining the search result, and the waiting time is long, resulting in a decrease in the user experience.
  • the end point of the speech is not detected or the recognition result is inaccurate.
  • an object of the present invention is to provide a speech recognition method capable of reducing user waiting time and improving user experience.
  • a second object of the present invention is to provide a speech recognition apparatus.
  • a third object of the invention is to propose an apparatus.
  • a fourth object of the present invention is to provide a non-volatile computer storage medium.
  • a first aspect of the present invention provides a voice recognition method, including the following steps: S1: receiving voice information input by a user, and identifying the voice information in real time; S2, when the voice information is When mute is generated, determining the type of the mute; S3, if the mute is short mute, obtaining a first recognition result, and displaying the first recognition result while continuing to perform step S1; and S4 if the mute To be muted for a long time, a second recognition result is obtained, and the second recognition result is displayed.
  • the voice recognition method of the embodiment of the present invention receives the voice information input by the user, and recognizes the voice information in real time.
  • the type of the silence is determined. If the voice is short, the first recognition result is obtained.
  • the first recognition result is displayed, and the voice information input by the user is continuously received. If the mute is long mute, the second recognition result is obtained, and the second recognition result is displayed, which can effectively reduce the waiting time of the user and improve the user experience.
  • the second aspect of the present invention provides a voice recognition apparatus, including: a receiving module, configured to receive voice information input by a user, and identify the voice information in real time; and a determining module, configured to generate the voice information When the mute is muted, the type of the mute is determined; the first identifying module is configured to obtain a first recognition result when the mute is short mute, and display the first recognition result, and the receiving module continues to receive the search user.
  • the input voice information; the second identification module is configured to obtain a second recognition result when the mute is long mute, and display the second recognition result.
  • the voice recognition device of the embodiment of the present invention receives the voice information input by the user, and recognizes the voice information in real time.
  • the type of the silence is determined. If the voice is short, the first recognition result is obtained.
  • the first recognition result is displayed, and the voice information input by the user is continuously received. If the mute is long mute, the second recognition result is obtained, and the second recognition result is displayed, which can effectively reduce the waiting time of the user and improve the user experience.
  • a third aspect of the present invention provides an apparatus comprising: one or more processors; a memory; one or more programs, the one or more programs being stored in the memory when The speech recognition method of the embodiment of the first aspect of the present invention is executed when a plurality of processors are executed.
  • a fourth aspect of the present invention provides a non-volatile computer storage medium storing one or more programs, when the one or more programs are executed by one device, causing the device A speech recognition method in accordance with an embodiment of the first aspect of the present invention is performed.
  • FIG. 1 is a flow chart of a speech recognition method in accordance with one embodiment of the present invention.
  • FIG. 2 is a flow chart of a speech recognition method in accordance with an embodiment of the present invention.
  • FIG. 3 is a schematic diagram of an effect of an initialization interface according to an embodiment of the present invention.
  • FIG. 4 is a schematic diagram showing the effect of a prompt interface according to an embodiment of the present invention.
  • FIG. 5 is a schematic diagram of an effect of receiving a voice information interface input by a user according to an embodiment of the present invention
  • FIG. 6 is a schematic diagram 1 showing an effect of displaying a recognition result interface according to an embodiment of the present invention
  • FIG. 7 is a second schematic diagram of an effect of displaying a recognition result interface according to an embodiment of the present invention.
  • FIG. 8 is a third schematic diagram of an effect of displaying a recognition result interface according to an embodiment of the present invention.
  • FIG. 9 is a schematic diagram of an interface effect of performing a search according to a recognition result according to an embodiment of the present invention.
  • FIG. 10 is a schematic diagram showing an interface effect of displaying search results according to an embodiment of the present invention.
  • FIG. 11 is a schematic diagram 1 of an interface effect of performing a search according to a recognition result according to an embodiment of the present invention
  • FIG. 12 is a second schematic diagram of an interface effect performed according to a recognition result according to an embodiment of the present invention.
  • FIG. 13 is a third schematic diagram of an interface effect performed according to a recognition result according to an embodiment of the present invention.
  • FIG. 14 is a fourth schematic diagram of an interface effect performed according to a recognition result according to an embodiment of the present invention.
  • FIG. 1 is a flow chart of a speech recognition method in accordance with one embodiment of the present invention.
  • the speech recognition method may include:
  • S1 Receive voice information input by the user, and identify the voice information in real time.
  • the voice information may be a phrase or a short sentence.
  • the mute detection algorithm may detect the mute and determine the type of mute.
  • the type of mute can include long mute and short mute. Short mute is a short pause for the user to input voice information, while long mute is the end point (tail point) for the user to input voice information.
  • voice samples can be collected in different environments and the tail point detection model can be trained. Then, when the voice information is recognized, the type of silence can be judged by the tail point detection model, and the type of silence can be accurately determined under the noise environment, and the noise resistance and accuracy are improved.
  • the server-side tail-point detection algorithm has more powerful computing power, and can continuously optimize the tail-point detection model.
  • the local tail point detection algorithm may be used for detection. If the end point of the voice information cannot be detected, the tail end detection algorithm of the server is used for detection. .
  • step S3 If the mute is short mute, the first recognition result is obtained, and the first recognition result is displayed, and step S1 is continued.
  • the voice information can be recognized in real time.
  • the silence occurs, if the currently occurring silence is short silence, that is, the user inputs a short pause of the voice information, the first recognition result can be obtained. Then, the first recognition result is displayed on the screen of the client and fed back to the user.
  • the first recognition result may be content between the start of the input voice information and the short silence, or may be the content between the two short silences.
  • the user Still continue to input voice messages. That is to say, the identification process is performed synchronously with the process of receiving the voice information, that is, two separate threads that do not interfere with each other are processed in parallel, which reduces the waiting time of the user.
  • the user While inputting the voice information, the user has already displayed a part of the recognition result on the screen of the client. Since the short silent time is short, the effect displayed on the screen of the client is equivalent to the user inputting the voice information while being dynamically continuous.
  • the recognition result is continuously displayed, which solves the problem that the waiting time for the overall recognition of the voice information is too long after waiting for the user to input the voice information in the traditional voice recognition, thereby improving the user experience.
  • the first recognition result may be searched as a keyword, and the first search result is acquired.
  • the recognition system is a voice search system
  • the search can be performed based on the recognition result recognized in real time.
  • the second recognition result is obtained, and then the second recognition result is displayed on the screen of the client, and is fed back to the user.
  • the second recognition result may be the content between the last short mute and the long mute.
  • the voice information input by the user is not short mute
  • the second recognition result may be the content between the input speech information start and the long mute.
  • the voice information input by the user is recognized in real time, and when the screen of the client displays the first recognition result, the voice information input by the user is also received, and the voice information is recognized in real time, thereby reducing the waiting time of the user. the goal of.
  • the first recognition result can also be compared with the second recognition result. If the first recognition result is consistent with the second recognition result, the first search result may be used as the final search result.
  • the first recognition result is a recognition result corresponding to when the voice information generates a short silence
  • the second recognition result is a recognition result corresponding to the voice information when the long silence is generated.
  • the second recognition result usually requires a long silence, and when it is judged whether the current silence is long silence, it has been silenced and voice recognition, the first recognition result is obtained, and the corresponding first search is obtained. result. After determining that the muting is long muted, if the first recognition result and the second recognition result are consistent, the first search result may be directly used as the final search result without searching the second recognition result again as a keyword, thereby Saves users time waiting.
  • the first recognition result and the second recognition result may be spliced to generate a final recognition result, and the recognition result is searched as a keyword to obtain a final search result.
  • search results can be displayed on the client's screen to feed back to the user.
  • the voice recognition method of the embodiment of the present invention receives the voice information input by the user, and recognizes the voice information in real time.
  • the type of the silence is determined. If the voice is short, the first recognition result is obtained.
  • the first recognition result is displayed, and the voice information input by the user is continuously received. If the mute is long mute, the second recognition result is obtained, and the second recognition result is displayed, which can effectively reduce the waiting time of the user and improve the user experience.
  • FIG. 2 is a flowchart of a voice recognition method according to an embodiment of the present invention. This embodiment uses a search APP as an example for detailed description.
  • the voice recognition method may include:
  • the running environment can be initialized.
  • the prompt interface as shown in FIG. 4 can be displayed.
  • S203 Receive voice information input by the user, and identify the voice information in real time.
  • the words "listening" may be displayed in the interface, indicating that the voice information input by the user is being received, and the input voice information is being recognized at the same time.
  • the voice information input by the user is “Baidu Voice Provisioning Technology”, and when a short silence is detected when inputting to “Baidu”, the corresponding recognition result “Baidu” can be obtained and displayed, as shown in FIG. 6.
  • the voice information input by the user is also received, and the voice information is recognized in real time.
  • the user inputs to "speech” a short mute is detected, and the corresponding recognition result "speech” can be obtained and displayed, as shown in FIG.
  • the method for detecting short mute and long mute used here is a cue point detection algorithm, which is consistent with the description in the previous embodiment, and therefore will not be described here.
  • the user when the user inputs the voice information “Baidu voice providing technology”, it can detect that a long mute is generated, and the corresponding recognition result “providing technology” can be obtained and displayed. Since “Baidu”, “Voice”, and “Providing Technology” are displayed one after another, and the time interval is short, the effect is equivalent to the user inputting the voice information while continuously displaying the recognition result on the screen of the client, and finally The “Baidu Voice Provisioning Technology” is displayed, as shown in Figure 8.
  • each segment of the recognition result may be spliced to generate a keyword "Baidu Voice Provisioning Technology", and a search request is sent to the search server.
  • the state in the interface can be displayed as "Processing”. Then, after obtaining the search result corresponding to the "Baidu Voice Providing Technology" by the search server, as shown in FIG. 10, the search result is displayed.
  • the speech recognition method of the embodiment of the present invention segments the voice information input by the user by using the tail point detection algorithm, and can accurately determine the pause point or the end point of the voice information input by the user, thereby improving the noise resistance of the voice recognition and Accuracy; by recognizing the voice information in real time, the recognized part can be displayed while the user inputs the voice information, reducing the waiting time of the user; reducing the entire voice by processing the recognition process and the search process in parallel Identify the response time of the search system, which in turn improves the user experience.
  • the present invention also proposes a speech recognition apparatus.
  • FIG. 11 is a first schematic structural diagram of a voice recognition apparatus according to an embodiment of the present invention.
  • the voice recognition apparatus may include: a receiving module 110, a determining module 120, a first identifying module 130, and a second identifying module 140.
  • the receiving module 110 is configured to receive voice information input by the user, and identify the voice information in real time.
  • the voice information may be a phrase or a short sentence.
  • the judging module 120 is configured to determine the type of muting when the voice information is muted.
  • the determining module 120 may detect the silence according to the tail point detection algorithm and determine the type of the silence.
  • the type of mute can include long mute and short mute. Short mute is a short pause for the user to input voice information, while long mute is the end point (tail point) for the user to input voice information.
  • voice samples can be collected in different environments and the tail point detection model can be trained. Then, when the voice information is recognized, the type of silence can be judged by the tail point detection model, and the type of silence can be accurately determined under the noise environment, and the noise resistance and accuracy are improved.
  • the server-side tail-point detection algorithm has more powerful computing power, and can continuously optimize the tail-point detection model.
  • the local tail point detection algorithm may be used for detection. If the end point of the voice information cannot be detected, the tail end detection algorithm of the server is used for detection. .
  • the first identification module 130 is configured to obtain a first recognition result when the mute is short mute, and display the first recognition result, and the receiving module continues to receive the voice information input by the search user.
  • the voice information can be recognized in real time.
  • the mute is present, if the currently occurring mute is short mute, that is, the user inputs a short pause of the voice information, the first identification module 130 can The first recognition result is obtained, and then the first recognition result is displayed on the screen of the client, and is fed back to the user.
  • the first recognition result may be content between the start of the input voice information and the short silence, or may be the content between the two short silences.
  • the user continues to input voice messages.
  • the identification process is performed synchronously with the process of receiving the voice information, that is, two separate threads that do not interfere with each other are processed in parallel, which reduces the waiting time of the user.
  • the user While inputting the voice information, the user has already displayed a part of the recognition result on the screen of the client. Since the short silent time is short, the effect displayed on the screen of the client is equivalent to the user inputting the voice information while being dynamically continuous. continuously The recognition result is displayed, which solves the problem that the waiting time of the voice information is too long after waiting for the user to input the voice information in the traditional voice recognition, thereby improving the user experience. .
  • the second identification module 140 is configured to obtain a second recognition result when the mute is long mute, and display the second recognition result.
  • the second recognition module 140 may obtain the second recognition result, and then display the second recognition result on the screen of the client, and feed back to the user.
  • the second recognition result may be the content between the last short mute and the long mute.
  • the voice information input by the user is not short mute
  • the second recognition result may be the content between the input speech information start and the long mute.
  • the voice information input by the user is recognized in real time, and when the screen of the client displays the first recognition result, the voice information input by the user is also received, and the voice information is recognized in real time, thereby reducing the waiting time of the user. the goal of.
  • the voice recognition apparatus of the embodiment of the present invention may further include a search module 150.
  • the searching module 150 is configured to search the first recognition result as a keyword after the first identification module 130 obtains the first recognition result, and acquire the first search result. For example, when the recognition system is a voice search system, the search can be performed based on the recognition result recognized in real time.
  • the voice recognition apparatus of the embodiment of the present invention may further include a processing module 160.
  • the processing module 160 is configured to compare the first recognition result with the second recognition result, and if the first recognition result is consistent with the second recognition result, the first search result is used as a final search result, and if the first recognition result and the first If the two recognition results are inconsistent, the first recognition result is spliced with the second recognition result to generate a final recognition result, and the recognition result is searched as a keyword to obtain a final search result.
  • the first recognition result is a recognition result corresponding to when the voice information generates a short silence
  • the second recognition result is a recognition result corresponding to the voice information when the long silence is generated.
  • the second recognition result usually requires a long silence, and when it is judged whether the current silence is long silence, it has been silenced and voice recognition, the first recognition result is obtained, and the corresponding first search is obtained. result. After determining that the muting is long muted, if the first recognition result and the second recognition result are consistent, the first search result may be directly used as the final search result without searching the second recognition result again as a keyword, thereby Saves users time waiting.
  • the voice recognition apparatus of the embodiment of the present invention may further include a display module 170.
  • the display module 170 is configured to display the search result after obtaining the final search result.
  • the voice recognition device of the embodiment of the present invention receives the voice information input by the user, and recognizes the voice information in real time.
  • the type of the silence is determined. If the voice is short, the first recognition result is obtained.
  • the first recognition result is displayed, and the voice information input by the user is continuously received. If the mute is long mute, the second recognition result is obtained, and the second recognition result is displayed, which can effectively reduce the waiting time of the user and improve the user experience.
  • first and second are used for descriptive purposes only and are not to be construed as indicating or implying a relative importance or implicitly indicating the number of technical features indicated.
  • features defining “first” or “second” may include at least one of the features, either explicitly or implicitly.
  • the meaning of "a plurality” is at least two, such as two, three, etc., unless specifically defined otherwise.
  • the terms “installation”, “connected”, “connected”, “fixed” and the like shall be understood broadly, and may be either a fixed connection or a detachable connection, unless explicitly stated and defined otherwise. , or integrated; can be mechanical or electrical connection; can be directly connected, or indirectly connected through an intermediate medium, can be the internal communication of two elements or the interaction of two elements, unless otherwise specified Limited.
  • the specific meanings of the above terms in the present invention can be understood on a case-by-case basis.
  • the first feature "on” or “under” the second feature may be a direct contact of the first and second features, or the first and second features may be indirectly through an intermediate medium, unless otherwise explicitly stated and defined. contact.
  • the first feature "above”, “above” and “above” the second feature may be that the first feature is directly above or above the second feature, or merely that the first feature level is higher than the second feature.
  • the first feature “below”, “below” and “below” the second feature may be that the first feature is directly below or obliquely below the second feature, or merely that the first feature level is less than the second feature.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • User Interface Of Digital Computer (AREA)
  • Telephonic Communication Services (AREA)

Abstract

一种语音识别方法和装置,其中,该方法包括:接收用户输入的语音信息,并实时对语音信息进行识别(S1);当语音信息产生静音时,判断静音的类型(S2);如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续执行步骤S1(S3);以及如果静音为长静音,则获得第二识别结果,并显示第二识别结果(S4)。该语音识别方法和装置,通过实时对用户输入的语音信息进行识别,能够降低用户等待时间,提升用户使用体验。

Description

语音识别方法和装置
相关申请的交叉引用
本申请要求百度在线网络技术(北京)有限公司于2015年7月22日提交的、发明名称为“语音识别方法和装置”的、中国专利申请号“201510435887.1”的优先权。
技术领域
本发明涉及语音识别技术领域,尤其涉及一种语音识别方法和装置。
背景技术
随着科技的不断进步,语音识别技术的应用也越来越广泛,例如工业、家电、通信、汽车电子、医疗、家庭服务、消费电子产品等领域,都会应用到语音识别技术。目前,语音识别系统主要通过接收用户输入的语音,对语音进行识别,从而获得语音识别结果。其中,语音搜索类产品不仅可以对用户输入的语音进行识别,还可根据语音识别结果向搜索服务器发送搜索请求,进一步获取搜索结果。
但是,有时候用户输入语音时,内容可能很多,则需要在用户输入语音结束后,等待很长时间才能获取到识别结果。如果是语音搜索类产品,则需要先等待获得识别结果的过程,再等待获取搜索结果的过程,等待时间长,导致用户体验降低。另外,在噪声环境中,由于噪声干扰,有可能出现检测不到语音结束点或者识别结果不准确的情况。
发明内容
本发明旨在至少在一定程度上解决相关技术中的技术问题之一。为此,本发明的一个目的在于提出一种语音识别方法,该方法能够降低用户等待时间,提升用户使用体验。
本发明的第二个目的在于提出一种语音识别装置。
本发明的第三个目的在于提出一种设备。
本发明的第四个目的在于提出一种非易失性计算机存储介质。
为了实现上述目的,本发明第一方面实施例提出了一种语音识别方法,包括以下步骤:S1、接收用户输入的语音信息,并实时对所述语音信息进行识别;S2、当所述语音信息产生静音时,判断所述静音的类型;S3、如果所述静音为短静音,则获得第一识别结果,并显示所述第一识别结果,同时继续执行步骤S1;以及S4、如果所述静音为长静音,则获得第二识别结果,并显示所述第二识别结果。
本发明实施例的语音识别方法,通过接收用户输入的语音信息,并实时对语音信息进行识别,当语音信息产生静音时,判断静音的类型,如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续接收用户输入的语音信息,如果静音为长静音,则获得第二识别结果,并显示第二识别结果,能够有效地降低用户等待时间,提升用户使用体验。
本发明第二方面实施例提出了一种语音识别装置,包括:接收模块,用于接收用户输入的语音信息,并实时对所述语音信息进行识别;判断模块,用于当所述语音信息产生静音时,判断所述静音的类型;第一识别模块,用于当所述静音为短静音时,获得第一识别结果,并显示所述第一识别结果,同时所述接收模块继续接收搜索用户输入的语音信息;第二识别模块,用于当所述静音为长静音时,获得第二识别结果,并显示所述第二识别结果。
本发明实施例的语音识别装置,通过接收用户输入的语音信息,并实时对语音信息进行识别,当语音信息产生静音时,判断静音的类型,如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续接收用户输入的语音信息,如果静音为长静音,则获得第二识别结果,并显示第二识别结果,能够有效地降低用户等待时间,提升用户使用体验。
本发明第三方面实施例提供了一种设备,包括:一个或者多个处理器;存储器;一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,执行本发明第一方面实施例的语音识别方法。
本发明第四方面实施例提供了一种非易失性计算机存储介质,所述计算机存储介质存储有一个或者多个程序,当所述一个或者多个程序被一个设备执行时,使得所述设备执行以本发明第一方面实施例的语音识别方法。
附图说明
图1是根据本发明一个实施例的语音识别方法的流程图;
图2是根据本发明一个具体实施例的语音识别方法的流程图;
图3是根据本发明一个具体实施例的初始化界面效果示意图;
图4是根据本发明一个具体实施例的提示界面效果示意图;
图5是根据本发明一个具体实施例的接收用户输入的语音信息界面效果示意图;
图6是根据本发明一个具体实施例的显示识别结果界面效果示意图一;
图7是根据本发明一个具体实施例的显示识别结果界面效果示意图二;
图8是根据本发明一个具体实施例的显示识别结果界面效果示意图三;
图9是根据本发明一个具体实施例的根据识别结果进行搜索的界面效果示意图;
图10是根据本发明一个具体实施例的显示搜索结果的界面效果示意图;
图11是根据本发明一个具体实施例的根据识别结果进行搜索的界面效果示意图一;
图12是根据本发明一个具体实施例的根据识别结果进行搜索的界面效果示意图二;
图13是根据本发明一个具体实施例的根据识别结果进行搜索的界面效果示意图三;
图14是根据本发明一个具体实施例的根据识别结果进行搜索的界面效果示意图四。
具体实施方式
下面详细描述本发明的实施例,所述实施例的示例在附图中示出,其中自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面通过参考附图描述的实施例是示例性的,旨在用于解释本发明,而不能理解为对本发明的限制。
下面参考附图描述本发明实施例的语音识别方法和装置。
图1是根据本发明一个实施例的语音识别方法的流程图。
如图1所示,语音识别方法可包括:
S1、接收用户输入的语音信息,并实时对语音信息进行识别。
其中,语音信息可以为词组,也可以为短句。
S2、当语音信息产生静音时,判断静音的类型。
在本发明的实施例中,为解决在噪声环境中,静音检测不准确的问题,可根据尾点检测算法检测出静音,并判断静音的类型。其中,静音的类型可包括长静音和短静音。短静音为用户输入语音信息的短暂停顿,而长静音则为用户输入语音信息的结束点(尾点)。
具体地,可先在不同环境下采集语音样本,并训练尾点检测模型。然后在对语音信息进行识别时,可通过尾点检测模型判断静音的类型,在噪声环境下能够准确地判断出静音的类型,提高了抗噪性和准确率。相对于本地的尾点检测算法,服务器端的尾点检测算法具有更强大的计算能力,可不断地对尾点检测模型进行优化。在本发明一个实施例中,在对语音信息识别的过程中,可先通过本地的尾点检测算法进行检测,如果无法检测出语音信息的结束点,则再通过服务器端的尾点检测算法进行检测。
S3、如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续执行步骤S1。
具体地,在用户输入语音信息开始时,可实时地对语音信息进行识别,当出现静音时,如果当前出现的静音为短静音,即用户输入语音信息的短暂停顿,则可获得第一识别结果,然后将第一识别结果显示在客户端的屏幕上,反馈给用户。其中,第一识别结果可以为输入语音信息开始至短静音之间的内容,也可以是两个短静音之间的内容。与此同时,用户 还在继续输入语音信息。也就是说,识别过程与接收语音信息过程同步进行,即两个单独且互不干扰的线程并行处理,减少了用户等待的时间。用户在输入语音信息的同时,已经在客户端的屏幕上显示出了一部分的识别结果,由于短静音时间很短,因此在客户端的屏幕上显示的效果相当于用户一边输入语音信息,同时动态地连续不断地显示出识别结果,解决了传统的语音识别中,等待用户输入语音信息结束后,再对语音信息进行整体识别所带来的等待时间过长的问题,提升了用户使用体验。
此外,在获得第一识别结果之后,还可将第一识别结果作为关键词进行搜索,并获取第一搜索结果。例如:识别系统为语音搜索系统时,可根据实时识别出的识别结果进行搜索。
S4、如果静音为长静音,则获得第二识别结果,并显示第二识别结果。
具体地,如果当前出现的静音为长静音,即用户输入语音信息结束,则可获得第二识别结果,然后将第二识别结果显示在客户端的屏幕上,反馈给用户。其中,第二识别结果可以是最后一个短静音与长静音之间的内容,如果用户输入的语音信息没有短静音,则第二识别结果可以为输入语音信息开始与长静音之间的内容。举例来说,实时地对用户输入的语音信息进行识别,当客户端的屏幕显示第一识别结果时,同时还在接收用户输入的语音信息,并实时地对语音信息识别,从而达到减少用户等待时间的目的。
另外,还可将第一识别结果与第二识别结果进行对比。若第一识别结果与第二识别结果一致,则可将第一搜索结果作为最终搜索结果。具体地,第一识别结果为语音信息产生短静音时对应的识别结果,第二识别结果为语音信息产生长静音时对应的识别结果。而获得第二识别结果通常需要一个长静音,而在判断当前静音是否为长静音时,已经将其作为短静音并进行了语音识别,获得了第一识别结果,并获取了对应的第一搜索结果。当确定该静音为长静音后,如果第一识别结果和第二识别结果一致,则可直接将第一搜索结果作为最终的搜索结果,而无需将第二识别结果作为关键词再次进行搜索,从而节省了用户等待的时间。
若第一识别结果与第二识别结果不一致,则可将第一识别结果与第二识别结果进行拼接,生成最终的识别结果,并将识别结果作为关键词进行搜索,以获取最终的搜索结果。
在确定最终的搜索结果后,可在客户端的屏幕显示搜索结果,以反馈给用户。
本发明实施例的语音识别方法,通过接收用户输入的语音信息,并实时对语音信息进行识别,当语音信息产生静音时,判断静音的类型,如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续接收用户输入的语音信息,如果静音为长静音,则获得第二识别结果,并显示第二识别结果,能够有效地降低用户等待时间,提升用户使用体验。
图2是根据本发明一个具体实施例的语音识别方法的流程图,本实施例以搜索APP为例进行详细描述。
如图2所示,语音识别方法可包括:
S201,开启搜索APP,并进行初始化。
如图3所示,在开启终端中的搜索APP时,可对运行环境进行初始化。
S202,显示提示界面。
在初始化结束后,可显示如图4所示的提示界面。
S203,接收用户输入的语音信息,并实时对语音信息进行识别。
当检测到有用户输入语音信息时,如图5所示,可在界面中显示如“倾听中”字样,表示正在接收用户输入的语音信息,与此同时正在对输入的语音信息进行识别。
S204,当产生短静音时,获得并显示第一识别结果。
例如,用户输入的语音信息为“百度语音提供技术”,而输入到“百度”时,检测到一个短静音,则可获得并显示对应的识别结果“百度”,如图6所示。与此同时,还在接收用户输入的语音信息,且实时地对语音信息进行识别。依此类推,当用户输入到“语音”时,又检测到一个短静音,此时可获得并显示对应的识别结果“语音”,如图7所示。
此外,在识别出“百度”的同时,还可以“百度”为关键词,向搜索服务器发送搜索请求,获得“百度”对应的搜索结果。以此类推,在识别出“语音”的同时,还可以“百度语音”为关键词,向搜索服务器发送搜索请求,获得“百度语音”对应的搜索结果。
此处检测短静音和长静音使用的方法为尾点检测算法,与上一实施例中的描述一致,故此处不赘述。
S205,当产生长静音时,显示第二识别结果。
例如:当用户输入语音信息“百度语音提供技术”结束时,可检测到产生长静音,则可获得并显示对应的识别结果“提供技术”。由于“百度”、“语音”、“提供技术”是先后显示的,且时间间隔很短,则其效果相当于用户一边输入语音信息,一边连续不断地在客户端的屏幕上显示出识别结果,最终显示出“百度语音提供技术”,如图8所示。
S206,将第一识别结果和第二识别结果进行拼接,以生成搜索词,并进行搜索。
在识别结束后,可将每段识别结果进行拼接,生成关键词“百度语音提供技术”,并向搜索服务器发送搜索请求。
S207,获得搜索词对应的搜索结果,并显示搜索结果。
具体地,如图9所示,在根据关键词“百度语音提供技术”进行搜索时,界面中的状态可显示为“处理中”。然后,在通过搜索服务器获得“百度语音提供技术”对应的搜索结果后,如图10所示,显示该搜索结果。
本发明实施例的语音识别方法,通过尾点检测算法对用户输入的语音信息进行分段,能够准确地判断出用户输入的语音信息的暂停点或结束点,提升了语音识别的抗噪性和准确性;通过实时地对语音信息进行识别,可在用户输入语音信息的同时即可显示出已识别的部分,减少了用户等待的时间;通过将识别过程和搜索过程并行处理,降低了整个语音识别搜索系统的响应时间,进而提高了用户使用体验。
为实现上述目的,本发明还提出一种语音识别装置。
图11是根据本发明一个实施例的语音识别装置的结构示意图一。
如图11所示,该语音识别装置可包括:接收模块110、判断模块120、第一识别模块130和第二识别模块140。
其中,接收模块110用于接收用户输入的语音信息,并实时对语音信息进行识别。
其中,语音信息可以为词组,也可以为短句。
判断模块120用于当语音信息产生静音时,判断静音的类型。
在本发明的实施例中,为解决在噪声环境中,静音检测不准确的问题,判断模块120可根据尾点检测算法检测出静音,并判断静音的类型。其中,静音的类型可包括长静音和短静音。短静音为用户输入语音信息的短暂停顿,而长静音则为用户输入语音信息的结束点(尾点)。
具体地,可先在不同环境下采集语音样本,并训练尾点检测模型。然后在对语音信息进行识别时,可通过尾点检测模型判断静音的类型,在噪声环境下能够准确地判断出静音的类型,提高了抗噪性和准确率。相对于本地的尾点检测算法,服务器端的尾点检测算法具有更强大的计算能力,可不断地对尾点检测模型进行优化。在本发明一个实施例中,在对语音信息识别的过程中,可先通过本地的尾点检测算法进行检测,如果无法检测出语音信息的结束点,则再通过服务器端的尾点检测算法进行检测。
第一识别模块130用于当静音为短静音时,获得第一识别结果,并显示第一识别结果,同时接收模块继续接收搜索用户输入的语音信息。
具体地,在用户输入语音信息开始时,可实时地对语音信息进行识别,当出现静音时,如果当前出现的静音为短静音,即用户输入语音信息的短暂停顿,则第一识别模块130可获得第一识别结果,然后将第一识别结果显示在客户端的屏幕上,反馈给用户。其中,第一识别结果可以为输入语音信息开始至短静音之间的内容,也可以是两个短静音之间的内容。与此同时,用户还在继续输入语音信息。也就是说,识别过程与接收语音信息过程同步进行,即两个单独且互不干扰的线程并行处理,减少了用户等待的时间。用户在输入语音信息的同时,已经在客户端的屏幕上显示出了一部分的识别结果,由于短静音时间很短,因此在客户端的屏幕上显示的效果相当于用户一边输入语音信息,同时动态地连续不断地 显示出识别结果,解决了传统的语音识别中,等待用户输入语音信息结束后,再对语音信息进行整体识别所带来的等待时间过长的问题,提升了用户使用体验。。
第二识别模块140用于当静音为长静音时,获得第二识别结果,并显示第二识别结果。
具体地,如果当前出现的静音为长静音,即用户输入语音信息结束,则第二识别模块140可获得第二识别结果,然后将第二识别结果显示在客户端的屏幕上,反馈给用户。其中,第二识别结果可以是最后一个短静音与长静音之间的内容,如果用户输入的语音信息没有短静音,则第二识别结果可以为输入语音信息开始与长静音之间的内容。举例来说,实时地对用户输入的语音信息进行识别,当客户端的屏幕显示第一识别结果时,同时还在接收用户输入的语音信息,并实时地对语音信息识别,从而达到减少用户等待时间的目的。
另外,如图12所示,本发明实施例的语音识别装置还可包括搜索模块150。
搜索模块150用于在第一识别模块130获得第一识别结果之后,将第一识别结果作为关键词进行搜索,并获取第一搜索结果。例如:识别系统为语音搜索系统时,可根据实时识别出的识别结果进行搜索。
此外,如图13所示,本发明实施例的语音识别装置还可包括处理模块160。
处理模块160用于将第一识别结果与第二识别结果进行对比,若第一识别结果与第二识别结果一致,则将第一搜索结果作为最终的搜索结果,以及若第一识别结果与第二识别结果不一致,则将第一识别结果与第二识别结果进行拼接,生成最终的识别结果,并将识别结果作为关键词进行搜索,以获取最终的搜索结果。
具体地,第一识别结果为语音信息产生短静音时对应的识别结果,第二识别结果为语音信息产生长静音时对应的识别结果。而获得第二识别结果通常需要一个长静音,而在判断当前静音是否为长静音时,已经将其作为短静音并进行了语音识别,获得了第一识别结果,并获取了对应的第一搜索结果。当确定该静音为长静音后,如果第一识别结果和第二识别结果一致,则可直接将第一搜索结果作为最终的搜索结果,而无需将第二识别结果作为关键词再次进行搜索,从而节省了用户等待的时间。
进一步地,如图14所示,本发明实施例的语音识别装置还可包括显示模块170。
显示模块170用于在获取最终的搜索结果之后,显示搜索结果。
本发明实施例的语音识别装置,通过接收用户输入的语音信息,并实时对语音信息进行识别,当语音信息产生静音时,判断静音的类型,如果静音为短静音,则获得第一识别结果,并显示第一识别结果,同时继续接收用户输入的语音信息,如果静音为长静音,则获得第二识别结果,并显示第二识别结果,能够有效地降低用户等待时间,提升用户使用体验。
在本发明的描述中,需要理解的是,术语“中心”、“纵向”、“横向”、“长度”、 “宽度”、“厚度”、“上”、“下”、“前”、“后”、“左”、“右”、“竖直”、“水平”、“顶”、“底”“内”、“外”、“顺时针”、“逆时针”、“轴向”、“径向”、“周向”等指示的方位或位置关系为基于附图所示的方位或位置关系,仅是为了便于描述本发明和简化描述,而不是指示或暗示所指的装置或元件必须具有特定的方位、以特定的方位构造和操作,因此不能理解为对本发明的限制。
此外,术语“第一”、“第二”仅用于描述目的,而不能理解为指示或暗示相对重要性或者隐含指明所指示的技术特征的数量。由此,限定有“第一”、“第二”的特征可以明示或者隐含地包括至少一个该特征。在本发明的描述中,“多个”的含义是至少两个,例如两个,三个等,除非另有明确具体的限定。
在本发明中,除非另有明确的规定和限定,术语“安装”、“相连”、“连接”、“固定”等术语应做广义理解,例如,可以是固定连接,也可以是可拆卸连接,或成一体;可以是机械连接,也可以是电连接;可以是直接相连,也可以通过中间媒介间接相连,可以是两个元件内部的连通或两个元件的相互作用关系,除非另有明确的限定。对于本领域的普通技术人员而言,可以根据具体情况理解上述术语在本发明中的具体含义。
在本发明中,除非另有明确的规定和限定,第一特征在第二特征“上”或“下”可以是第一和第二特征直接接触,或第一和第二特征通过中间媒介间接接触。而且,第一特征在第二特征“之上”、“上方”和“上面”可是第一特征在第二特征正上方或斜上方,或仅仅表示第一特征水平高度高于第二特征。第一特征在第二特征“之下”、“下方”和“下面”可以是第一特征在第二特征正下方或斜下方,或仅仅表示第一特征水平高度小于第二特征。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本发明的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不必须针对的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任一个或多个实施例或示例中以合适的方式结合。此外,在不相互矛盾的情况下,本领域的技术人员可以将本说明书中描述的不同实施例或示例以及不同实施例或示例的特征进行结合和组合。
尽管上面已经示出和描述了本发明的实施例,可以理解的是,上述实施例是示例性的,不能理解为对本发明的限制,本领域的普通技术人员在本发明的范围内可以对上述实施例进行变化、修改、替换和变型。

Claims (12)

  1. 一种语音识别方法,其特征在于,包括以下步骤:
    S1、接收用户输入的语音信息,并实时对所述语音信息进行识别;
    S2、当所述语音信息产生静音时,判断所述静音的类型;
    S3、如果所述静音为短静音,则获得第一识别结果,并显示所述第一识别结果,同时继续执行步骤S1;以及
    S4、如果所述静音为长静音,则获得第二识别结果,并显示所述第二识别结果。
  2. 如权利要求1所述的方法,其特征在于,在获得所述第一识别结果之后,还包括:
    将所述第一识别结果作为关键词进行搜索,并获取第一搜索结果。
  3. 如权利要求1或2所述的方法,其特征在于,还包括:
    将所述第一识别结果与所述第二识别结果进行对比;
    若所述第一识别结果与所述第二识别结果一致,则将所述第一搜索结果作为最终的搜索结果;
    若所述第一识别结果与所述第二识别结果不一致,则将所述第一识别结果与所述第二识别结果进行拼接,生成最终的识别结果,并将所述识别结果作为所述关键词进行搜索,以获取最终的所述搜索结果。
  4. 如权利要求3所述的方法,其特征在于,在获取最终的所述搜索结果之后,还包括:
    显示所述搜索结果。
  5. 如权利要求1-4任一项所述的方法,其特征在于,所述判断所述静音的类型,包括:
    根据尾点检测算法判断所述静音的类型。
  6. 一种语音识别装置,其特征在于,包括:
    接收模块,用于接收用户输入的语音信息,并实时对所述语音信息进行识别;
    判断模块,用于当所述语音信息产生静音时,判断所述静音的类型;
    第一识别模块,用于当所述静音为短静音时,获得第一识别结果,并显示所述第一识别结果,同时所述接收模块继续接收搜索用户输入的语音信息;
    第二识别模块,用于当所述静音为长静音时,获得第二识别结果,并显示所述第二识别结果。
  7. 如权利要求6所述的装置,其特征在于,还包括:
    搜索模块,用于在获得所述第一识别结果之后,将所述第一识别结果作为关键词进行搜索,并获取第一搜索结果。
  8. 如权利要求6或7所述的装置,其特征在于,还包括:
    处理模块,用于将所述第一识别结果与所述第二识别结果进行对比,若所述第一识别结果与所述第二识别结果一致,则将所述第一搜索结果作为最终的搜索结果,以及若所述第一识别结果与所述第二识别结果不一致,则将所述第一识别结果与所述第二识别结果进行拼接,生成最终的识别结果,并将所述识别结果作为所述关键词进行搜索,以获取最终的所述搜索结果。
  9. 如权利要求8所述的装置,其特征在于,还包括:
    显示模块,用于在获取最终的所述搜索结果之后,显示所述搜索结果。
  10. 如权利要求6-9任一项所述的装置,其特征在于,所述判断判断模块,具体用于:
    根据尾点检测算法判断所述静音的类型。
  11. 一种设备,其特征在于,包括:
    一个或者多个处理器;
    存储器;
    一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,执行如权利要求1-5任一项所述的语音识别方法。
  12. 一种非易失性计算机存储介质,其特征在于,所述计算机存储介质存储有一个或者多个程序,当所述一个或者多个程序被一个设备执行时,使得所述设备执行如权利要求1-5任一项所述的语音识别方法。
PCT/CN2015/096596 2015-07-22 2015-12-07 语音识别方法和装置 Ceased WO2017012242A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510435887.1A CN105139849B (zh) 2015-07-22 2015-07-22 语音识别方法和装置
CN201510435887.1 2015-07-22

Publications (1)

Publication Number Publication Date
WO2017012242A1 true WO2017012242A1 (zh) 2017-01-26

Family

ID=54725171

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/096596 Ceased WO2017012242A1 (zh) 2015-07-22 2015-12-07 语音识别方法和装置

Country Status (2)

Country Link
CN (1) CN105139849B (zh)
WO (1) WO2017012242A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111768800A (zh) * 2020-06-23 2020-10-13 中兴通讯股份有限公司 语音信号处理方法、设备及存储介质
EP4156179A1 (de) * 2021-09-23 2023-03-29 Siemens Healthcare GmbH Sprachsteuerung einer medizinischen vorrichtung
US12469492B2 (en) 2021-09-23 2025-11-11 Siemens Healthineers Ag Speech control of a medical apparatus
US12478360B2 (en) 2021-09-23 2025-11-25 Siemens Healthineers Ag Speech control of a medical apparatus

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106648530B (zh) * 2016-11-21 2020-09-08 海信集团有限公司 语音控制方法及终端
CN107644095A (zh) * 2017-09-28 2018-01-30 百度在线网络技术(北京)有限公司 用于搜索信息的方法和装置
WO2019084890A1 (en) * 2017-11-03 2019-05-09 Tencent Technology (Shenzhen) Company Limited Method and system for processing audio communications over a network
CN108962283B (zh) * 2018-01-29 2020-11-06 北京猎户星空科技有限公司 一种发问结束静音时间的确定方法、装置及电子设备
CN108847237A (zh) * 2018-07-27 2018-11-20 重庆柚瓣家科技有限公司 连续语音识别方法及系统
CN110517673B (zh) * 2019-07-18 2023-08-18 平安科技(深圳)有限公司 语音识别方法、装置、计算机设备及存储介质
CN110995943B (zh) * 2019-12-25 2021-05-07 携程计算机技术(上海)有限公司 多用户流式语音识别方法、系统、设备及介质
CN111261161B (zh) * 2020-02-24 2021-12-14 腾讯科技(深圳)有限公司 一种语音识别方法、装置及存储介质
CN114255757B (zh) * 2020-09-22 2025-09-16 阿尔卑斯阿尔派株式会社 语音信息处理装置及语音信息处理方法
CN112466302B (zh) * 2020-11-23 2022-09-23 北京百度网讯科技有限公司 语音交互的方法、装置、电子设备和存储介质
CN112927680B (zh) * 2021-02-10 2022-06-17 中国工商银行股份有限公司 一种基于电话信道的声纹有效语音的识别方法及装置
CN114898755B (zh) * 2022-07-14 2023-01-17 科大讯飞股份有限公司 语音处理方法及相关装置、电子设备、存储介质
CN115910043B (zh) * 2023-01-10 2023-06-30 广州小鹏汽车科技有限公司 语音识别方法、装置及车辆

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07230293A (ja) * 1994-02-17 1995-08-29 Sony Corp 音声認識装置
EP0770986A2 (en) * 1995-10-26 1997-05-02 Dragon Systems Inc. Modified discrete word recognition
JP2003255978A (ja) * 2002-03-05 2003-09-10 Masakazu Suzuki 数式音声認識方法及び装置
CN101042866A (zh) * 2006-03-22 2007-09-26 富士通株式会社 语音识别设备及方法,以及记录有计算机程序的记录介质
CN102231278A (zh) * 2011-06-10 2011-11-02 安徽科大讯飞信息科技股份有限公司 实现语音识别中自动添加标点符号的方法及系统
CN102280106A (zh) * 2010-06-12 2011-12-14 三星电子株式会社 用于移动通信终端的语音网络搜索方法及其装置
CN103871401A (zh) * 2012-12-10 2014-06-18 联想(北京)有限公司 一种语音识别的方法及电子设备

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9384736B2 (en) * 2012-08-21 2016-07-05 Nuance Communications, Inc. Method to provide incremental UI response based on multiple asynchronous evidence about user input
CN102903361A (zh) * 2012-10-15 2013-01-30 Itp创新科技有限公司 一种通话即时翻译系统和方法
CN103035243B (zh) * 2012-12-18 2014-12-24 中国科学院自动化研究所 长语音连续识别及识别结果实时反馈方法和系统
CN104538030A (zh) * 2014-12-11 2015-04-22 科大讯飞股份有限公司 一种可以通过语音控制家电的控制系统与方法

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07230293A (ja) * 1994-02-17 1995-08-29 Sony Corp 音声認識装置
EP0770986A2 (en) * 1995-10-26 1997-05-02 Dragon Systems Inc. Modified discrete word recognition
JP2003255978A (ja) * 2002-03-05 2003-09-10 Masakazu Suzuki 数式音声認識方法及び装置
CN101042866A (zh) * 2006-03-22 2007-09-26 富士通株式会社 语音识别设备及方法,以及记录有计算机程序的记录介质
CN102280106A (zh) * 2010-06-12 2011-12-14 三星电子株式会社 用于移动通信终端的语音网络搜索方法及其装置
CN102231278A (zh) * 2011-06-10 2011-11-02 安徽科大讯飞信息科技股份有限公司 实现语音识别中自动添加标点符号的方法及系统
CN103871401A (zh) * 2012-12-10 2014-06-18 联想(北京)有限公司 一种语音识别的方法及电子设备

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111768800A (zh) * 2020-06-23 2020-10-13 中兴通讯股份有限公司 语音信号处理方法、设备及存储介质
EP4156179A1 (de) * 2021-09-23 2023-03-29 Siemens Healthcare GmbH Sprachsteuerung einer medizinischen vorrichtung
US12469492B2 (en) 2021-09-23 2025-11-11 Siemens Healthineers Ag Speech control of a medical apparatus
US12478360B2 (en) 2021-09-23 2025-11-25 Siemens Healthineers Ag Speech control of a medical apparatus

Also Published As

Publication number Publication date
CN105139849B (zh) 2017-05-10
CN105139849A (zh) 2015-12-09

Similar Documents

Publication Publication Date Title
US12159622B2 (en) Text independent speaker recognition
US10438595B2 (en) Speaker identification and unsupervised speaker adaptation techniques
US9472196B1 (en) Developer voice actions system
CN106663427B (zh) 用于服务语音发音的高速缓存设备
CN108009303B (zh) 基于语音识别的搜索方法、装置、电子设备和存储介质
JP2019091012A (ja) 情報認識方法および装置
CN104144239B (zh) 一种语音辅助通讯方法和装置
CN105139849A (zh) 语音识别方法和装置
CN106415719A (zh) 使用说话者识别的语音信号的稳健端点指示
CN103811006A (zh) 用于语音识别的方法和装置
EP2950307A1 (en) Reducing the need for manual start/end-pointing and trigger phrases
CN110610701B (zh) 语音交互方法、语音交互提示方法、装置和设备
CN103488401A (zh) 一种语音助手激活方法和装置
CN103489444A (zh) 一种语音识别方法和装置
US9916831B2 (en) System and method for handling a spoken user request
CN105469801B (zh) 一种修复输入语音的方法及其装置
US20180367669A1 (en) Input during conversational session
CN108694941A (zh) 用于交互式会话的方法、信息处理装置及产品
CN108073275A (zh) 信息处理方法、信息处理设备及程序产品
WO2015014157A1 (zh) 基于移动终端的歌曲推荐方法与装置
US10950221B2 (en) Keyword confirmation method and apparatus
CN113470649A (zh) 语音交互方法及装置
WO2021136298A1 (zh) 一种语音处理方法、装置、智能设备及存储介质
CN106791226B (zh) 通话故障检测方法及系统
US20190065608A1 (en) Query input received at more than one device

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15898798

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15898798

Country of ref document: EP

Kind code of ref document: A1