WO2018006489A1 - 终端的语音交互方法及装置 - Google Patents
终端的语音交互方法及装置 Download PDFInfo
- Publication number
- WO2018006489A1 WO2018006489A1 PCT/CN2016/098147 CN2016098147W WO2018006489A1 WO 2018006489 A1 WO2018006489 A1 WO 2018006489A1 CN 2016098147 W CN2016098147 W CN 2016098147W WO 2018006489 A1 WO2018006489 A1 WO 2018006489A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- matching
- information
- terminal
- text information
- cloud server
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
- H04N21/42203—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS] sound input device, e.g. microphone
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/27—Server based end-user applications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
- H04N21/42204—User interfaces specially adapted for controlling a client device through a remote control device; Remote control devices therefor
- H04N21/42206—User interfaces specially adapted for controlling a client device through a remote control device; Remote control devices therefor characterized by hardware details
Definitions
- the present invention relates to the field of terminal technologies, and in particular, to a voice interaction method and apparatus for a terminal.
- the intelligent interactive system implementation mode on the existing television system is a customized mode, that is, the TV manufacturer puts forward the requirement and is customized by the third-party identification system.
- Voice and semantic recognition generally adopt a binding method. TV manufacturers can only select a service provider on a TV to complete voice and semantic recognition in voice interaction. This implementation mode is too limited for traditional TV companies. Large, unable to adjust according to demand, and poor flexibility.
- the main purpose of the present invention is to provide a method and device for voice interaction of a terminal, which aims to solve the problem that a TV manufacturer can only select a service provider on a television to complete voice and semantic recognition in voice interaction.
- a TV manufacturer can only select a service provider on a television to complete voice and semantic recognition in voice interaction.
- the limitations are too large to adjust according to demand and the flexibility is poor.
- the present invention provides a method for voice interaction of a terminal, including the steps of:
- the terminal receives the output information returned by the cloud server and outputs the output information.
- the method further includes:
- the terminal performs a matching operation according to the text information and the information pre-stored in the local database of the terminal;
- a response control operation corresponding to the control information is performed.
- the step of performing the matching operation according to the text information and the information pre-stored by the terminal local database comprises:
- the terminal calculates a matching parameter according to the text information and the information of the current page collected in advance;
- the calculating, by the terminal, the matching parameters according to the text information and the information of the current page collected in advance includes:
- the page space collection algorithm of the terminal in the background performs the collection of the text information of the controllable control of the current page of the television
- the terminal After acquiring the text information, the terminal calculates matching parameters for the text information and the text collected by the combined scene control.
- the method further includes:
- the terminal After the current page entry fails, the terminal matches the matching parameter with the global static term. After the global static term is successfully matched, the terminal sets a tag that matches the global static term.
- the matching parameter is matched with the application information, and after matching with the application information, setting a label matching the application information;
- the cloud server searches for output information corresponding to the text information, including:
- the cloud server After receiving the text information, the cloud server identifies the search parameter of the text information according to the semantic parsing engine loaded by itself;
- the method further comprises the steps of:
- the cloud server determines the service type corresponding to the search parameter, and accesses the information provider corresponding to the service type to provide the information service.
- the present invention further provides a voice interaction device for a terminal, including:
- a receiving module configured to receive an audio stream output by the voice input device
- An obtaining module configured to acquire text information corresponding to the audio stream
- a sending module configured to upload the text information to a cloud server built by the terminal corresponding operator, to search for output information corresponding to the text information by using the cloud server, and return to the terminal;
- the receiving module is further configured to receive output information returned by the cloud server;
- An output module for outputting output information returned by the cloud server.
- the method further comprises:
- a matching module configured to perform a matching operation according to the information stored in the terminal database according to the text information
- the obtaining module is further configured to: after the matching operation is successful, obtain control information corresponding to the matching operation;
- a response module configured to execute a response control operation corresponding to the control information
- the sending module is configured to upload the text information to a cloud server of the terminal after the matching operation fails.
- the matching module comprises:
- a calculating unit configured to calculate a matching parameter according to the text information and information of a pre-acquired current page
- a matching unit configured to match the matching parameter with a current page entry, after the current page entry is successfully matched
- a setting unit configured to set a label that matches the current page entry.
- the calculating unit is further configured to: in the process of transmitting the audio stream to the terminal, the page space collection algorithm in the background performs the collection of the text information of the current page controllable control of the television; the calculating unit is further used for
- the matching parameters are calculated for the text information and the text collected by the combined scene control.
- the matching module further includes: a prompting unit,
- the matching unit is further configured to: after the current page entry matching fails, match the matching parameter with a global static term;
- the setting unit is further configured to: after matching with the global static term, set a tag that matches the global static term;
- the matching unit is further configured to: after the matching with the global static term fails, match the matching parameter with the application information;
- the setting unit is further configured to: after matching the application information, set a label that matches the application information;
- the prompting unit is configured to prompt the matching operation operation to fail after the matching with the application information fails.
- the process of the cloud server acquiring the output information includes: after receiving the text information, identifying a search parameter of the text information according to a semantic analysis engine loaded by itself; and searching for corresponding output information according to the search parameter.
- the cloud server determines a service type corresponding to the search parameter, and accesses an information provider corresponding to the service type to provide an information service.
- the terminal of the invention sets up the voice interaction of its own terminal platform, and uses the television server as an interface to independently select the voice recognition service identification service and the semantic analysis engine, and separates the voice recognition from the semantic recognition, is not bound, and is semantically recognized.
- the operation is identified in the terminal's own server, without relying on the third-party service provider to provide services, and can be adjusted according to requirements, and the flexibility is greatly increased.
- FIG. 1 is a schematic flowchart of a first embodiment of a voice interaction method of a terminal according to the present invention
- FIG. 2 is a schematic flowchart of a second embodiment of a voice interaction method of a terminal according to the present invention
- FIG. 3 is a schematic flowchart of a matching operation according to an embodiment of the present invention.
- FIG. 4 is a schematic flowchart of a third embodiment of a voice interaction method of a terminal according to the present invention.
- FIG. 5 is a schematic diagram of functional modules of a first embodiment of a voice interaction device of a terminal according to the present invention.
- FIG. 6 is a schematic diagram of functional modules of a second embodiment of a voice interaction device of a terminal according to the present invention.
- FIG. 7 is a schematic diagram of a refinement function module of an embodiment of the matching module of FIG. 6;
- FIG. 8 is a schematic diagram of logic of a voice interaction service according to an embodiment of the present invention.
- FIG. 9 is a schematic flowchart of voice interaction according to an embodiment of the present invention.
- the main solution of the embodiment of the present invention is: the terminal establishes the voice interaction of its own terminal platform, and uses the TV server as an interface to independently select the voice recognition service identification service and the semantic analysis engine to separate the voice recognition from the semantic recognition.
- the binding and semantic recognition operations are identified in the terminal's own server, without relying on third-party service providers to provide services, which can be adjusted according to requirements, and the flexibility is greatly increased.
- the present invention provides a voice interaction method for a terminal.
- FIG. 1 is a schematic flowchart diagram of a first embodiment of a voice interaction method of a terminal according to the present invention.
- the voice interaction method of the terminal includes:
- Step S10 The terminal receives the audio stream output by the voice input device, and acquires text information corresponding to the audio stream.
- the voice input device is a mobile phone or a remote controller, and the mobile phone can input voice to the terminal by using a WeChat voice or a multi-screen interactive voice module;
- the remote controller is a remote controller capable of supporting a voice input function.
- the terminal is preferably a television set, and may also be a controlled display device.
- the user When the user needs to interact with the television, the user is connected to the television through a mobile phone, and the connection may be a wireless or wired connection. After the connection is established, the user enters the voice through the mobile phone, and the mobile phone converts the recorded voice into an audio stream in real time, transmits it to the television, or transmits the converted audio stream to the television after a voice recording is completed.
- the television acquires text information corresponding to the audio stream.
- the obtaining process includes but is not limited to: 1) the television uploads the audio stream to a third-party voice recognition server, and the third-party voice recognition server identifies the audio stream to obtain text information of the audio stream, and feeds the text information to the television.
- the TV customizes or purchases the voice recognition service, and saves the customized or purchased voice recognition service database on the TV local end.
- the television After receiving the audio stream, the television recognizes the text information of the audio stream through the local database, and completes the text on the TV local end.
- the process of audio streaming text information is merely exemplary, and the present invention is not limited to the scope of the above description.
- Step S20 the terminal uploads the text information to a cloud server constructed by the terminal corresponding to the operator, to search for output information corresponding to the text information by using the cloud server, and return to the terminal;
- the television has its own cloud server loaded with a semantic parsing engine for identifying the semantics of the textual information of the audio stream.
- the text information is uploaded to the cloud server of the terminal.
- the cloud server identifies the search parameter of the text information according to the semantic parsing engine loaded by itself (for semantic recognition, and identifies the semantics of the user from the voice sent by the user through the voice input device).
- the search parameter is keyword information of text information or user demand information, for example, an on-demand service, a song search service, or an e-commerce service.
- the search parameter takes the keyword information as an example, and searches for output information corresponding to the text information according to the keyword information, where the output information may be a resource stored in a local server database, or may be provided by a third-party service provider. of. After searching for the information, return the searched output information to the TV.
- the output information may be e-commerce push information, product advertisement information, or the like.
- Step S30 the terminal receives the information returned by the cloud server and outputs the information.
- the television receives the output information returned by the cloud server and outputs the output, including direct display, or push to other terminals (eg, mobile phones, pads, etc.) connected to the television or played.
- the terminal establishes a voice interaction of its own terminal platform, and uses the TV server as an interface to independently select an access voice recognition service identification service and a semantic resolution engine to separate the voice recognition from the semantic recognition, without binding, and semantic recognition.
- the operation is identified in the server of the terminal itself, without relying on the service provided by the third-party service provider, and can be adjusted according to requirements, and the flexibility is greatly increased.
- FIG. 2 is a schematic flowchart diagram of a second embodiment of a voice interaction method of a terminal according to the present invention. Based on the first embodiment of the voice interaction method of the terminal, after the step S10, the method further includes:
- Step S40 The terminal performs a matching operation according to the text information and information pre-stored in the local database of the terminal;
- Step S50 After the matching operation is successful, obtain control information corresponding to the matching operation;
- step S60 a response control operation corresponding to the control information is performed.
- the matching operation of the television control is performed first.
- the television stores a database of control information, such as control information including volume addition and subtraction, up and down left and right control, play, pause, fast forward or rewind.
- Matching the text information with the information pre-stored by the terminal local database after the matching operation is successful, acquiring control information corresponding to the matching operation; performing a response control operation corresponding to the control information; after the matching operation fails, A process of uploading the text information to a cloud server of the terminal is performed.
- the process of performing the matching operation according to the text information and the information pre-stored by the terminal local database includes:
- Step S41 the terminal calculates a matching parameter according to the text information and the information of the current page collected in advance;
- Step S42 Match the matching parameter with the current page entry, and after the current page term is successfully matched, set a tag that matches the current page entry.
- Step S43 After the current page term matching fails, the matching parameter is matched with the global static term, and after the global static term is successfully matched, the tag matching the global static term is set;
- Step S44 After the matching with the global static term fails, the matching parameter is matched with the application information, and after the matching with the application information is successful, setting a label matching the application information;
- step S45 after the matching with the application information fails, the prompt matching operation fails.
- the page space collection algorithm of the TV in the background will start collecting the text information of the controllable control of the current page of the TV.
- the local fuzzy matching operation of the TV is performed, and the fuzzy matching is performed through the fuzzy matching.
- the algorithm calculates the character matching number, source matching degree, and target matching degree as the matching parameters for the text information and the text collected by the combined scene control.
- the target matching degree is also set differently, according to requirements and performance settings, for example. It can be 0.67 or 1 and so on.
- the matching priority order is also set.
- the current page entry is matched first, and the matching degree requirement is 0.67; after the current page term matching fails, the global static term is matched, and the global static term includes some preset global control. Commands, for example, volume addition and subtraction, up and down left and right control, etc.; global static term matching fails to match playback control terms, such as pause, play, fast forward or rewind; the last match is the application term, and the machine
- the current page matching success label is FUZZY_MATCH
- the global static matching success label is GLOBAL_MATCH
- the local application matching success label is APP_MATCH
- the playback control matching success label is PLAYER_MATCH
- the fuzzy matching failure label is FAIL_MATCH.
- the matching success label is FUZZY_MATCH
- the matching success label is PLAYER_MATCH
- the matching is completed in the playback control entry, and the corresponding playback control instruction is processed
- the local fuzzy matching succeeds. After the corresponding control command is completed, the voice interaction ends.
- the matching operation fails, and the interaction with the cloud service is performed to obtain the information required by the user.
- the matching operation is performed locally in the terminal, and the third party is not required to complete the matching operation and control, thereby effectively improving the local control efficiency of the terminal.
- FIG. 4 is a schematic flowchart diagram of a third embodiment of a voice interaction method of a terminal according to the present invention. The method also includes the steps of:
- Step S70 After identifying the search parameter, the cloud server determines a service type corresponding to the search parameter, and accesses an information provider corresponding to the service type to provide an information service.
- the cloud server determines the service type corresponding to the search parameter, for example, the on-demand service is required.
- the song search business is still e-commerce business.
- the cloud server selects and accesses an information provider providing information service corresponding to the service type according to the service type corresponding to the identified search parameter.
- the service type may be customized according to requirements, and the server of the terminal selects an appropriate information provider to provide the service.
- the step S70 can also be performed before or after other steps, and the order can be adjusted according to actual needs.
- the embodiment is based on the cloud platform of the terminal, and the cloud platform provides an interface, customizes the extended service type, and selects an appropriate information provider to provide information services, avoids the limitation caused by outsourcing by a voice service provider, and improves the flexibility of the funding control. Sex.
- the invention further provides a voice interaction device for a terminal.
- FIG. 5 is a schematic diagram of functional modules of a first embodiment of a voice interaction apparatus of a terminal according to the present invention.
- the device includes: a receiving module 10, an obtaining module 20, a sending module 30, and an output module 40.
- the receiving module 10 is configured to receive an audio stream output by the voice input device.
- the obtaining module 20 is configured to acquire text information corresponding to the audio stream
- the voice input device is a mobile phone or a remote controller, and the mobile phone can input voice to the terminal by using a WeChat voice or a multi-screen interactive voice module;
- the remote controller is a remote controller capable of supporting a voice input function.
- the terminal is preferably a television set, and may also be a controlled display device.
- the user When the user needs to interact with the television, the user is connected to the television through a mobile phone, and the connection may be a wireless or wired connection. After the connection is established, the user enters the voice through the mobile phone, and the mobile phone converts the recorded voice into an audio stream in real time, transmits it to the television, or transmits the converted audio stream to the television after a voice recording is completed.
- the receiving module 10 receives the audio stream output by the voice input device, and the acquiring module 20 acquires the text information corresponding to the audio stream.
- the acquisition module 20 acquisition process includes, but is not limited to: 1) uploading the audio stream to a third-party voice recognition server, and the third-party voice recognition server will identify the audio stream to obtain text information of the audio stream, and feedback the text information.
- the TV is customized or purchased the voice recognition service, and the customized or purchased voice recognition service database is saved on the TV local end.
- the text information of the audio stream is identified through the local database, and is completed at the local end.
- the process of audio streaming text information is merely exemplary, and the present invention is not limited to the scope of the above description.
- the sending module 30 is configured to upload the text information to a cloud server that is configured by the terminal corresponding to the operator, to search for output information corresponding to the text information by using the cloud server, and return the information;
- the television has its own cloud server loaded with a semantic parsing engine for identifying the semantics of the textual information of the audio stream.
- the text information is uploaded to the cloud server of the terminal.
- the cloud server identifies the search parameter of the text information according to the semantic parsing engine loaded by itself (for semantic recognition, and identifies the semantics of the user from the voice sent by the user through the voice input device).
- the search parameter is keyword information of text information or user demand information, for example, an on-demand service, a song search service, or an e-commerce service.
- the search parameter takes the keyword information as an example, and searches for output information corresponding to the text information according to the keyword information, where the output information may be a resource stored in a local server database, or may be provided by a third-party service provider. of. After searching for the output information, return the searched output information to the TV.
- the information may be e-commerce push information, product advertisement information, and the like.
- the receiving module 10 is further configured to receive output information returned by the cloud server;
- the output module 40 is configured to output output information returned by the cloud server.
- the receiving module 10 receives the information returned by the cloud server and outputs it through the output module 40.
- the output manner includes direct display or push to other terminals (eg, mobile phones, pads, etc.) connected to the output module 40 or played.
- the terminal establishes a voice interaction of its own terminal platform, and uses the TV server as an interface to independently select an access voice recognition service identification service and a semantic resolution engine to separate the voice recognition from the semantic recognition, without binding, and semantic recognition.
- the operation is identified in the server of the terminal itself, without relying on the service provided by the third-party service provider, and can be adjusted according to requirements, and the flexibility is greatly increased.
- FIG. 6 is a schematic diagram of functional modules of a second embodiment of a voice interaction apparatus of a terminal according to the present invention. Also included: a matching module 50 and a response module 60,
- the matching module 50 is configured to perform a matching operation according to the text information and information pre-stored in the local database of the terminal;
- the obtaining module 20 is further configured to: after the matching operation is successful, obtain control information corresponding to the matching operation;
- the response module 60 is configured to perform a response control operation corresponding to the control information.
- the matching operation of the television control is performed first.
- a database in which control information is stored in advance such as control information including volume addition and subtraction, up and down left and right control, play, pause, fast forward or rewind.
- Matching the text information with the information pre-stored by the terminal local database after the matching operation is successful, acquiring control information corresponding to the matching operation; performing a response control operation corresponding to the control information; after the matching operation fails, A process of uploading the text information to a cloud server of the terminal is performed.
- the matching module 50 includes:
- the calculating unit 51 is configured to calculate a matching parameter according to the text information and the information of the current page collected in advance;
- the matching unit 52 is configured to match the matching parameter with the current page entry, after the current page entry is successfully matched;
- the setting unit 53 is configured to set a label that matches the current page entry.
- the matching unit 52 is further configured to: after the current page entry matching fails, match the matching parameter with a global static term;
- the setting unit 53 is further configured to: after matching with the global static term, set a tag that matches the global static term;
- the matching unit 52 is further configured to: after the matching with the global static term fails, match the matching parameter with the application information;
- the setting unit 53 is further configured to: after matching the application information, set a label that matches the application information;
- the prompting unit 54 is configured to prompt the matching operation operation to fail after the matching with the application information fails.
- the page space collection algorithm of the computing unit 51 in the background starts to collect the text information of the current page steerable control of the television.
- the matching unit 52 performs the television localization.
- the fuzzy matching operation the calculating unit 51 calculates the character matching number, the source matching degree and the target matching degree as the matching parameters by using the fuzzy matching algorithm for the text information and the text collected by the combined scene control, and setting the target matching degree in different scenarios. Also, depending on the requirements and performance settings, for example, it can be 0.67 or 1 or the like. Similarly, the matching priority order is also set.
- the current page entry is matched first, and the matching degree requirement is 0.67; after the current page term matching fails, the global static term is matched, and the global static term includes some preset global control. Commands, for example, volume addition and subtraction, up and down left and right control, etc.; global static term matching fails to match playback control terms, such as pause, play, fast forward or rewind; the last match is the application term, and the machine
- the current page matching success label is FUZZY_MATCH
- the global static matching success label is GLOBAL_MATCH
- the local application matching success label is APP_MATCH
- the playback control matching success label is PLAYER_MATCH
- the fuzzy matching failure label is FAIL_MATCH.
- the matching success label is FUZZY_MATCH
- the matching success label is PLAYER_MATCH
- the matching is completed in the playback control entry, and the corresponding playback control instruction is processed
- the local fuzzy matching succeeds. After the corresponding control command is completed, the voice interaction ends.
- the prompting unit 54 prompts that the matching operation fails, and transfers to the interaction with the cloud service to obtain the information required by the user.
- the matching operation is performed locally in the terminal, and the third party is not required to complete the matching operation and control, thereby effectively improving the local control efficiency of the terminal.
- the cloud server determines the service type corresponding to the search parameter, and accesses the information provider corresponding to the service type to provide the information service.
- the cloud server determines the service type corresponding to the search parameter, for example, the on-demand service is required.
- the song search business is still e-commerce business.
- the cloud server selects and accesses an information provider providing information service corresponding to the service type according to the service type corresponding to the identified search parameter.
- the service type may be customized according to requirements, and the server of the terminal selects an appropriate information provider to provide the service.
- the embodiment is based on the cloud platform of the terminal, and the cloud platform provides an interface, customizes the extended service type, and selects an appropriate information provider to provide information services, avoids the limitation caused by outsourcing by a voice service provider, and improves the flexibility of the funding control. Sex.
- the service logic diagram of the voice interaction includes:
- the system includes a plurality of parts, including: a voice input module, a local fuzzy matching module, a local control module, a service display module, and a cloud service module;
- Voice input is a voice input device.
- the voice input devices supported by this system include a mobile phone and a remote control.
- the mobile phone input device can input voice via WeChat voice or multi-screen interactive voice module; the remote control supports all remote controllers that support voice input.
- the local fuzzy matching module is the key to implement local control, including the collection of local terms and the fuzzy matching algorithm of terms. After the user's voice input is converted into a voice text, the fuzzy matching algorithm is firstly sent to determine whether the current instruction of the user matches the local term, and if the matching succeeds, the matching type and the matching ID are returned. When doing local fuzzy matching, we set the matching priority of the local scene. First, match the current page control entry. If the matching is unsuccessful, it will continue to match the preset static entry. If the matching is unsuccessful, the matching control entry will continue to match. If the success is successful, the local application terms will continue to be matched. If the matching is unsuccessful, the cloud platform will be submitted for semantic understanding.
- the local control module is the module that performs the local control function. According to the result of the fuzzy matching, the control corresponding to the matching result is found, and the control operation is completed.
- the local control module includes a lookup algorithm and control instructions.
- the business presentation module refers to the presentation of the results fed back by the cloud platform in addition to local control. Such as movie list, song list, product list, etc.;
- the cloud platform module includes all server-side processing.
- the cloud platform includes a local server and a third-party server.
- the local server is responsible for interfacing with the terminal service and connecting with the third-party server.
- the third-party server includes a voice recognition server, a semantic understanding service, and a third-party content provider.
- Step S100 The user inputs a voice command, and the collection algorithm collects the term information of the controllable control of the current page of the system.
- the text information of the voice recognition is transmitted to the local fuzzy matching algorithm for matching, and if the matching succeeds, the local control module is executed, and the control function of the response is performed to complete a voice interaction experience; if the matching is unsuccessful, the text information of the voice recognition is transmitted to the cloud.
- the semantic understanding server feeds back the result of the semantic understanding to the local server; the local server goes to the service display module according to the keyword of the semantic understanding feedback to the resource library search content; the service display module performs the content fed back by the local server at the terminal Reasonable display. Thereby completing a voice interaction experience.
- the system built a set of standard frameworks for speech recognition, semantic understanding and business content access on the TV platform. As a traditional TV manufacturer, you can choose partner access independently. We can choose the voice recognition service engine independently, or we can independently plan the terminal service access type.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Telephonic Communication Services (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本发明公开了一种终端的语音交互方法,包括步骤:终端接收语音输入设备输出的音频流,获取所述音频流对应的文本信息;所述终端将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;所述终端接收所述云服务器返回的输出信息并输出。本发明还公开了一种终端的语音交互装置。本发明语义识别的操作在终端自身的服务器中识别,无需依赖第三方服务商提供服务,可以根据需求进行调整,灵活性大大增加。
Description
技术领域
本发明涉及终端技术领域,尤其涉及终端的语音交互方法及装置。
背景技术
随着科学技术的不断发展,越来越多的智能终端进入人们的日常生活和工作当中。例如,以智能电视为例,用户对智能电视的智能化要求越来越高,用户期望通过语音的方式与智能电视交流,达到想要的目的(电视的控制、节目推送或信息推送等)。然而,目前智能电视在智能语音助手这方面还处于初级阶段,需要依赖语音识别技术和语义识别技术。在现有的电视系统上的智能交互系统实现模式都是定制模式,即电视厂商提出需求,由第三方的识别系统定制实现。语音与语义识别一般采取绑定的方式,电视厂商只能在一台电视上选择一个服务提供商来完成语音交互中的语音和语义识别,这种实现模式对于传统的电视企业来说局限性太大,无法根据需求进行调整,灵活性差。
上述内容仅用于辅助理解本发明的技术方案,并不代表承认上述内容是现有技术。
发明内容
本发明的主要目的在于提供一种终端的语音交互方法及装置,旨在解决目前电视厂商只能在一台电视上选择一个服务提供商来完成语音交互中的语音和语义识别,这种实现模式对于传统的电视企业来说局限性太大,无法根据需求进行调整,灵活性差的问题。
为实现上述目的,本发明提供的一种终端的语音交互方法,包括步骤:
终端接收语音输入设备输出的音频流,获取所述音频流对应的文本信息;
所述终端将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;
所述终端接收所述云服务器返回的输出信息并输出。
优选地,所述获取所述音频流对应的文本信息的步骤之后,还包括:
终端按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作;
在所述匹配操作成功后,获取匹配操作对应的控制信息;
执行与所述控制信息对应的响应控制操作。
优选地,所述按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作的步骤包括:
终端根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;
将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后,设置与所述当前页面词条匹配的标签。
优选地,所述终端根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数包括:
在传输音频流至终端的过程中,所述终端在后台的页面空间收集算法会进行电视当前页面可操控控件文本信息的收集;
所述终端获取到文本信息后,对文本信息和结合场景控件采集的文本计算匹配参数。
优选地,所述将所述匹配参数与当前页面词条匹配的步骤之后,还包括:
终端在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配,在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;
在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配,在与应用信息匹配成功后,设置与应用信息匹配的标签;
在与应用信息匹配失败后,提示匹配操作操作失败。
优选地,所述云服务器搜索与所述文本信息对应的输出信息包括:
所述云服务器在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数;
根据所述搜索参数搜索对应的输出信息。
优选地,所述方法还包括步骤:
在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
此外,为实现上述目的,本发明还提供一种终端的语音交互装置,包括:
接收模块,用于接收语音输入设备输出的音频流;
获取模块,用于获取所述音频流对应的文本信息;
发送模块,用于将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;
所述接收模块,还用于接收所述云服务器返回的输出信息;
输出模块,用于输出云服务器返回的输出信息。
优选地,还包括:
匹配模块,用于按照所述文本信息与终端数据库存储的信息进行匹配操作;
所述获取模块,还用于在所述匹配操作成功后,获取匹配操作对应的控制信息;
响应模块,用于执行与所述控制信息对应的响应控制操作;
所述发送模块,用于在匹配操作失败后,将所述文本信息上传至所述终端的云服务器。
优选地,所述匹配模块包括:
计算单元,用于根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;
匹配单元,用于将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后;
设置单元,用于设置与所述当前页面词条匹配的标签。
优选地,所述计算单元,还用于在传输音频流至终端的过程中,在后台的页面空间收集算法会进行电视当前页面可操控控件文本信息的收集;计算单元还用于
获取到文本信息后,对文本信息和结合场景控件采集的文本计算匹配参数。
优选地,所述匹配模块还包括:提示单元,
所述匹配单元,还用于在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配;
所述设置单元,还用于在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;
所述匹配单元,还用于在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配;
所述设置单元,还用于在与应用信息匹配成功后,设置与应用信息匹配的标签;
所述提示单元,用于在与应用信息匹配失败后,提示匹配操作操作失败。
优选地,所述云端服务器获取输出信息的过程包括:在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数;根据所述搜索参数搜索对应的输出信息。
优选地,在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
本发明终端搭建自己的终端平台的语音交互,通过电视服务器作为接口,自主选择接入的语音识别服务识别服务和语义解析引擎,将语音识别与语义识别分开,不绑定起来,且语义识别的操作在终端自身的服务器中识别,无需依赖第三方服务商提供服务,可以根据需求进行调整,灵活性大大增加。
附图说明
图1为本发明终端的语音交互方法的第一实施例的流程示意图;
图2为本发明终端的语音交互方法的第二实施例的流程示意图;
图3为本发明一实施例中匹配操作的流程示意图;
图4为本发明终端的语音交互方法的第三实施例的流程示意图;
图5为本发明终端的语音交互装置的第一实施例的功能模块示意图;
图6为本发明终端的语音交互装置的第二实施例的功能模块示意图;
图7为图6中匹配模块一实施例的细化功能模块示意图;
图8为本发明一实施例中语音交互业务逻辑示意图;
图9为本发明一实施例中语音交互的流程示意图。
本发明目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
具体实施方式
应当理解,此处所描述的具体实施例仅仅用以解释本发明,并不用于限定本发明。
本发明实施例的主要解决方案是:终端搭建自己的终端平台的语音交互,通过电视服务器作为接口,自主选择接入的语音识别服务识别服务和语义解析引擎,将语音识别与语义识别分开,不绑定起来,且语义识别的操作在终端自身的服务器中识别,无需依赖第三方服务商提供服务,可以根据需求进行调整,灵活性大大增加。
目前存在电视厂商只能在一台电视上选择一个服务提供商来完成语音交互中的语音和语义识别,这种实现模式对于传统的电视企业来说局限性太大,无法根据需求进行调整,灵活性差的问题
基于上述问题,本发明提供一种终端的语音交互方法。
参照图1,图1为本发明终端的语音交互方法的第一实施例的流程示意图。
在一实施例中,所述终端的语音交互方法包括:
步骤S10,终端接收语音输入设备输出的音频流,获取所述音频流对应的文本信息;
在本实施例中,所述语音输入设备为手机或遥控器等,手机可借助微信语音或多屏互动语音模块向终端输入语音;遥控器则为能够支持语音输入功能的遥控器。所述终端优选为电视机,也还可以是被控显示设备。
用户在需要与电视交互时,通过手机与电视连接,所述连接可以是无线或有线的连接。在建立连接后,用户通过手机录入语音,同时手机会将录入的语音实时转化为音频流,传输给电视,或者在一段语音录入结束后,将转换的音频流传输给电视。电视获取所述音频流对应的文本信息。获取过程包括但不限于:1)电视将所述音频流上传到第三方的语音识别服务器,第三方语音识别服务器将对音频流进行识别得到音频流的文本信息,将所述文本信息反馈至电视;2)电视定制或购买了语音识别服务,在电视本端保存定制或者购买的语音识别服务数据库,电视接收到音频流后,通过本地的数据库识别出音频流的文本信息,在电视本端完成音频流转文本信息的过程。以上获取所述音频流对应的文本信息的方式仅仅为示例性说明,而不代表本发明仅仅局限于上述记载的范围。
步骤S20,所述终端将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;
电视有自身的云服务器,所述云服务器上加载有语义解析引擎,用以识别音频流的文本信息的语义。在获取到文本信息后,将所述文本信息上传至所述终端的云服务器,例如,在电视为A服务商的,则将文本信息上传至A服务商的云服务器。所述云服务器在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数(进行语义识别,从用户通过语音输入设备发出的语音中识别用户的语义即需求),所述搜索参数为文本信息的关键字信息或者用户的需求信息,例如,是点播服务、歌曲搜索服务或电商服务等。所述搜索参数以关键字信息为例,根据所述关键字信息搜索与所述文本信息对应的输出信息,所述输出信息可以是本地服务器数据库存储的资源,也可以是通过第三方服务商提供的。在搜索到信息后,将所搜索到的输出信息返回至电视。所述输出信息可以是电商推送信息、产品广告信息等。
步骤S30,所述终端接收所述云服务器返回的信息并输出。
电视接收所述云服务器返回的输出信息并输出,所述输出的方式包括直接显示,或者推送至其他与电视连接的终端(例如,手机、pad等)或者播放。本实施例终端搭建自己的终端平台的语音交互,通过电视服务器作为接口,自主选择接入的语音识别服务识别服务和语义解析引擎,将语音识别与语义识别分开,不绑定起来,且语义识别的操作在终端自身的服务器中识别,无需依赖第三方服务商提供服务,可以根据需求进行调整,灵活性大大增加。
参照图2,图2为本发明终端的语音交互方法的第二实施例的流程示意图。基于上述终端的语音交互方法的第一实施例,所述步骤S10之后,还包括:
步骤S40,终端按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作;
步骤S50,在所述匹配操作成功后,获取匹配操作对应的控制信息;
步骤S60,执行与所述控制信息对应的响应控制操作。
在本实施例中,在获取到文本信息后,先进行电视控制的匹配操作。电视存储了控制信息的数据库,例如包括音量加减,上下左右控制,播放、暂停、快进或快退等控制信息。将所述文本信息与终端本地数据库预先存储的信息进行匹配操作在所述匹配操作成功后,获取匹配操作对应的控制信息;执行与所述控制信息对应的响应控制操作;在匹配操作失败后,执行将所述文本信息上传至所述终端的云服务器的过程。
具体的,参考图3,所述按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作的过程包括:
步骤S41,终端根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;
步骤S42,将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后,设置与所述当前页面词条匹配的标签。
步骤S43,在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配,在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;
步骤S44,在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配,在与应用信息匹配成功后,设置与应用信息匹配的标签;
步骤S45,在与应用信息匹配失败后,提示匹配操作操作失败。
在手机传输音频流至电视的过程中,电视在后台的页面空间收集算法会开始进行电视当前页面可操控控件文本信息的收集,电视获取到文本信息后,进行电视本地模糊匹配操作,通过模糊匹配算法对文本信息和结合场景控件采集的文本计算得到字符匹配数、源匹配度和目标匹配度等数据作为匹配参数,不同的场景下,目标匹配度的设置也不同,根据需求和性能设置,例如可以是0.67或1等。同样也设置匹配的优先顺序,首先优先匹配的是当前页面词条,匹配度要求达到0.67;当前页面词条匹配失败后,就匹配全局静态词条,全局静态词条包括预设的一些全局控制命令,例如,音量加减,上下左右控制等;全局静态词条匹配失效后就匹配播放控制词条,比如暂停、播放、快进或快退等;最后匹配的是应用词条,及本机安装的所有应用名的匹配;除当前词条匹配的目标匹配度为0.67外,其余场景的匹配度都为1,即,必须全匹配才算匹配成功。不同场景的标签定义如下:当前页面匹配成功标签为FUZZY_MATCH,全局静态匹配成功标签为GLOBAL_MATCH,本地应用匹配成功标签为APP_MATCH,播放控制匹配成功标签为PLAYER_MATCH,模糊匹配失败标签为FAIL_MATCH。例如匹配成功标签为FUZZY_MATCH时,代表在当前页面词条中完成匹配,处理当前页面的控制指令;当匹配成功标签为PLAYER_MATCH时,代表在播放控制词条中完成匹配,处理对应的播放控制指令;本地模糊匹配成功,完成对应的控制指令后本次语音交互结束;在匹配失败后,提示匹配操作操作失败,转入与云服务的交互操作,获取用户需求的信息。本实施例通过将匹配操作放在终端本地执行,无需连接第三方去完成匹配操作和控制,有效的提高了终端本地控制效率。
参照图4,图4为本发明终端的语音交互方法的第三实施例的流程示意图。所述方法还包括步骤:
步骤S70,在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
在本实施例中,在识别到用户通过语音输入设备输出的音频流的搜索参数后,即,识别出用户的需求后,由云服务器确定搜索参数对应的业务类型,例如,是需要点播业务、歌曲搜索业务还是电商业务等。云服务器根据识别的搜索参数对应的业务类型,选择接入与所述业务类型对应的信息商提供信息服务。在本发明其他实施例中也还可以是根据需求自定义业务类型,通过终端的服务器这个接口选择合适的信息提供商提供服务。图中仅仅为一个实施例的执行顺序,在本发明其他实施例中,所述步骤S70也可以执行在其他的步骤之前或者之后,可以根据实际需求进行顺序的调整。本实施例基于终端的云平台,有云平台提供接口,自定义扩展业务类型并选择合适的信息商提供信息服务,避免由一个语音服务厂商外包完成所导致的局限性,提高了资助控制的灵活性。
本发明进一步提供一种终端的语音交互装置。
参照图5,图5为本发明终端的语音交互装置的第一实施例的功能模块示意图。
在一实施例中,所述装置包括:接收模块10、获取模块20、发送模块30及输出模块40。
所述接收模块10,用于接收语音输入设备输出的音频流;
所述获取模块20,用于获取所述音频流对应的文本信息;
在本实施例中,所述语音输入设备为手机或遥控器等,手机可借助微信语音或多屏互动语音模块向终端输入语音;遥控器则为能够支持语音输入功能的遥控器。所述终端优选为电视机,也还可以是被控显示设备。
用户在需要与电视交互时,通过手机与电视连接,所述连接可以是无线或有线的连接。在建立连接后,用户通过手机录入语音,同时手机会将录入的语音实时转化为音频流,传输给电视,或者在一段语音录入结束后,将转换的音频流传输给电视。接收模块10接收语音输入设备输出的音频流,获取模块20获取所述音频流对应的文本信息。获取模块20获取过程包括但不限于:1)将所述音频流上传到第三方的语音识别服务器,第三方语音识别服务器将对音频流进行识别得到音频流的文本信息,将所述文本信息反馈至电视;2)电视定制或购买了语音识别服务,在电视本端保存定制或者购买的语音识别服务数据库,接收到音频流后,通过本地的数据库识别出音频流的文本信息,在本端完成音频流转文本信息的过程。以上获取所述音频流对应的文本信息的方式仅仅为示例性说明,而不代表本发明仅仅局限于上述记载的范围。
所述发送模块30,用于将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回;
电视有自身的云服务器,所述云服务器上加载有语义解析引擎,用以识别音频流的文本信息的语义。在获取到文本信息后,将所述文本信息上传至所述终端的云服务器,例如,在电视为A服务商的,则将文本信息上传至A服务商的云服务器。所述云服务器在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数(进行语义识别,从用户通过语音输入设备发出的语音中识别用户的语义即需求),所述搜索参数为文本信息的关键字信息或者用户的需求信息,例如,是点播服务、歌曲搜索服务或电商服务等。所述搜索参数以关键字信息为例,根据所述关键字信息搜索与所述文本信息对应的输出信息,所述输出信息可以是本地服务器数据库存储的资源,也可以是通过第三方服务商提供的。在搜索到输出信息后,将所搜索到的输出信息返回至电视。所述信息可以是电商推送信息、产品广告信息等。
所述接收模块10,还用于接收所述云服务器返回的输出信息;
所述输出模块40,用于输出所述云服务器返回的输出信息。
接收模块10接收所述云服务器返回的信息并通过输出模块40输出,所述输出的方式包括直接显示,或者推送至其他与输出模块40连接的终端(例如,手机、pad等)或者播放。本实施例终端搭建自己的终端平台的语音交互,通过电视服务器作为接口,自主选择接入的语音识别服务识别服务和语义解析引擎,将语音识别与语义识别分开,不绑定起来,且语义识别的操作在终端自身的服务器中识别,无需依赖第三方服务商提供服务,可以根据需求进行调整,灵活性大大增加。
参照图6,图6为本发明终端的语音交互装置的第二实施例的功能模块示意图。还包括:匹配模块50和响应模块60,
所述匹配模块50,用于按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作;
所述获取模块20,还用于在所述匹配操作成功后,获取匹配操作对应的控制信息;
所述响应模块60,用于与所述控制信息对应的响应控制操作。
在本实施例中,在获取到文本信息后,先进行电视控制的匹配操作。预先存储了控制信息的数据库,例如包括音量加减,上下左右控制,播放、暂停、快进或快退等控制信息。将所述文本信息与终端本地数据库预先存储的信息进行匹配操作在所述匹配操作成功后,获取匹配操作对应的控制信息;执行与所述控制信息对应的响应控制操作;在匹配操作失败后,执行将所述文本信息上传至所述终端的云服务器的过程。
参考图7,所述匹配模块50包括:
计算单元51,用于根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;
匹配单元52,用于将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后;
设置单元53,用于设置与所述当前页面词条匹配的标签。
所述匹配单元52,还用于在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配;
所述设置单元53,还用于在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;
所述匹配单元52,还用于在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配;
所述设置单元53,还用于在与应用信息匹配成功后,设置与应用信息匹配的标签;
所述提示单元54,用于在与应用信息匹配失败后,提示匹配操作操作失败。
在手机传输音频流至电视的过程中,计算单元51在后台的页面空间收集算法会开始进行电视当前页面可操控控件文本信息的收集,获取模块20获取到文本信息后,匹配单元52进行电视本地模糊匹配操作,计算单元51通过模糊匹配算法对文本信息和结合场景控件采集的文本计算得到字符匹配数、源匹配度和目标匹配度等数据作为匹配参数,不同的场景下,目标匹配度的设置也不同,根据需求和性能设置,例如可以是0.67或1等。同样也设置匹配的优先顺序,首先优先匹配的是当前页面词条,匹配度要求达到0.67;当前页面词条匹配失败后,就匹配全局静态词条,全局静态词条包括预设的一些全局控制命令,例如,音量加减,上下左右控制等;全局静态词条匹配失效后就匹配播放控制词条,比如暂停、播放、快进或快退等;最后匹配的是应用词条,及本机安装的所有应用名的匹配;除当前词条匹配的目标匹配度为0.67外,其余场景的匹配度都为1,即,必须全匹配才算匹配成功。不同场景的标签定义如下:当前页面匹配成功标签为FUZZY_MATCH,全局静态匹配成功标签为GLOBAL_MATCH,本地应用匹配成功标签为APP_MATCH,播放控制匹配成功标签为PLAYER_MATCH,模糊匹配失败标签为FAIL_MATCH。例如匹配成功标签为FUZZY_MATCH时,代表在当前页面词条中完成匹配,处理当前页面的控制指令;当匹配成功标签为PLAYER_MATCH时,代表在播放控制词条中完成匹配,处理对应的播放控制指令;本地模糊匹配成功,完成对应的控制指令后本次语音交互结束;在匹配失败后,提示单元54提示匹配操作操作失败,转入与云服务的交互操作,获取用户需求的信息。本实施例通过将匹配操作放在终端本地执行,无需连接第三方去完成匹配操作和控制,有效的提高了终端本地控制效率。
进一步地,在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
在本实施例中,在识别到用户通过语音输入设备输出的音频流的搜索参数后,即,识别出用户的需求后,由云服务器确定搜索参数对应的业务类型,例如,是需要点播业务、歌曲搜索业务还是电商业务等。云服务器根据识别的搜索参数对应的业务类型,选择接入与所述业务类型对应的信息商提供信息服务。在本发明其他实施例中也还可以是根据需求自定义业务类型,通过终端的服务器这个接口选择合适的信息提供商提供服务。本实施例基于终端的云平台,有云平台提供接口,自定义扩展业务类型并选择合适的信息商提供信息服务,避免由一个语音服务厂商外包完成所导致的局限性,提高了资助控制的灵活性。
为了更好的描述本发明的实现过程,参考图8,语音交互的业务逻辑图,包括:
本系统(包括上述运行过程的系统,也为云平台)包括几大部分,包括:语音输入模块、本地模糊匹配模块、本地控制模块、业务展示模块、云服务模块;
语音输入即为语音输入设备,本系统支持的语音输入设备有手机和遥控器。手机输入设备可借助微信语音或多屏互动语音模块输入语音;遥控器则支持所有支持语音输入功能遥控器。
本地模糊匹配模块是实现本地控制的关键,包括本地词条的收集和词条的模糊匹配算法。用户的语音输入转换成语音文本后首先给到模糊匹配算法,判断用户当前指令是否匹配到本地词条,如匹配成功返回匹配类型及匹配ID。在做本地模糊匹配时我们设置了本地场景的匹配优先次序,首先匹配当前页面控件词条,匹配不成功则继续匹配预设的静态词条,匹配不成功则继续匹配播放控制词条,匹配不成功则继续匹配本地应用词条,匹配不成功则提交云平台进行语义理解;
本地控制模块则是完成本地控制功能的模块。依据模糊匹配的结果,找到匹配结果对应的控件,完成控制操作。本地控制模块包括查找算法和控制指令。
业务展示模块则指除本地控制外由云平台反馈的结果的展示。例如影视列表、歌曲列表、商品列表等;
云平台模块则包括所有服务器端的处理。在本系统中云平台包括本地服务器和第三方服务器。本地服务器负责和终端业务对接以及和第三方服务器对接,第三方服务器则包括语音识别服务器、语义理解服务和第三方的内容提供商。
本系统的执行流程图如图9所示,现结合图9将整个系统的操作流程详细描述如下:
步骤S100:用户输入语音命令,同时收集算法收集系统当前页面可控控件的词条信息。将语音识别的文本信息传给本地模糊匹配算法进行匹配,匹配成功则进入本地控制模块,执行响应的控制功能,完成一次语音交互体验;匹配不成功则将语音识别的文本信息传给云端的语义理解服务器,语义理解服务器将语义理解的结果反馈给本地服务器;本地服务器根据语义理解反馈的关键字去资源库搜素对应内容交给业务展示模块;业务展示模块将本地服务器反馈的内容在终端进行合理的展示。从而完成一次语音交互体验。
该系统在电视平台搭建了一套语音识别、语义理解和业务内容接入标准框架。作为传统电视厂商,可以自主选择合作方接入,我们可以自主选择语音识别服务引擎,也可以自主规划终端业务接入类型。
以上仅为本发明的优选实施例,并非因此限制本发明的专利范围,凡是利用本发明说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本发明的专利保护范围内。
Claims (14)
- 一种终端的语音交互方法,其特征在于,包括步骤:终端接收语音输入设备输出的音频流,获取所述音频流对应的文本信息;所述终端将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;所述终端接收所述云服务器返回的输出信息并输出。
- 如权利要求1所述的终端的语音交互方法,其特征在于,所述获取所述音频流对应的文本信息的步骤之后,还包括:终端按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作;在所述匹配操作成功后,获取匹配操作对应的控制信息;执行与所述控制信息对应的响应控制操作。
- 如权利要求2所述的终端的语音交互方法,其特征在于,所述按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作的步骤包括:终端根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后,设置与所述当前页面词条匹配的标签。
- 如权利要求3所述的终端的语音交互方法,其特征在于,所述终端根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数包括:在传输音频流至终端的过程中,所述终端在后台的页面空间收集算法会进行电视当前页面可操控控件文本信息的收集;所述终端获取到文本信息后,对文本信息和结合场景控件采集的文本计算匹配参数。
- 如权利要求3所述的终端的语音交互方法,其特征在于,所述将所述匹配参数与当前页面词条匹配的步骤之后,还包括:终端在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配,在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配,在与应用信息匹配成功后,设置与应用信息匹配的标签;在与应用信息匹配失败后,提示匹配操作操作失败。
- 如权利要求1所述的终端的语音交互方法,其特征在于,所述云服务器搜索与所述文本信息对应的输出信息包括:所述云服务器在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数;根据所述搜索参数搜索对应的输出信息。
- 如权利要求6所述的终端的语音交互方法,其特征在于,所述方法还包括步骤:在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
- 一种终端的语音交互装置,其特征在于,包括:接收模块,用于接收语音输入设备输出的音频流;获取模块,用于获取所述音频流对应的文本信息;发送模块,用于将所述文本信息上传至所述终端对应运营商构建的云服务器,以通过所述云服务器搜索与所述文本信息对应的输出信息并返回至所述终端;所述接收模块,还用于接收所述云服务器返回的输出信息;输出模块,用于输出云服务器返回的输出信息。
- 如权利要求8所述的终端的语音交互装置,其特征在于,还包括:匹配模块,用于按照所述文本信息与终端本地数据库预先存储的信息进行匹配操作;所述获取模块,还用于在所述匹配操作成功后,获取匹配操作对应的控制信息;响应模块,用于执行与所述控制信息对应的响应控制操作;所述发送模块,用于在匹配操作失败后,将所述文本信息上传至所述终端的云服务器。
- 如权利要求9所述的终端的语音交互装置,其特征在于,所述匹配模块包括:计算单元,用于根据所述文本信息以及预先采集的当前页面的信息计算出匹配参数;匹配单元,用于将所述匹配参数与当前页面词条匹配,在当前页面词条匹配成功后;设置单元,用于设置与所述当前页面词条匹配的标签。
- 如权利要求10所述的终端的语音交互装置,其特征在于,所述计算单元,还用于在传输音频流至终端的过程中,在后台的页面空间收集算法会进行电视当前页面可操控控件文本信息的收集;计算单元还用于获取到文本信息后,对文本信息和结合场景控件采集的文本计算匹配参数。
- 如权利要求10所述的终端的语音交互装置,其特征在于,所述匹配模块还包括:提示单元,所述匹配单元,还用于在当前页面词条匹配失败后,将所述匹配参数与全局静态词条匹配;所述设置单元,还用于在与全局静态词条匹配成功后,设置与全局静态词条匹配的标签;所述匹配单元,还用于在与全局静态词条匹配失败后,将所述匹配参数与应用信息匹配;所述设置单元,还用于在与应用信息匹配成功后,设置与应用信息匹配的标签;所述提示单元,用于在与应用信息匹配失败后,提示匹配操作操作失败。
- 如权利要求8所述的终端的语音交互装置,其特征在于,所述云端服务器获取输出信息的过程包括:在接收到文本信息后,根据自身加载的语义解析引擎识别出所述文本信息的搜索参数;根据所述搜索参数搜索对应的输出信息。
- 如权利要求13所述的终端的语音交互装置,其特征在于,在识别出搜索参数后,由云服务器确定搜索参数对应的业务类型,接入与所述业务类型对应的信息商提供信息服务。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610529267.9A CN106101789B (zh) | 2016-07-06 | 2016-07-06 | 终端的语音交互方法及装置 |
| CN201610529267.9 | 2016-07-06 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018006489A1 true WO2018006489A1 (zh) | 2018-01-11 |
Family
ID=57213435
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/098147 Ceased WO2018006489A1 (zh) | 2016-07-06 | 2016-09-06 | 终端的语音交互方法及装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN106101789B (zh) |
| WO (1) | WO2018006489A1 (zh) |
Cited By (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109584870A (zh) * | 2018-12-04 | 2019-04-05 | 安徽精英智能科技有限公司 | 一种智能语音交互服务方法及系统 |
| CN111176607A (zh) * | 2019-12-27 | 2020-05-19 | 国网山东省电力公司临沂供电公司 | 一种基于电力业务的语音交互系统及方法 |
| CN111223485A (zh) * | 2019-12-19 | 2020-06-02 | 深圳壹账通智能科技有限公司 | 智能交互方法、装置、电子设备及存储介质 |
| CN111367492A (zh) * | 2020-03-04 | 2020-07-03 | 深圳市腾讯信息技术有限公司 | 网页页面展示方法及装置、存储介质 |
| CN111801731A (zh) * | 2019-01-22 | 2020-10-20 | 京东方科技集团股份有限公司 | 语音控制方法、语音控制装置以及计算机可执行非易失性存储介质 |
| CN113921003A (zh) * | 2021-07-27 | 2022-01-11 | 歌尔科技有限公司 | 语音识别方法、本地语音识别装置及智能电子设备 |
| CN115396709A (zh) * | 2022-08-22 | 2022-11-25 | 海信视像科技股份有限公司 | 显示设备、服务器及免唤醒语音控制方法 |
| CN116416985A (zh) * | 2022-01-05 | 2023-07-11 | 博泰车联网(南京)有限公司 | 应用于语音交互的页面扫描控制方法、系统、设备和介质 |
Families Citing this family (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108109618A (zh) * | 2016-11-25 | 2018-06-01 | 宇龙计算机通信科技(深圳)有限公司 | 语音交互方法、系统以及终端设备 |
| CN106782561A (zh) * | 2016-12-09 | 2017-05-31 | 深圳Tcl数字技术有限公司 | 语音识别方法和系统 |
| CN106792047B (zh) * | 2016-12-20 | 2020-05-05 | Tcl科技集团股份有限公司 | 一种智能电视的语音控制方法及系统 |
| CN107845384A (zh) * | 2017-10-30 | 2018-03-27 | 江西博瑞彤芸科技有限公司 | 一种语音识别方法 |
| CN109785844A (zh) * | 2017-11-15 | 2019-05-21 | 青岛海尔多媒体有限公司 | 用于智能电视交互操作的方法及装置 |
| CN109741749B (zh) * | 2018-04-19 | 2020-03-27 | 北京字节跳动网络技术有限公司 | 一种语音识别的方法和终端设备 |
| CN110444200B (zh) * | 2018-05-04 | 2024-05-24 | 北京京东尚科信息技术有限公司 | 信息处理方法、电子设备、服务器、计算机系统及介质 |
| CN108877797A (zh) * | 2018-06-26 | 2018-11-23 | 上海早糯网络科技有限公司 | 主动交互式的智能语音系统 |
| CN110164411A (zh) * | 2018-07-18 | 2019-08-23 | 腾讯科技(深圳)有限公司 | 一种语音交互方法、设备及存储介质 |
| CN110795175A (zh) * | 2018-08-02 | 2020-02-14 | Tcl集团股份有限公司 | 模拟控制智能终端的方法、装置及智能终端 |
| CN109979449A (zh) * | 2019-02-15 | 2019-07-05 | 江门市汉的电气科技有限公司 | 一种智能灯具的语音控制方法、装置、设备和存储介质 |
| CN109859761A (zh) * | 2019-02-22 | 2019-06-07 | 安徽卓上智能科技有限公司 | 一种智能语音交互控制方法 |
| CN109785840B (zh) * | 2019-03-05 | 2021-01-29 | 湖北亿咖通科技有限公司 | 自然语言识别的方法、装置及车载多媒体主机、计算机可读存储介质 |
| CN110335602A (zh) * | 2019-07-10 | 2019-10-15 | 青海中水数易信息科技有限责任公司 | 一种具有语音识别功能的河长制信息化系统 |
| CN110517690A (zh) * | 2019-08-30 | 2019-11-29 | 四川长虹电器股份有限公司 | 语音控制功能的引导方法及系统 |
| CN110600003A (zh) * | 2019-10-18 | 2019-12-20 | 北京云迹科技有限公司 | 机器人的语音输出方法、装置、机器人和存储介质 |
| CN111475241B (zh) * | 2020-04-02 | 2022-03-11 | 深圳创维-Rgb电子有限公司 | 一种界面的操作方法、装置、电子设备及可读存储介质 |
| CN111627440A (zh) * | 2020-05-25 | 2020-09-04 | 红船科技(广州)有限公司 | 一种基于三维虚拟人物和语音识别实现交互的学习系统 |
| CN112767943A (zh) * | 2021-02-26 | 2021-05-07 | 湖北亿咖通科技有限公司 | 一种语音交互系统 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102740014A (zh) * | 2011-04-07 | 2012-10-17 | 青岛海信电器股份有限公司 | 语音控制电视机、电视系统及通过语音控制电视机的方法 |
| CN102855872A (zh) * | 2012-09-07 | 2013-01-02 | 深圳市信利康电子有限公司 | 基于终端及互联网语音交互的家电控制方法及系统 |
| CN102957711A (zh) * | 2011-08-16 | 2013-03-06 | 广州欢网科技有限责任公司 | 在电视上通过语音进行网址定位的方法及系统 |
| CN103093755A (zh) * | 2012-09-07 | 2013-05-08 | 深圳市信利康电子有限公司 | 基于终端及互联网语音交互的网络家电控制方法及系统 |
| CN103176591A (zh) * | 2011-12-21 | 2013-06-26 | 上海博路信息技术有限公司 | 一种基于语音识别的文本定位和选择方法 |
| CN105609104A (zh) * | 2016-01-22 | 2016-05-25 | 北京云知声信息技术有限公司 | 一种信息处理方法、装置及智能语音路由控制器 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103188409A (zh) * | 2011-12-29 | 2013-07-03 | 上海博泰悦臻电子设备制造有限公司 | 语音自动应答云端服务器、系统及方法 |
| CN104506901B (zh) * | 2014-11-12 | 2018-06-15 | 科大讯飞股份有限公司 | 基于电视场景状态及语音助手的语音辅助方法及系统 |
| CN104599669A (zh) * | 2014-12-31 | 2015-05-06 | 乐视致新电子科技(天津)有限公司 | 一种语音控制方法和装置 |
| CN105161106A (zh) * | 2015-08-20 | 2015-12-16 | 深圳Tcl数字技术有限公司 | 智能终端的语音控制方法、装置及电视机系统 |
| CN105512182B (zh) * | 2015-11-25 | 2019-03-12 | 深圳Tcl数字技术有限公司 | 语音控制方法及智能电视 |
| CN105551488A (zh) * | 2015-12-15 | 2016-05-04 | 深圳Tcl数字技术有限公司 | 语音控制方法及系统 |
-
2016
- 2016-07-06 CN CN201610529267.9A patent/CN106101789B/zh active Active
- 2016-09-06 WO PCT/CN2016/098147 patent/WO2018006489A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102740014A (zh) * | 2011-04-07 | 2012-10-17 | 青岛海信电器股份有限公司 | 语音控制电视机、电视系统及通过语音控制电视机的方法 |
| CN102957711A (zh) * | 2011-08-16 | 2013-03-06 | 广州欢网科技有限责任公司 | 在电视上通过语音进行网址定位的方法及系统 |
| CN103176591A (zh) * | 2011-12-21 | 2013-06-26 | 上海博路信息技术有限公司 | 一种基于语音识别的文本定位和选择方法 |
| CN102855872A (zh) * | 2012-09-07 | 2013-01-02 | 深圳市信利康电子有限公司 | 基于终端及互联网语音交互的家电控制方法及系统 |
| CN103093755A (zh) * | 2012-09-07 | 2013-05-08 | 深圳市信利康电子有限公司 | 基于终端及互联网语音交互的网络家电控制方法及系统 |
| CN105609104A (zh) * | 2016-01-22 | 2016-05-25 | 北京云知声信息技术有限公司 | 一种信息处理方法、装置及智能语音路由控制器 |
Cited By (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109584870A (zh) * | 2018-12-04 | 2019-04-05 | 安徽精英智能科技有限公司 | 一种智能语音交互服务方法及系统 |
| CN111801731A (zh) * | 2019-01-22 | 2020-10-20 | 京东方科技集团股份有限公司 | 语音控制方法、语音控制装置以及计算机可执行非易失性存储介质 |
| CN111801731B (zh) * | 2019-01-22 | 2024-02-13 | 京东方科技集团股份有限公司 | 语音控制方法、语音控制装置以及计算机可执行非易失性存储介质 |
| CN111223485A (zh) * | 2019-12-19 | 2020-06-02 | 深圳壹账通智能科技有限公司 | 智能交互方法、装置、电子设备及存储介质 |
| CN111176607A (zh) * | 2019-12-27 | 2020-05-19 | 国网山东省电力公司临沂供电公司 | 一种基于电力业务的语音交互系统及方法 |
| CN111367492A (zh) * | 2020-03-04 | 2020-07-03 | 深圳市腾讯信息技术有限公司 | 网页页面展示方法及装置、存储介质 |
| CN111367492B (zh) * | 2020-03-04 | 2023-07-18 | 深圳市腾讯信息技术有限公司 | 网页页面展示方法及装置、存储介质 |
| CN113921003A (zh) * | 2021-07-27 | 2022-01-11 | 歌尔科技有限公司 | 语音识别方法、本地语音识别装置及智能电子设备 |
| CN116416985A (zh) * | 2022-01-05 | 2023-07-11 | 博泰车联网(南京)有限公司 | 应用于语音交互的页面扫描控制方法、系统、设备和介质 |
| CN115396709A (zh) * | 2022-08-22 | 2022-11-25 | 海信视像科技股份有限公司 | 显示设备、服务器及免唤醒语音控制方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106101789B (zh) | 2020-04-24 |
| CN106101789A (zh) | 2016-11-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017143692A1 (zh) | 智能电视及其语音控制方法 | |
| WO2019051902A1 (zh) | 终端控制方法、空调器及计算机可读存储介质 | |
| WO2020060325A1 (ko) | 전자 장치, 시스템 및 음성 인식 서비스 이용 방법 | |
| WO2019061613A1 (zh) | 贷款资质筛选方法、装置及计算机可读存储介质 | |
| WO2014003283A1 (en) | Display apparatus, method for controlling display apparatus, and interactive system | |
| WO2017028601A1 (zh) | 智能终端的语音控制方法、装置及电视机系统 | |
| WO2018023926A1 (zh) | 电视与移动终端的互动方法及系统 | |
| WO2019104876A1 (zh) | 保险产品的推送方法、系统、终端、客户终端及存储介质 | |
| WO2019062113A1 (zh) | 家电设备的控制方法、装置、家电设备及可读存储介质 | |
| WO2019085543A1 (zh) | 电视机系统及电视机控制方法 | |
| WO2017101266A1 (zh) | 语音控制方法及系统 | |
| WO2017054488A1 (zh) | 电视播放控制方法、服务器及电视播放控制系统 | |
| WO2016032021A1 (ko) | 음성 명령 인식을 위한 장치 및 방법 | |
| WO2016058258A1 (zh) | 终端远程控制方法和系统 | |
| WO2019114262A1 (zh) | 加载用户界面的方法、智能电视及计算机可读存储介质 | |
| WO2019041851A1 (zh) | 家电售后咨询方法、电子设备和计算机可读存储介质 | |
| WO2019114127A1 (zh) | 空气调节器的语音播报方法及装置 | |
| WO2017206377A1 (zh) | 同步播放节目的方法和装置 | |
| WO2018233221A1 (zh) | 多窗口声音输出方法、电视机以及计算机可读存储介质 | |
| WO2018036057A1 (zh) | 软件后台自适应升级方法及装置 | |
| WO2017036208A1 (zh) | 显示界面中的信息提取方法及系统 | |
| WO2019062112A1 (zh) | 空调器控制方法、装置、空调器及计算机可读存储介质 | |
| WO2018006581A1 (zh) | 智能电视的播放方法及装置 | |
| WO2019210574A1 (zh) | 消息处理方法、装置、设备及可读存储介质 | |
| WO2017185482A1 (zh) | 多媒体会话方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16907993 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 27.05.2019) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16907993 Country of ref document: EP Kind code of ref document: A1 |