WO2020042530A1 - 自然语言的数据查询意图确定方法、装置和计算机设备 - Google Patents
自然语言的数据查询意图确定方法、装置和计算机设备 Download PDFInfo
- Publication number
- WO2020042530A1 WO2020042530A1 PCT/CN2019/071606 CN2019071606W WO2020042530A1 WO 2020042530 A1 WO2020042530 A1 WO 2020042530A1 CN 2019071606 W CN2019071606 W CN 2019071606W WO 2020042530 A1 WO2020042530 A1 WO 2020042530A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- query
- range
- data query
- presentation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/903—Querying
- G06F16/9032—Query formulation
Definitions
- the present application relates to the field of data query technology, and in particular, to a method, an apparatus, a computer device, and a storage medium for determining a data query intention in natural language.
- the server end forms the query items selected by the user into the user's query content, and obtains the query results based on the query content.
- the purpose of the present application is to provide a method, an apparatus, a computer device, and a storage medium for determining a data query intention in natural language, which are used to solve the foregoing problems in the prior art.
- the present application provides a method for determining data query intent in natural language.
- the method includes: presetting a plurality of keyword filters representing a query range, wherein each of the keyword filters corresponds to a range word set, and the range word set includes a plurality of range words; obtaining a natural-based Language data query request; segment the data query request to obtain a first word set; use each of the keyword filters to filter words in the first word set that match the range word; according to all the filtered words
- the words obtain the query range of the data query request; remove the words that match the range words in the first word set to obtain a second word set; and semantically mark the words in the second word set according to the semantic knowledge base Generating a semantic analysis result; determining a data presentation dimension and a data presentation method corresponding to the data query request according to the semantic analysis result; and outputting the query range, the data presentation dimension, and the data presentation method
- Three parameters are used as the standardized data query intent corresponding to the data query request.
- the present application provides a device for determining data query intention in natural language.
- the device includes: a plurality of keyword filters, wherein the keyword filters are used to characterize a query range, and each of the keyword filters corresponds to a range word set, and the range word set includes a plurality of range words; A module for obtaining a natural language-based data query request to be analyzed; a word segmentation module for segmenting the data query request to obtain a first word set; a calling module for calling each of the keyword filter filters Words in the first word set that match the range words; a first determining module configured to obtain a query range of the data query request according to all filtered words; a filtering module configured to convert the first words The words in the second word set are collected by removing the words that match the scope word, and a second word set is obtained.
- the tagging module is used for semantically tagging the words in the second word set according to the semantic knowledge base to generate a semantic analysis result.
- the second determining module uses For determining the data presentation dimension and data presentation method corresponding to the data query request according to the semantic analysis result; and an output module for outputting all data Range query, data rendering method of three kinds of parameters of the dimensions and the presentation of data, as the data corresponding to the query request normalized data query intentions.
- the present application further provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor.
- the processor implements data query in natural language when the processor executes the program.
- each of the keyword filters corresponds to a range word set, the range word set includes a plurality of range words; obtaining a natural language-based data query to be analyzed Request; word segmentation of the data query request to obtain a first word set; using each of the keyword filters to filter words in the first word set that match the range word; to obtain the word based on all the filtered words Query range of a data query request; removing words that match the range word in the first word set to obtain a second word set; performing semantic annotation on the words in the second word set according to the semantic knowledge base to generate a semantic analysis Result; determining, according to the result of the semantic analysis, a data presentation dimension and a data presentation method corresponding to the data query request; and outputting three parameters of the query range, the data presentation dimension, and the data presentation method, As a normalized data query intent corresponding to a data query request.
- the present application also provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, implements the following steps of a method for determining data query intent in natural language:
- each of the keyword filters corresponds to a range word set, the range word set includes a plurality of range words; obtaining a natural language-based data query to be analyzed Request; word segmentation of the data query request to obtain a first word set; using each of the keyword filters to filter words in the first word set that match the range word; to obtain the word based on all the filtered words Query range of a data query request; removing words that match the range word in the first word set to obtain a second word set; performing semantic annotation on the words in the second word set according to the semantic knowledge base to generate a semantic analysis Result; determining, according to the result of the semantic analysis, a data presentation dimension and a data presentation method corresponding to the data query request; and outputting three parameters of the query range, the data presentation dimension, and the data presentation method, As a normalized data query intent corresponding to a data query request.
- Method, device, computer equipment and storage medium for determining natural language data query intent provided by the present application preset keyword filter, and the scope filter in the scope word set corresponding to the keyword filter can be filtered by the keyword filter Out, therefore, by setting a range word combination as required, a predetermined range word can be filtered.
- a natural language-based data query request After obtaining a natural language-based data query request, first segment the data query request to obtain a word set, and then filter the word set through a preset keyword filter to match the word set with the range word. The words are filtered out, and the query range can be formed based on the filtered words.
- the words filtered in the word set are removed, the remaining words are semantically labeled, and the dimensions and method of data presentation are determined according to the results of the semantic annotation.
- the three parameters of the query range, the dimension of the data presentation, and the method of the data presentation are used as the standardized data query intent corresponding to the data query request, so that when performing data query, the unified query logic can be used to perform the query according to the standardized query intent. To improve query efficiency.
- FIG. 1 is a flowchart of a method for determining a data query intention of a natural language according to Embodiment 1 of the present application;
- FIG. 1 is a flowchart of a method for determining a data query intention of a natural language according to Embodiment 1 of the present application;
- FIG. 2 is a block diagram of an apparatus for determining a data query intention of a natural language provided in Embodiment 2 of the present application.
- FIG. 3 is a hardware structural diagram of a computer device according to Embodiment 3 of the present application.
- FIG. 1 is a natural language A flowchart of a data query intent determination method is shown in FIG. 1.
- the natural language data query intent determination method includes the following steps S101 to S109:
- Step S101 preset a plurality of keyword filters representing a query range.
- Each keyword filter corresponds to a range word set, and the range word set includes multiple range words.
- setting a range word set corresponding to a keyword filter includes range words including Shanghai, Tianjin, Beijing, Nanchang, and Shenyang. Etc.
- setting a range word set corresponding to another keyword filter includes range words such as old age, infants, adolescents, and middle-aged. How to set a range word set corresponding to a keyword filter can be based on data. Content settings.
- Step S102 Acquire a natural language-based data query request to be analyzed.
- the specific receiving method can be text input or voice input.
- voice input it can be recognized as text in the background, regardless of the type.
- the natural language-based data query request is directly reflected in natural language.
- the natural language-based data query request is "how about the overall overdue situation of customers in Shanghai", and the natural language-based data query request can be "What was the total sales of the first sales team in March?"
- Step S103 segment the data query request to obtain a first word set.
- the first word set obtained is “Shanghai”, “region”, “of”, “male”, “user”, “ “”, “educationion”, “yes”, “how”, “distribution”, “of”, preferably, the first word set obtained by the segmentation can be filtered to filter out useless words, such as "", “ Yes "and so on.
- Step S104 Use each keyword filter to filter the words in the first word set that match the range words.
- the keywords in the first word set are filtered through preset keyword filters to obtain words in the first word set that match the range words.
- Step S105 Obtain the query range of the data query request according to all the filtered words.
- a word matching the scope word can be obtained through one filter, and a plurality of words matching the scope word can be obtained through multiple filters.
- Words each word that matches the range word is combined to get the query range of the data query request. For example, when a keyword filter is used to filter the first word set corresponding to "how are the academic qualifications of male users in Shanghai distributed", the words that match the scope word include “Shanghai” and "male", so , The query range of the obtained data query request is "Shanghai male".
- Step S106 Remove the words that match the range words in the first word set to obtain a second word set.
- the second word set obtained after removing these two words includes "region", “user”, “educational level”, “how” And “distribution.”
- Step S107 semantically mark the words in the second word set according to the semantic knowledge base to generate a semantic analysis result.
- the semantic analysis result includes a plurality of word-meaning pairs, and the word-meaning pairs include a word and the semantics of the words in the second word set.
- the semantics of words include parts of speech and meanings.
- the parameters corresponding to each word are determined according to the semantic knowledge base.
- the words in the second word set include "region”, “user”, and “education” , “How” and “distribution”, according to the semantic knowledge base, "region is a noun, administrative division”, “user is a noun, the category of person”, “education is a noun, a description of education Way ",” how is a question word “and” distribution is a verb ", then the results of the semantic analysis can be shown in the following way:
- Step S108 Determine the dimensions of the data presentation corresponding to the data query request and the method of the data presentation according to the result of the semantic analysis.
- the semantics corresponding to the dimensions of each data presentation in the data content are preset.
- the data content is census data for young people between the ages of 25 and 30.
- the dimensions of data presentation include "educational qualifications”,
- the semantics corresponding to "annual income”, “annual consumption”, and “fertility status” are preset for "educational background”, “annual income”, “annual consumption”, and “fertility status”, respectively.
- determining the dimensions of the data presentation corresponding to the data query request according to the semantic analysis results specifically includes:
- the dimensions presented by the first data are used as the dimensions presented by the data query request.
- the method of data presentation refers to a way of expressing data, for example, the age data of a school student.
- the method of data presentation includes presenting the data by the average of all ages, and the ratio of each age group. Present data, present data by the distribution of each age value, and more.
- the semantics corresponding to each data presentation method in the data content are preset.
- the data content is census data for young people between the ages of 25 and 30.
- the data presentation methods include "distribution”, “Average”, etc., preset the semantics corresponding to "distribution” and "average”, respectively.
- the method for determining the data presentation corresponding to the data query request according to the result of the semantic analysis specifically includes:
- the first data presentation method is used as the dimension presented by the data query request corresponding method.
- Step S109 Output three parameters of the query range, the dimension of the data presentation, and the method of the data presentation as the standardized data query intent corresponding to the data query request.
- Each output data query intent includes three parameters: query range, dimension of data presentation, and method of data presentation. It becomes a standardized and uniform data query intent.
- a keyword filter is preset, and the scope word in the scope word set corresponding to the keyword filter can be filtered out through the keyword filter.
- Set the scope word combination to filter the predetermined scope words.
- the three parameters of the query range, the dimension of the data presentation, and the method of the data presentation are used as the standardized data query intent corresponding to the data query request, so that when the data query is performed, the unified query logic can be used to perform the query in accordance with the standardized query intent. To improve query efficiency.
- the range word can be obtained at most one word that matches the range word.
- query through the data A corresponding number of data query intents can be obtained in the request, where the query scope of each data query intent is different, and the data presentation dimensions and data presentation methods are the same.
- the data query request is "the male user in Shanghai and Beijing "How are academic degrees distributed?"
- the query range includes "Males in Shanghai” and "Males in Beijing”.
- the present application may determine an incomplete data query request, that is, a data query request that does not include the entire content of the query range, the dimension of the data presentation, and the method of the data presentation Determine the data query intent.
- an incomplete data query request that is, a data query request that does not include the entire content of the query range, the dimension of the data presentation, and the method of the data presentation Determine the data query intent.
- the embodiments provided in the present application are supplemented first, and then the data query intent is determined, so that the incomplete data query requests can also output standardized and uniform data query intents.
- the historical data query intent when each keyword filter fails to filter the words in the first word set that match the range words, the historical data query intent is obtained; the dimensions and data presentation methods are determined.
- the historical data query intent with the highest matching degree specifically, the historical data query intent including the dimensions of the data presentation and the method of data presentation among multiple historical data query intents.
- the historical data query intent is For historical data query intent that has the highest degree of matching with the dimensions and method of data presentation, if multiple and multiple historical data query intents include historical data query intents with the same content, query historical data with the same content
- the historical data query intent with the largest number of intents is used as the historical data query intent with the highest degree of matching with the dimensions and methods of data presentation; the historical query range in the historical data query intent with the highest matching is obtained as the query for the data query request Scope, for example, the data query request to be analyzed is "Average annual income", the dimension of the available data is “annual income”, and the method of data presentation is "average”. At this time, the history including the most "annual income” and "average” is found in multiple historical data query intents.
- the data query intention is the historical query query with the highest matching degree.
- the historical query range of the historical query query with the highest matching degree is "Beijing Teen", and the "Beijing Teen” is the corresponding data query request "average annual income” Query range.
- the color of the text font representing the query range is set to gray, that is, for the above The data query request "average annual income”.
- the text font color of "Beijing Teen” is set to gray to remind the user to make the user determine the query range.
- the historical data query intent is obtained; the historical data query intent that has the highest degree of matching with the query range and the method of data presentation is determined Specifically, look for historical data query intent including query range and data presentation method among multiple historical data query intents. If one query is found, the historical data query intent is matched with the query range and data presentation method. The highest degree of historical data query intent.
- the historical data query intent with the highest number of historical data query intents with the same content is used as The query range and data presentation method have the most matching historical data query intent; obtain the historical data query intent with the highest matching historical data query intent as the data presentation request data query dimension, for example, the data query request to be analyzed "The average male love in Shanghai ", It can be obtained that the query range is" Shanghai men "and the data presentation method is" average ".
- the historical data query intents including" Shanghai male "and” average are found as matches The highest historical data query intent.
- the dimension of the data with the highest matching historical data query intent is "annual income", then the "annual income” is presented as the data corresponding to the "average situation of Shanghai men” in the data query request.
- Dimensions when the three parameters of the query range, the dimension of the data presentation and the method of the data presentation are output as the steps of the standardized data query intention corresponding to the data query request, the color of the text font characterizing the dimension of the data presentation is set to gray, that is, For the above data query request "average situation of Shanghai men", when outputting the standardized data query intent, the text font color of "Annual Income” is set to gray to remind the user in particular, so that the user can determine the dimensions of the data presentation.
- the historical data query intent is obtained; the historical data query intent that has the highest degree of matching with the query range and the dimensions of the data presentation is determined , Specifically, look for historical data query intent including query range and dimensions of data presentation among multiple historical data query intents. If one query is found, the historical data query intent is a method related to query range and data presentation The most matching historical data query intent.
- the historical data query intent with the highest number of historical data query intents with the same content is taken as The historical data query intent that matches the query range and the dimensions of the data presentation with the highest matching degree; the method of obtaining historical data in the query intent with the highest matching historical data is used as the data presentation method for the data query request, for example, the data query to be analyzed Request for "Shanghai Men's Annual Income ", The query scope is" Shanghai men ", and the dimension of the data presentation is" annual income ". At this time, the historical data query including" Shanghai men “and” annual income "is found in multiple historical data query intents. The intent is the query with the highest matching historical data.
- the method of presenting the data in the query with the highest matching historical data is "average”. Then, use the "average” as the data corresponding to the "Shanghai male annual income” data query request.
- Method of rendering when the three parameters of the query range, the dimension of the data presentation, and the method of the data presentation are output as the steps of the standardized data query intention corresponding to the data query request, the color of the text font characterizing the method of data presentation is set to gray, that is, For the above data query request "Shanghai Male Annual Income", when the standardized data query intent is output, the "average” text font color is set to gray to remind the user specifically, so that the user determines the method of data presentation.
- Embodiment 2 of the present application provides a data query device for natural language, and relevant parts can be cross-referenced with the above-mentioned Embodiment 1.
- 2 is a block diagram of a natural language data query intent determination apparatus provided in Embodiment 2 of the present application. As shown in FIG. 2, the apparatus includes multiple keyword filters 201, an acquisition module 202, a word segmentation module 203, a call module 204, The first determination module 205, the screening module 206, the labeling module 207, the second determination module 208, and the output module 209.
- the keyword filter 201 is used to characterize the query range, and each keyword filter corresponds to a range word set, and the range word set includes multiple range words.
- the acquisition module 202 is used to obtain a natural language-based data query request to be analyzed.
- the word segmentation module 203 is used to segment the data query request to obtain the first word set;
- the call module 204 is used to call each keyword filter to filter the words in the first word set that match the range words;
- the first determination module 205 is used to All the filtered words get the query range of the data query request;
- the filtering module 206 is used to remove the words that match the range word in the first word set to get the second word set;
- the labeling module 207 is used to the words in the second word set
- the semantic annotation is performed according to the semantic knowledge base to generate a semantic analysis result;
- the second determination module 208 is used to determine the data presentation dimension and data presentation method corresponding to the data query request according to the semantic analysis result;
- the output module 209 is used
- the natural language data query intent determining device is used to preset a keyword filter.
- the keyword filter can filter out the range words in the range word set corresponding to the keyword filter. Therefore, as needed, Set the scope word combination to filter the predetermined scope words.
- the word segmentation module After the acquisition module obtains the natural language-based data query request, the word segmentation module performs word segmentation on the data query request to obtain a word set.
- the calling module calls a preset keyword filter to filter the word set, and associates the word set with the word set. Words with matching range words are filtered out, and the first determining module can form a query range based on the filtered words.
- the filtering module removes the words filtered from the word set, the labeling module performs semantic labeling on the remaining words, and the second determination module determines the dimensions and method of data presentation according to the results of the semantic labeling.
- the output module uses the three parameters of the query range, the dimension of the data presentation, and the data presentation method as the standardized data query intent corresponding to the data query request, so that when performing data query, the unified query logic can be used in accordance with the standardized query intent. Perform queries to improve query efficiency.
- the device further includes a first supplementary module.
- the first supplementary module is configured to perform the following steps: obtaining historical data query Intent; determine the historical data query intent with the highest degree of matching with the dimensions and method of data presentation; obtain the historical query range in the historical data query intent with the highest matching as the query range of the data query request.
- the output module sets the color of the text font representing the query range to gray when outputting the three parameters of the query range, the dimension of the data presentation, and the data presentation method as the steps of the standardized data query intent corresponding to the data query request.
- the device further includes a second supplementary module.
- the second supplementary module is configured to perform the following steps: obtaining historical data query Intent; determine the historical data query intent that has the highest degree of matching with the query range and the method of data presentation; obtain the dimension of the historical data that is present in the query query with the highest degree of matching as the dimension of the data presentation requested by the data query.
- the output module sets three parameters of the query range, the dimension of the data presentation, and the method of the data presentation as the steps of the standardized data query intention corresponding to the data query request, and sets the color of the text font representing the dimension of the data to gray.
- the device further includes a third supplementary module.
- the third supplementary module is configured to perform the following steps: obtaining historical data query Intent; determine the historical data query intent that has the highest degree of matching with the query range and the dimension of the data presentation; obtain the historical data presentation method in the historical intent that has the highest matching query as the data presentation method for the data query request, where the output module When the three parameters of the query range, the dimension of the data presentation, and the method of the data presentation are output as the steps of the standardized data query intent corresponding to the data query request, the text font color representing the method of the data presentation is set to gray.
- the semantic analysis result generated by the annotation module includes a plurality of word-sense pairs, and the word-sense pair includes a word in the second word set and the semantics of the words.
- the device further includes a first preset module for presetting the semantics corresponding to the dimensions of the data presentation in the data content, and the second determination module specifically executes the dimensions of the data presentation corresponding to the data query request according to the semantic analysis result.
- the following steps match the semantics of the word-sense pair with the semantics corresponding to the dimensions presented by the data content; when the semantics of the word-sense pair with the semantics corresponding to the dimension presented by the first data in the data content Using the dimension presented by the first data as the dimension presented by the data query request.
- the device further includes a second preset module for presetting the semantics corresponding to each data presentation method in the data content.
- the second determination module specifically executes the method for determining the data presentation method corresponding to the data query request according to the semantic analysis result. The following steps: match the semantics in the word-sense pair with the semantics corresponding to the data presentation method; when the semantics in the word-sense pair are the same as the semantics corresponding to the first data presentation method in the data content Using the method of presenting the first data as a method of presenting the data corresponding to the data query request.
- This embodiment also provides a computer device, such as a smart phone, tablet computer, notebook computer, desktop computer, rack server, blade server, tower server, or rack server (including a stand-alone server, or Server cluster consisting of multiple servers).
- the computer device 20 of this embodiment includes, but is not limited to, a memory 21 and a processor 22 that can be communicatively connected to each other through a system bus, as shown in FIG. 3.
- FIG. 3 only shows the computer device 20 with components 21-22, but it should be understood that it is not required to implement all the illustrated components, and more or fewer components may be implemented instead.
- the memory 21 (ie, a readable storage medium) includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), Read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disks, optical disks, etc.
- the memory 21 may be an internal storage unit of the computer device 20, such as a hard disk or a memory of the computer device 20.
- the memory 21 may also be an external storage device of the computer device 20, for example, a plug-in hard disk, a smart media card (SMC), and a secure digital (Secure Digital, SD) card, flash card, etc.
- the memory 21 may also include both the internal storage unit of the computer device 20 and its external storage device.
- the memory 21 is generally used to store an operating system and various types of application software installed in the computer device 20, for example, a program code of the natural language data query device of the second embodiment.
- the memory 21 may also be used to temporarily store various types of data that have been output or are to be output.
- the processor 22 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments.
- the processor 22 is generally used to control the overall operation of the computer device 20.
- the processor 22 is configured to run program code or process data stored in the memory 21, such as a data query device for natural language.
- This embodiment also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), Read memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disks, optical disks, servers, App application stores, etc., which have computer programs stored on them, When the program is executed by the processor, the corresponding function is realized.
- the computer-readable storage medium of this embodiment is used for a natural language data query device, and when executed by a processor, implements the natural language data query method of embodiment 1.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Databases & Information Systems (AREA)
- Theoretical Computer Science (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Machine Translation (AREA)
Abstract
一种自然语言的数据查询意图确定方法、装置、计算机设备及存储介质,涉及自然语言处理(NLP,Natural Language Processing)技术领域。该方法包括预设多个关键词过滤器(S101);获取基于自然语言的数据查询请求(S102);对该请求进行分词得到第一词集(S103);采用各个关键词过滤器过滤第一词集(S104);根据过滤到的词得到查询范围(S105);将第一词集中与范围词相符的词去除得到第二词集(S106);对第二词集中的词进行语义标注生成语义分析结果(S107);根据语义分析结果确定数据呈现的维度和数据呈现的方法(S108);输出查询范围、数据呈现的维度和数据呈现的方法作为标准化数据查询意图(S109)。能够提升基于自然语言的数据查询效率。
Description
本申请申明享有2018年8月31日递交的申请号为CN 201811021831.1、名称为“自然语言的数据查询意图确定方法、装置和计算机设备”的中国专利申请的优先权,该中国专利申请的整体内容以参考的方式结合在本申请中。
本申请涉及数据查询技术领域,尤其涉及一种自然语言的数据查询意图确定方法、装置、计算机设备及存储介质。
目前,用户进行数据的查询时,通常需要用户按照自己的查询意图,先选定查询界面上设置的选择项,然后服务器一端将用户选择的查询项形成用户的查询内容,基于查询内容得到查询结果。
然而,因为是通过选择查询项的方式进行搜索,这就将导致,用户只能按照选择项中所提供的选择项进行选择,在选择项太少的时候,用户的选择范围会受到限制,在选择项过多的时候,用户在选择选择项的时候,选择的操作比较复杂。为了使用户的搜索和查询更加方便,基于自然语言的查询成为未来的重要的查询方式,其中,识别自然语言的数据查询意图是基于自然语言的查询的基础。
因而,提供一种自然语言的数据查询意图确定方法、装置、计算机设备及存储介质,以准确的从自然语言中捕捉用户数据查询意图,明确的提取出提问范围和分析维度,是本领域需要解决的技术问题。
发明内容
本申请的目的是提供一种自然语言的数据查询意图确定方法、装置、计算机设备及存储介质,用于解决现有技术存在的上述问题。
为实现上述目的,本申请提供一种自然语言的数据查询意图确定方法。
该方法包括:预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求 的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
为实现上述目的,本申请提供一种自然语言的数据查询意图确定装置。
该装置包括:多个关键词过滤器,其中,所述关键词过滤器用于表征查询范围,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取模块,用于获取待分析的基于自然语言的数据查询请求;分词模块,用于对所述数据查询请求进行分词,得到第一词集;调用模块,用于调用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;第一确定模块,用于根据所有过滤到的词得到所述数据查询请求的查询范围;筛除模块,用于将所述第一词集中与所述范围词相符的词去除,得到第二词集;标注模块,用于对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;第二确定模块,用于根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出模块,用于输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
为实现上述目的,本申请还提供一种计算机设备,包括存储器、处理器以及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现自然语言的数据查询意图确定方法的以下步骤:
预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
为实现上述目的,本申请还提供计算机可读存储介质,其上存储有计算机程序,所述程序被处理器执行时实现自然语言的数据查询意图确定方法的以下步骤:
预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一 词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
本申请提供的自然语言的数据查询意图确定方法、装置、计算机设备及存储介质,预设关键词过滤器,通过该关键词过滤器能够将关键词过滤器对应的范围词集合中的范围词过滤出来,因而,根据需要设置范围词结合,就可将预定的范围词进行过滤。在获取到基于自然语言的数据查询请求之后,首先对数据查询请求进行分词,得到一个词集,然后通过预设的关键词过滤器对该词集进行过滤,将该词集中与范围词相符的词过滤出来,根据过滤出来的词能够形成查询范围。得到查询范围后,将词集中过滤出来的词去除掉,对剩余的词进行语义标注,按照语义标注的结果确定出数据呈现的维度和数据呈现的方法。最后,以查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图,从而在进行数据查询时,能够利用统一的查询逻辑按照标准化的查询意图进行查询,提高查询效率。
图1为本申请实施例1提供的自然语言的数据查询意图确定方法的流程图;
图2为本申请实施例2提供的自然语言的数据查询意图确定装置的框图。
图3为本申请实施例3提供的计算机设备的硬件结构图。
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施例仅用以解释本申请,并不用于限定本申请。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
实施例1
为了使用户能够基于自然语言表达查询内容,提高查询效率,对于服务器而言直接根据用户的自然语言生成查询结果,该实施例1提供一种自然语言的数据查询意图确定方法,该方法将基于自然语言的数据查询请求转化为标准化数据查询意图,其中,数据查询意图包括查询范围、数据呈现的维度和数据呈现的方法三种参数,具体地,图1为本申请实施 例1提供的自然语言的数据查询意图确定方法的流程图,如图1所示,该自然语言的数据查询意图确定方法包括如下的步骤S101至步骤S109:
步骤S101:预设多个表征查询范围的关键词过滤器。
其中,每个关键词过滤器对应一个范围词集合,范围词集合包括多个范围词,例如,设置一个关键词过滤器对应的范围词集合包括的范围词有上海、天津、北京、南昌和沈阳等,又如,设置另一个关键词过滤器对应的范围词集合包括的范围词有老年、婴幼儿、青少年和中年等,具体怎样设置一个关键词过滤器对应的范围词集合,可根据数据内容进行设定。
步骤S102:获取待分析的基于自然语言的数据查询请求。
提供接收用户查询请求的接口,以接收用户输入的基于自然语言的数据查询请求,具体接收方式可以为文字输入,也可以为语音输入,对于语音输入,可在后台识别为文字,无论以何种输入方式,基于自然语言的数据查询请求直接是通过自然语言体现的,例如,基于自然语言的数据查询请求为“上海地区客户的整体逾期情况怎么样”,又如基于自然语言的数据查询请求可以为“三月份第一销售小组的销售总额是多少”等。
步骤S103:对数据查询请求进行分词,得到第一词集。
例如,对于数据查询请求“上海地区的男性用户的学历是如何分布的”进行分词,得到的第一词集为“上海”、“地区”、“的”、“男性”、“用户”、“的”、“学历”、“是”、“如何”、“分布”、“的”,优选地,可将分词得到的第一词集进行过滤,过滤掉无用的词,例如“的”、“是”等。
步骤S104:采用各个关键词过滤器过滤第一词集中与范围词相符的词。
在该步骤中,通过预先设置的各个关键词过滤器,对第一词集中的词进行过滤,以得到第一词集中与范围词相符的词。
步骤S105:根据所有过滤到的词得到数据查询请求的查询范围。
在上述步骤S104中,可以通过一个过滤器得到一个与范围词相符的词,也可以通过多个过滤器得到多个与范围词相符的词,当通过多个过滤器得到多个与范围词相符的词时,各个与范围词相符的词组合得到数据查询请求的查询范围。例如,采用关键词过滤器对“上海地区的男性用户的学历是如何分布的”所对应的第一词集进行过滤时,得到的与范围词相符的词包括“上海”和“男性”,因而,得到的数据查询请求的查询范围为“上海男性”。
步骤S106:将第一词集中与范围词相符的词去除,得到第二词集。
当通过步骤S104得到的与范围词相符的词包括“上海”和“男性”时,去除这两个词之后得到的第二词集包括“地区”、“用户”、“学历”、“如何”和“分布”。
步骤S107:对第二词集中的词根据语义知识库进行语义标注,生成语义分析结果。
例如,语义分析结果包括多个词-义对,词-义对包括第二词集中的一个词和词的语义。
其中,词的语义包括词性和词义,在进行语义标注时,根据语义知识库,来确定各个词所对应的参数,例如,第二词集中的词包括“地区”、“用户”、“学历”、“如何”和“分布”,根据语义知识库,可确定“地区是一个名词,行政划分单位”,“用户是一个名词,人的类别”,“学历是一个名词,一种描述受教育程度的方式”、“如何是一个疑问词”和“分布是一个动词”,则语义分析结果可通过如下方式展示:
<地区,名词,行政划分单位>;
<用户,名词,人的类别>;
<学历,名词,描述受教育程度的方式>;
<如何,代词,表示疑问>;
<分布,动词,表示某种客体在一定范围散布>。
步骤S108:根据语义分析结果确定数据查询请求对应的数据呈现的维度和数据呈现的方法。
对于数据呈现的维度的确定,预设数据内容中各个数据呈现的维度所对应的语义,例如,数据内容为针对全国25至30岁之间青年的普查数据,数据呈现的维度包括“学历”、“年收入”、“年消费”和“生育情况”等,分别预设“学历”、“年收入”、“年消费”和“生育情况”所对应的语义。
在该步骤中,根据语义分析结果确定数据查询请求对应的数据呈现的维度具体包括:
匹配词-义对中的语义和数据内容中各个数据呈现的维度所对应的语义;
当词-义对中的语义与数据内容中的第一数据呈现的维度所对应的语义相同时,将第一数据呈现的维度作为数据查询请求对应的数据呈现的维度。
其中,数据呈现的方法是指对数据进行表达的一种方式,例如针对某学校学生的年龄数据,数据呈现的方法包括通过所有年龄的平均值来呈现数据、通过各年龄段所占的比值来呈现数据、通过每个年龄值的分布情况来呈现数据等等。
对于数据呈现的方法的确定,预设数据内容中各个数据呈现的方法所对应的语义,例如,数据内容为针对全国25至30岁之间青年的普查数据,数据呈现的方法包括“分布”、“平均”等,分别预设“分布”、“平均”所对应的语义。在该步骤中,根据语义分析结果确定数据查询请求对应的数据呈现的方法具体包括:
匹配词-义对中的语义和数据内容中各个数据呈现的方法所对应的语义;
当词-义对中的语义与数据内容中的第一数据呈现的方法所对应的语义相同时,将第一数据呈现的方法作为数据查询请求对应的方法呈现的维度。
步骤S109:输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查 询请求对应的标准化数据查询意图。
输出的每个数据查询意图均包括查询范围、数据呈现的维度和数据呈现的方法三种参数,成为标准化的、统一的数据查询意图。
采用该实施例提供的自然语言的数据查询意图确定方法,预设关键词过滤器,通过该关键词过滤器能够将关键词过滤器对应的范围词集合中的范围词过滤出来,因而,根据需要设置范围词结合,就可将预定的范围词进行过滤。在获取到基于自然语言的数据查询请求之后,首先对数据查询请求进行分词,得到一个词集,然后通过预设的关键词过滤器对该词集进行过滤,将该词集中与范围词相符的词过滤出来,根据过滤出来的词能够形成查询范围。得到查询范围后,将词集中过滤出来的词去除掉,对剩余的词进行语义标注,按照语义标注的结果确定出数据呈现的维度和数据呈现的方法。最后,以查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图,从而在进行数据查询时,能够利用统一的查询逻辑按照标准化的查询意图进行查询,提高查询效率。
通常情况下,对于一个过滤器,至多可得到一个与范围词相符的词,在一种具体的实施例中,如果得到两个或两个以上的与范围词相符的词时,通过该数据查询请求可得到相应个数的数据查询意图,其中,每个数据查询意图的查询范围不同,数据呈现的维度和数据呈现的方法相同,例如,数据查询请求为“上海地区和北京地区的男性用户的学历是如何分布的”,得到的查询范围包括“上海地区的男性”和“北京地区的男性”。
可选地,本申请在确定数据查询意图时,可对不完整的数据查询请求进行确定,也即,可对不包括查询范围、数据呈现的维度和数据呈现的方法的全部内容的数据查询请求进行数据查询意图的确定。对于不完整的数据查询请求,本申请提供的实施例首先进行补充,然后再进行数据查询意图的确定,以使不完整的数据查询请求也能输出标准化的、统一的数据查询意图。
具体地,在一种实施例中,当采用各个关键词过滤器均过滤不到第一词集中与范围词相符的词时,获取历史数据查询意图;确定与数据呈现的维度和数据呈现的方法匹配度最高的历史数据查询意图,具体地,在多条历史数据查询意图中查找包括数据呈现的维度和数据呈现的方法的历史数据查询意图,如果查询到一条,则该条历史数据查询意图即为与数据呈现的维度和数据呈现的方法匹配度最高的历史数据查询意图,如果查询到多条且多条历史数据查询意图中包括内容相同的历史数据查询意图,则将内容相同的历史数据查询意图中条数最多的历史数据查询意图作为与数据呈现的维度和数据呈现的方法匹配度最高的历史数据查询意图;获取匹配度最高的历史数据查询意图中的历史查询范围作为数据查询请求的查询范围,例如,待分析的数据查询请求为“平均年收入”,可得到数据呈现的维 度为“年收入”,数据呈现的方法为“平均”,此时,在多条历史数据查询意图中找到包括“年收入”和“平均”最多的历史数据查询意图作为匹配度最高的历史数据查询意图,该匹配度最高的历史数据查询意图中的历史查询范围为“北京青年”,则将“北京青年”作为数据查询请求“平均年收入”所对应的查询范围。其中,在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征查询范围的文字字体颜色设置为灰色,也即对于上述数据查询请求“平均年收入”,在输出标准化数据查询意图时,将“北京青年”的文字字体颜色设置为灰色,以向用户特别提醒,以使用户对查询范围进行确定。
在另一种实施例中,当根据语义分析结果无法确定出数据查询请求对应的数据呈现的维度时,获取历史数据查询意图;确定与查询范围和数据呈现的方法匹配度最高的历史数据查询意图,具体地,在多条历史数据查询意图中查找包括查询范围和数据呈现的方法的历史数据查询意图,如果查询到一条,则该条历史数据查询意图即为与查询范围和数据呈现的方法匹配度最高的历史数据查询意图,如果查询到多条且多条历史数据查询意图中包括内容相同的历史数据查询意图,则将内容相同的历史数据查询意图中条数最多的历史数据查询意图作为与查询范围和数据呈现的方法匹配度最高的历史数据查询意图;获取匹配度最高的历史数据查询意图中的历史数据呈现的维度作为数据查询请求的数据呈现的维度,例如,待分析的数据查询请求为“上海男性平均情况”,可得到查询范围为“上海男性”,数据呈现的方法为“平均”,此时,在多条历史数据查询意图中找到包括“上海男性”和“平均”最多的历史数据查询意图作为匹配度最高的历史数据查询意图,该匹配度最高的历史数据查询意图中的数据呈现的维度为“年收入”,则将“年收入”作为数据查询请求“上海男性平均情况”所对应的数据呈现的维度。其中,在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征数据呈现的维度的文字字体颜色设置为灰色,也即对于上述数据查询请求“上海男性平均情况”,在输出标准化数据查询意图时,将“年收入”的文字字体颜色设置为灰色,以向用户特别提醒,以使用户对数据呈现的维度进行确定。
在又一种实施例中,当根据语义分析结果无法确定出数据查询请求对应的数据呈现的方法时,获取历史数据查询意图;确定与查询范围和数据呈现的维度匹配度最高的历史数据查询意图,,具体地,在多条历史数据查询意图中查找包括查询范围和数据呈现的维度的历史数据查询意图,如果查询到一条,则该条历史数据查询意图即为与查询范围和数据呈现的方法匹配度最高的历史数据查询意图,如果查询到多条且多条历史数据查询意图中包括内容相同的历史数据查询意图,则将内容相同的历史数据查询意图中条数最多的历史数据查询意图作为与查询范围和数据呈现的维度匹配度最高的历史数据查询意图;获取匹配 度最高的历史数据查询意图中的历史数据呈现的方法作为数据查询请求的数据呈现的方法,例如,待分析的数据查询请求为“上海男性年收入”,可得到查询范围为“上海男性”,数据呈现的维度为“年收入”,此时,在多条历史数据查询意图中找到包括“上海男性”和“年收入”最多的历史数据查询意图作为匹配度最高的历史数据查询意图,该匹配度最高的历史数据查询意图中的数据呈现的方法为“平均”,则将“平均”作为数据查询请求“上海男性年收入”所对应的数据呈现的方法。其中,在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征数据呈现的方法的文字字体颜色设置为灰色,也即对于上述数据查询请求“上海男性年收入”,在输出标准化数据查询意图时,将“平均”的文字字体颜色设置为灰色,以向用户特别提醒,以使用户对数据呈现的方法进行确定。
实施例2
对应于上述实施例1,本申请实施例2提供了一种自然语言的数据查询装置,相关部分可与上述实施例1相互参考。图2为本申请实施例2提供的自然语言的数据查询意图确定装置的框图,如图2所示,该装置包括多个关键词过滤器201、获取模块202、分词模块203、调用模块204、第一确定模块205、筛除模块206、标注模块207、第二确定模块208和输出模块209。
其中,关键词过滤器201用于表征查询范围,每个关键词过滤器对应一个范围词集合,范围词集合包括多个范围词;获取模块202用于获取待分析的基于自然语言的数据查询请求;分词模块203用于对数据查询请求进行分词,得到第一词集;调用模块204用于调用各个关键词过滤器过滤第一词集中与范围词相符的词;第一确定模块205用于根据所有过滤到的词得到数据查询请求的查询范围;筛除模块206用于将第一词集中与范围词相符的词去除,得到第二词集;标注模块207用于对第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;第二确定模块208用于根据语义分析结果确定数据查询请求对应的数据呈现的维度和数据呈现的方法;输出模块209用于输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
采用该实施例提供的自然语言的数据查询意图确定装置,预设关键词过滤器,通过该关键词过滤器能够将关键词过滤器对应的范围词集合中的范围词过滤出来,因而,根据需要设置范围词结合,就可将预定的范围词进行过滤。在获取模块获取到基于自然语言的数据查询请求之后,分词模块对数据查询请求进行分词,得到一个词集,调用模块调用预设的关键词过滤器对该词集进行过滤,将该词集中与范围词相符的词过滤出来,第一确定模块根据过滤出来的词能够形成查询范围。得到查询范围后,筛除模块将词集中过滤出来的 词去除掉,标注模块对剩余的词进行语义标注,第二确定模块按照语义标注的结果确定出数据呈现的维度和数据呈现的方法。最后,输出模块以查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图,从而在进行数据查询时,能够利用统一的查询逻辑按照标准化的查询意图进行查询,提高查询效率。
优选地,该装置还包括第一补充模块,当采用各个关键词过滤器均过滤不到第一词集中与范围词相符的词时,该第一补充模块用于执行以下步骤:获取历史数据查询意图;确定与数据呈现的维度和数据呈现的方法匹配度最高的历史数据查询意图;获取匹配度最高的历史数据查询意图中的历史查询范围作为数据查询请求的查询范围。其中,输出模块在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征查询范围的文字字体颜色设置为灰色。
优选地,该装置还包括第二补充模块,当第二确定模块根据语义分析结果无法确定出数据查询请求对应的数据呈现的维度时,该第二补充模块用于执行以下步骤:获取历史数据查询意图;确定与查询范围和数据呈现的方法匹配度最高的历史数据查询意图;获取匹配度最高的历史数据查询意图中的历史数据呈现的维度作为数据查询请求的数据呈现的维度。其中,输出模块在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征数据呈现的维度的文字字体颜色设置为灰色。
优选地,该装置还包括第三补充模块,当第二确定模块根据语义分析结果无法确定出数据查询请求对应的数据呈现的方法时,该第三补充模块用于执行以下步骤:获取历史数据查询意图;确定与查询范围和数据呈现的维度匹配度最高的历史数据查询意图;获取匹配度最高的历史数据查询意图中的历史数据呈现的方法作为数据查询请求的数据呈现的方法,其中,输出模块在输出查询范围、数据呈现的维度和数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征数据呈现的方法的文字字体颜色设置为灰色。
优选地,标注模块生成的语义分析结果包括多个词-义对,词-义对包括第二词集中的一个词和词的语义。
该装置还包括第一预设模块,用于预设数据内容中各个数据呈现的维度所对应的语义,第二确定模块在根据语义分析结果确定数据查询请求对应的数据呈现的维度时,具体执行以下步骤:匹配词-义对中的语义和数据内容中各个数据呈现的维度所对应的语义;当词-义对中的语义与数据内容中的第一数据呈现的维度所对应的语义相同时,将第一数据呈现的维度作为数据查询请求对应的数据呈现的维度。
该装置还包括第二预设模块,用于预设数据内容中各个数据呈现的方法所对应的语义, 第二确定模块在根据语义分析结果确定数据查询请求对应的数据呈现的方法时,具体执行以下步骤:匹配词-义对中的语义和数据内容中各个数据呈现的方法所对应的语义;当词-义对中的语义与数据内容中的第一数据呈现的方法所对应的语义相同时,将第一数据呈现的方法作为数据查询请求对应的数据呈现的方法。
实施例3
本实施例还提供一种计算机设备,如可以执行程序的智能手机、平板电脑、笔记本电脑、台式计算机、机架式服务器、刀片式服务器、塔式服务器或机柜式服务器(包括独立的服务器,或者多个服务器所组成的服务器集群)等。如图3所示,本实施例的计算机设备20至少包括但不限于:可通过系统总线相互通信连接的存储器21、处理器22,如图3所示。需要指出的是,图3仅示出了具有组件21-22的计算机设备20,但是应理解的是,并不要求实施所有示出的组件,可以替代的实施更多或者更少的组件。
本实施例中,存储器21(即可读存储介质)包括闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等。在一些实施例中,存储器21可以是计算机设备20的内部存储单元,例如该计算机设备20的硬盘或内存。在另一些实施例中,存储器21也可以是计算机设备20的外部存储设备,例如该计算机设备20上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。当然,存储器21还可以既包括计算机设备20的内部存储单元也包括其外部存储设备。本实施例中,存储器21通常用于存储安装于计算机设备20的操作系统和各类应用软件,例如实施例2的自然语言的数据查询装置的程序代码等。此外,存储器21还可以用于暂时地存储已经输出或者将要输出的各类数据。
处理器22在一些实施例中可以是中央处理器(Central Processing Unit,CPU)、控制器、微控制器、微处理器、或其他数据处理芯片。该处理器22通常用于控制计算机设备20的总体操作。本实施例中,处理器22用于运行存储器21中存储的程序代码或者处理数据,例如自然语言的数据查询装置等。
实施例4
本实施例还提供一种计算机可读存储介质,如闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘、服务器、App应用商城等等,其上存储有计算机程序,程序被 处理器执行时实现相应功能。本实施例的计算机可读存储介质用于自然语言的数据查询装置,被处理器执行时实现实施例1的自然语言的数据查询方法。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。
以上仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。
Claims (20)
- 一种自然语言的数据查询意图确定方法,其特征在于,包括:预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
- 根据权利要求1所述的自然语言的数据查询意图确定方法,其特征在于,所述方法还包括:当采用各个所述关键词过滤器均过滤不到所述第一词集中与所述范围词相符的词时,获取历史数据查询意图;确定与所述数据呈现的维度和所述数据呈现的方法匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史查询范围作为所述数据查询请求的查询范围;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述查询范围的文字字体颜色设置为灰色。
- 根据权利要求1所述的自然语言的数据查询意图确定方法,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的维度时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的方法匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的维度作为所述数据查询请求的数据呈现的维度;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的维度的文字字体颜色设置为灰色。
- 根据权利要求1所述的自然语言的数据查询意图确定方法,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的方法时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的维度匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的方法作为所述数据查询请求的数据呈现的方法;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的方法的文字字体颜色设置为灰色。
- 根据权利要求1所述的自然语言的数据查询意图确定方法,其特征在于,所述语义分析结果包括多个词-义对,所述词-义对包括所述第二词集中的一个词和所述词的语义。
- 根据权利要求5所述的自然语言的数据查询意图确定方法,其特征在于,所述方法还包括:预设数据内容中各个数据呈现的维度所对应的语义;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度包括:匹配所述词-义对中的语义和所述数据内容中各个数据呈现的维度所对应的语义;当所述词-义对中的语义与所述数据内容中的第一数据呈现的维度所对应的语义相同时,将所述第一数据呈现的维度作为所述数据查询请求对应的数据呈现的维度。
- 根据权利要求5所述的自然语言的数据查询意图确定方法,其特征在于,所述方法还包括:预设数据内容中各个数据呈现的方法所对应的语义;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的方法包括:匹配所述词-义对中的语义和所述数据内容中各个数据呈现的方法所对应的语义;当所述词-义对中的语义与所述数据内容中的第一数据呈现的方法所对应的语义相同时,将所述第一数据呈现的方法作为所述数据查询请求对应的方法呈现的维度。
- 一种自然语言的数据查询装置,其特征在于,包括:多个关键词过滤器,其中,所述关键词过滤器用于表征查询范围,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取模块,用于获取待分析的基于自然语言的数据查询请求;分词模块,用于对所述数据查询请求进行分词,得到第一词集;调用模块,用于调用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;第一确定模块,用于根据所有过滤到的词得到所述数据查询请求的查询范围;筛除模块,用于将所述第一词集中与所述范围词相符的词去除,得到第二词集;标注模块,用于对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;第二确定模块,用于根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出模块,用于输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
- 一种计算机设备,所述计算机设备包括存储器、处理器以及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,所述处理器执行所述程序时实现自然语言的数据查询意图确定方法的以下步骤:预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
- 根据权利要求9所述的计算机设备,其特征在于,所述方法还包括:当采用各个所述关键词过滤器均过滤不到所述第一词集中与所述范围词相符的词 时,获取历史数据查询意图;确定与所述数据呈现的维度和所述数据呈现的方法匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史查询范围作为所述数据查询请求的查询范围;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述查询范围的文字字体颜色设置为灰色。
- 根据权利要求9所述的计算机设备,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的维度时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的方法匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的维度作为所述数据查询请求的数据呈现的维度;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的维度的文字字体颜色设置为灰色。
- 根据权利要求9所述的计算机设备,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的方法时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的维度匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的方法作为所述数据查询请求的数据呈现的方法;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的方法的文字字体颜色设置为灰色。
- 根据权利要求9所述的计算机设备,其特征在于,所述语义分析结果包括多个词-义对,所述词-义对包括所述第二词集中的一个词和所述词的语义。
- 根据权利要求13所述的计算机设备,其特征在于,所述方法还包括:预设数据内容中各个数据呈现的维度所对应的语义;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度包括:匹配所述词-义对中的语义和所述数据内容中各个数据呈现的维度所对应的语义;当所述词-义对中的语义与所述数据内容中的第一数据呈现的维度所对应的语义相同时,将所述第一数据呈现的维度作为所述数据查询请求对应的数据呈现的维度。
- 一种计算机可读存储介质,其上存储有计算机程序,其特征在于:所述程序被处理器执行时实现自然语言的数据查询意图确定方法的以下步骤:预设多个表征查询范围的关键词过滤器,其中,每个所述关键词过滤器对应一个范围词集合,所述范围词集合包括多个范围词;获取待分析的基于自然语言的数据查询请求;对所述数据查询请求进行分词,得到第一词集;采用各个所述关键词过滤器过滤所述第一词集中与所述范围词相符的词;根据所有过滤到的词得到所述数据查询请求的查询范围;将所述第一词集中与所述范围词相符的词去除,得到第二词集;对所述第二词集中的词根据语义知识库进行语义标注,生成语义分析结果;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度和数据呈现的方法;以及输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图。
- 根据权利要求15所述的计算机可读存储介质,其特征在于,所述方法还包括:当采用各个所述关键词过滤器均过滤不到所述第一词集中与所述范围词相符的词时,获取历史数据查询意图;确定与所述数据呈现的维度和所述数据呈现的方法匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史查询范围作为所述数据查询请求的查询范围;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述查询范围的文字字体颜色设置为灰色。
- 根据权利要求15所述的计算机可读存储介质,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的维度时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的方法匹配度最高的所述历史数据查询意 图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的维度作为所述数据查询请求的数据呈现的维度;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的维度的文字字体颜色设置为灰色。
- 根据权利要求15所述的计算机可读存储介质,其特征在于,所述方法还包括:当根据所述语义分析结果无法确定出所述数据查询请求对应的数据呈现的方法时,获取历史数据查询意图;确定与所述查询范围和所述数据呈现的维度匹配度最高的所述历史数据查询意图;获取所述匹配度最高的历史数据查询意图中的历史数据呈现的方法作为所述数据查询请求的数据呈现的方法;其中,在输出所述查询范围、所述数据呈现的维度和所述数据呈现的方法三种参数,作为数据查询请求对应的标准化数据查询意图的步骤时,将表征所述数据呈现的方法的文字字体颜色设置为灰色。
- 根据权利要求15所述的计算机可读存储介质,其特征在于,所述语义分析结果包括多个词-义对,所述词-义对包括所述第二词集中的一个词和所述词的语义。
- 根据权利要求19所述的计算机可读存储介质,其特征在于,所述方法还包括:预设数据内容中各个数据呈现的维度所对应的语义;根据所述语义分析结果确定所述数据查询请求对应的数据呈现的维度包括:匹配所述词-义对中的语义和所述数据内容中各个数据呈现的维度所对应的语义;当所述词-义对中的语义与所述数据内容中的第一数据呈现的维度所对应的语义相同时,将所述第一数据呈现的维度作为所述数据查询请求对应的数据呈现的维度。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| SG11201914037QA SG11201914037QA (en) | 2018-08-31 | 2019-01-14 | Natural-language data query intention determining method and apparatus, computer device, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811021831.1 | 2018-08-31 | ||
| CN201811021831.1A CN109344300A (zh) | 2018-08-31 | 2018-08-31 | 自然语言的数据查询意图确定方法、装置和计算机设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020042530A1 true WO2020042530A1 (zh) | 2020-03-05 |
Family
ID=65292417
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/071606 Ceased WO2020042530A1 (zh) | 2018-08-31 | 2019-01-14 | 自然语言的数据查询意图确定方法、装置和计算机设备 |
Country Status (3)
| Country | Link |
|---|---|
| CN (1) | CN109344300A (zh) |
| SG (1) | SG11201914037QA (zh) |
| WO (1) | WO2020042530A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12499141B2 (en) | 2021-09-10 | 2025-12-16 | International Business Machines Corporation | Ontology-based data visualization |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111523062B (zh) * | 2020-04-24 | 2024-02-27 | 浙江口碑网络技术有限公司 | 多维度信息展示方法及装置 |
| US11537660B2 (en) * | 2020-06-18 | 2022-12-27 | International Business Machines Corporation | Targeted partial re-enrichment of a corpus based on NLP model enhancements |
| CN112015921B (zh) * | 2020-09-15 | 2024-04-16 | 重庆广播电视大学重庆工商职业学院 | 一种基于学习辅助知识图谱的自然语言处理方法 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090307194A1 (en) * | 2005-06-03 | 2009-12-10 | Delefevre Patrick Y | Neutral sales consultant |
| CN102737049A (zh) * | 2011-04-11 | 2012-10-17 | 腾讯科技(深圳)有限公司 | 一种数据库的查询方法和系统 |
| CN107729336A (zh) * | 2016-08-11 | 2018-02-23 | 阿里巴巴集团控股有限公司 | 数据处理方法、设备及系统 |
| CN107748784A (zh) * | 2017-10-26 | 2018-03-02 | 邢加和 | 一种通过自然语言实现结构化数据搜索的方法 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005182280A (ja) * | 2003-12-17 | 2005-07-07 | Ibm Japan Ltd | 情報検索システム、検索結果加工システム及び情報検索方法並びにプログラム |
| CN103092979B (zh) * | 2013-01-31 | 2016-01-27 | 中国科学院对地观测与数字地球科学中心 | 遥感数据检索自然语言的处理方法 |
| CN104933100B (zh) * | 2015-05-28 | 2018-05-04 | 北京奇艺世纪科技有限公司 | 关键词推荐方法和装置 |
| CN107798032B (zh) * | 2017-02-17 | 2020-05-19 | 平安科技(深圳)有限公司 | 自助语音会话中的应答消息处理方法和装置 |
| CN106980689B (zh) * | 2017-03-31 | 2020-07-14 | 江苏赛睿信息科技股份有限公司 | 一种通过语音交互实现数据可视化的方法 |
-
2018
- 2018-08-31 CN CN201811021831.1A patent/CN109344300A/zh not_active Withdrawn
-
2019
- 2019-01-14 SG SG11201914037QA patent/SG11201914037QA/en unknown
- 2019-01-14 WO PCT/CN2019/071606 patent/WO2020042530A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090307194A1 (en) * | 2005-06-03 | 2009-12-10 | Delefevre Patrick Y | Neutral sales consultant |
| CN102737049A (zh) * | 2011-04-11 | 2012-10-17 | 腾讯科技(深圳)有限公司 | 一种数据库的查询方法和系统 |
| CN107729336A (zh) * | 2016-08-11 | 2018-02-23 | 阿里巴巴集团控股有限公司 | 数据处理方法、设备及系统 |
| CN107748784A (zh) * | 2017-10-26 | 2018-03-02 | 邢加和 | 一种通过自然语言实现结构化数据搜索的方法 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12499141B2 (en) | 2021-09-10 | 2025-12-16 | International Business Machines Corporation | Ontology-based data visualization |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109344300A (zh) | 2019-02-15 |
| SG11201914037QA (en) | 2020-04-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11392775B2 (en) | Semantic recognition method, electronic device, and computer-readable storage medium | |
| US9971967B2 (en) | Generating a superset of question/answer action paths based on dynamically generated type sets | |
| CN104216942B (zh) | 查询建议模板 | |
| US10698956B2 (en) | Active knowledge guidance based on deep document analysis | |
| US20160328650A1 (en) | Mining Forums for Solutions to Questions | |
| CN118981527B (zh) | 基于大模型的问答方法、装置、电子设备、存储介质、智能体和程序产品 | |
| WO2019091026A1 (zh) | 知识库文档快速检索方法、应用服务器及计算机可读存储介质 | |
| US20150278345A1 (en) | Method, apparatus, and server for acquiring recommended topic | |
| CN104657346A (zh) | 智能交互系统中的问题匹配方法和系统 | |
| WO2020042530A1 (zh) | 自然语言的数据查询意图确定方法、装置和计算机设备 | |
| US20140379719A1 (en) | System and method for tagging and searching documents | |
| US20090112845A1 (en) | System and method for language sensitive contextual searching | |
| WO2020056979A1 (zh) | 知识库搜索方法、装置及计算机可读存储介质 | |
| WO2019200700A1 (zh) | 一种公文处理的方法、装置、终端设备及存储介质 | |
| CN111401034A (zh) | 文本的语义分析方法、语义分析装置及终端 | |
| CN110609959B (zh) | 基于项目生命周期的检索方法、存储介质及电子设备 | |
| CN112183074B (zh) | 一种数据增强方法、装置、设备及介质 | |
| WO2024183711A1 (zh) | 搜索结果的展示方法、装置、电子设备和存储介质 | |
| KR20190109628A (ko) | 개인화된 기사 컨텐츠 제공 방법 및 장치 | |
| CN114911898A (zh) | 基于知识图谱的搜索方法、装置及电子设备 | |
| CN110263312B (zh) | 文章生成方法、装置、服务器和计算机可读介质 | |
| JP7045970B2 (ja) | リスク特定装置、リスク特定方法、およびプログラム | |
| CN119049700B (zh) | 推荐理由生成方法、装置、电子设备及存储介质 | |
| CN112668286A (zh) | 一种动态标签生成方法、装置、电子设备及存储介质 | |
| HK40000586A (zh) | 自然语言的数据查询意图确定方法、装置和计算机设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19855945 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 09.06.2021) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19855945 Country of ref document: EP Kind code of ref document: A1 |