WO2020233345A1 - 基于自然语言处理的数据图表生成方法和相关装置 - Google Patents

基于自然语言处理的数据图表生成方法和相关装置 Download PDF

Info

Publication number
WO2020233345A1
WO2020233345A1 PCT/CN2020/086680 CN2020086680W WO2020233345A1 WO 2020233345 A1 WO2020233345 A1 WO 2020233345A1 CN 2020086680 W CN2020086680 W CN 2020086680W WO 2020233345 A1 WO2020233345 A1 WO 2020233345A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
phrase
chart
natural language
keyword
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/086680
Other languages
English (en)
French (fr)
Inventor
刘利
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
OneConnect Smart Technology Co Ltd
Original Assignee
OneConnect Smart Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by OneConnect Smart Technology Co Ltd filed Critical OneConnect Smart Technology Co Ltd
Publication of WO2020233345A1 publication Critical patent/WO2020233345A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/36Creation of semantic tools, e.g. ontology or thesauri
    • G06F16/367Ontology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/253Grammatical analysis; Style critique
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/289Phrasal analysis, e.g. finite state techniques or chunking
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis

Definitions

  • This application relates to the field of natural language processing, and in particular to a data chart generation method and related devices based on natural language processing.
  • the embodiments of the present application provide a data chart generation method and related devices based on natural language processing (NLP) to solve the problem of complicated manual chart generation operations.
  • NLP natural language processing
  • a data chart generation method based on natural language processing including:
  • Target natural language data is natural language data related to generating a data chart
  • the keyword sequence includes at least one Data chart keywords
  • the data chart function templates in the data chart function template set are sequentially invoked and executed to generate a data chart corresponding to the target natural language data.
  • a data chart generation device based on natural language processing, including: a data acquisition module for acquiring target natural language data input by a target user, and the target natural language data is a natural language related to generating a data chart Data; analysis module for segmentation and semantic analysis of the target natural language data based on natural language processing to determine the grammatical structure features of the target natural language data and the keyword sequence corresponding to the target natural language data, so
  • the keyword sequence includes at least one data chart keyword;
  • a function template determining module is used to determine at least one data chart function template corresponding to the keyword sequence; an assembly module is used to compare the at least one data chart function template according to the grammatical structure feature
  • a data chart function template is assembled to determine the data chart function template set corresponding to the target natural language data; the chart generation module is used to sequentially call and execute the data chart function templates in the data chart function template set to generate all The data chart corresponding to the target natural language data.
  • another natural language processing-based data chart generation device which includes a processor, a memory, and an input-output interface.
  • the processor, the memory and the input-output interface are connected to each other, wherein the input-output interface is used Input or output data, the memory is used to store the application code of the data chart generation device based on natural language processing to execute the above method, and the processor is configured to execute the data chart generation method, including: obtaining the target input by the target user Natural language data, where the target natural language data is natural language data related to generating data charts; performing word segmentation and semantic analysis on the target natural language data based on natural language processing to determine the grammatical structure characteristics of the target natural language data A keyword sequence corresponding to the target natural language data, the keyword sequence includes at least one data chart keyword; at least one data chart function template corresponding to the keyword sequence is determined; The at least one data chart function template is assembled to determine the data chart function template set corresponding to the target natural language data; the data chart function templates in the data chart function template set
  • a computer storage medium stores a computer program
  • the computer program includes program instructions that, when executed by a processor, cause the processor to execute the data graph generation method
  • the method includes: obtaining target natural language data input by a target user, where the target natural language data is natural language data related to generating a data chart; and performing word segmentation and semantic analysis on the target natural language data based on natural language processing to determine the The grammatical structure feature of the target natural language data and the keyword sequence corresponding to the target natural language data, the keyword sequence including at least one data chart keyword; determining at least one data chart function template corresponding to the keyword sequence; Assemble the at least one data chart function template according to the grammatical structure feature to determine the data chart function template set corresponding to the target natural language data; call and execute the data chart function template in the data chart function template set in turn , To generate a data chart corresponding to the target natural language data.
  • the above solution has the following beneficial effects: eliminating the need for users to manually set the parameters of the chart, and so on, improving the efficiency of chart production.
  • FIG. 1 is a schematic diagram of the architecture of a communication system provided by an embodiment of the present application.
  • FIG. 2 is a schematic flowchart of a data chart generation method based on natural language processing provided by an embodiment of the present application
  • 3A is a schematic diagram of a phrase structure tree provided by an embodiment of the present application.
  • 3B is a schematic diagram of another phrase structure tree provided by an embodiment of the present application.
  • FIG. 4 is a schematic flowchart of another data chart generation method based on natural language processing provided by an embodiment of the present application
  • FIG. 5 is a schematic diagram of the composition structure of a data chart generation device based on natural language processing provided by an embodiment of the present application
  • FIG. 6 is a schematic diagram of the composition structure of another data chart generation device based on natural language processing provided by an embodiment of the present application.
  • the technical solutions of the embodiments of the present application can be applied to a communication system composed of a terminal device and a server.
  • the communication system may be as shown in FIG. 1, and the communication system 100 may include one or more terminal devices 101 and one or more servers 102.
  • the terminal device 101 is used to interact with the user, and the terminal device 10 can be used to obtain natural language data input by the target user and submit the natural language data input by the user to the server 102; the terminal device 101 can also be used to receive
  • the data chart is generated from the data, and the data chart is displayed to the user.
  • the terminal device includes but is not limited to a personal computer, a tablet computer, a mobile phone, an IPAD, etc.
  • the one or more servers 102 can form a data processing background system for providing background business support for terminal devices, such as providing business support for generating data charts for terminal devices, and the server can be used to receive user input obtained by terminal device 101 According to the natural language data of, a data chart corresponding to the natural language data is generated according to the natural language data; the server 102 may also be used to send the data chart to the terminal device 101.
  • the communication system may be a website system based on a browser/server (B/S) mode or a client and server mode, and the website system may include a website client and a website Server.
  • the website client can run on the terminal device 101 to provide services to users.
  • the website client can be a general-purpose client, and the general-purpose client can provide services for multiple website servers.
  • the client can be, for example, a browser; the website client can also be a specific client.
  • the specific client is only used to provide services for a specific website.
  • the specific client can be designed for generating data charts, for example. Client.
  • the specific client may refer to a computer client running on a computer, or may refer to an application client (application, APP) running on a mobile phone, a tablet computer, etc.
  • the website server is composed of a server 102, which is used to manage and provide resources of the website system to the website client.
  • the website server is used to provide various data to the website client so that the website client can display various pages to users.
  • the technical solutions of the embodiments of the present application can also be applied to an independent device that can generate data charts.
  • the independent device may be the aforementioned terminal device 101 or server 102, and the independent device may also be other
  • the device for generating the data chart is not limited in the embodiment of this application.
  • FIG. 2 is a schematic flow chart of a data chart generation method based on natural language processing provided by an embodiment of the present application.
  • the method can be implemented in the above-mentioned communication system 100 or an independent device that can generate data charts, as shown in FIG. As shown, the method includes the following steps:
  • S201 Obtain target natural language data input by a target user, where the target natural language data is natural language data related to generating a data chart.
  • the target natural language data may be voice data or text data.
  • the user can speak voice data to a terminal device or a server or other device that interacts with the user.
  • the voice data is a natural language
  • the voice data is the target natural language data.
  • the user can say "use the data in Table 1 as the data source, generate a histogram with the X-axis as the month and the Y-axis as the sales amount", then the "use the data in Table 1 as the data source and generate the X-axis as the The speech data corresponding to the month and the Y-axis is the histogram of sales is the natural language data related to the generated data chart, that is, the target natural language data.
  • the user may also input text data into a terminal device or server or other device that interacts with the user through text input, and the text data is the target language data.
  • the user enters the content of the demand for the chart on the display interface of the device interacting with the user.
  • the demand content is specifically "using the data in Table 1 as the data source, generate a histogram with the X axis as month and Y axis as sales ", then, the "use the data in Table 1 as the data source, generate a histogram with the X axis as the month and the Y axis as the sales" is the natural language data related to the generated data chart, that is, the target natural language data.
  • the target text information corresponding to the target language data may be generated based on the voice recognition technology, and the target text information may be Chinese text information.
  • the acquired target language data of the target user is not natural language data related to generating data charts
  • the current process is ended.
  • prompts such as "input error”, “input error”, “please input again”, “please rewrite input” and the like may also be issued to the user.
  • S202 Perform word segmentation and semantic analysis on the target natural language data based on NLP to determine the grammatical structure feature of the target natural language data and the keyword sequence corresponding to the target natural language data, the keyword sequence includes at least one data chart keyword.
  • performing word segmentation and semantic analysis on the target natural language data based on NLP to determine the grammatical structure feature of the target natural language data and the target and the keyword sequence corresponding to the target natural language data include the following steps:
  • performing word segmentation processing on the target natural language data refers to segmenting the text information corresponding to the target natural language data.
  • Word segmentation may refer to dividing a text information sequence into one or more word sequences.
  • the embodiment of the application will perform word segmentation on the text information
  • the multiple word sequences obtained after word segmentation are called multiple phrases.
  • word segmentation algorithm can be used to segment the text information corresponding to the target natural language data.
  • the word segmentation algorithm used to segment the text information corresponding to the target natural language data may include a word segmentation method based on string matching, a word segmentation method based on understanding, a word segmentation method based on statistics, etc., and are not limited to the description here.
  • the part-of-speech tagging for each phrase refers to the process of tagging each phrase with the most suitable part-of-speech, that is, the process of determining each phrase as a noun, verb, adjective, or other part of speech.
  • each phrase After performing part-of-speech tagging, each phrase has a part-of-speech tag, where the part-of-speech tag is used to identify the part of speech of the phrase.
  • the part-of-speech tag of each phrase can be any of the following: nouns, verbs, adjectives, numerals, quantifiers, pronouns, adverbs, prepositions, conjunctions, auxiliary words, interjections, and onomatopoeia.
  • nouns, verbs, adjectives, numerals, quantifiers, and pronouns are content words
  • adverbs, prepositions, conjunctions, auxiliary words, interjections and onomatopoeias are function words.
  • the text information corresponding to the target natural language data is "take the data in Table 1 as the data source, generate a histogram with the month on the X axis and the sales on the Y axis", and the phrases obtained by segmenting the text information are “ ⁇ ” , "Table”, “1”, “of”, “data”, “as”, “data source”, “generate”, “X-axis”, “as”, “month”, “and”, “Y-axis” "”, “ ⁇ ”, “Sales”, “ ⁇ ”, “Histogram”, respectively mark each phrase as part of speech, and get the part-of-speech tag of each phrase: the part-of-speech tag of "Yi” is the preposition; “table” The part-of-speech tag of "1” is a noun; the part-of-speech tag of "1” is a quantifier; the part-of-speech tag of " ⁇ ” is an auxiliary word; the part-of-speech tag of "data” is "noun"
  • the part of speech can be performed on each phrase in the phrase sequence corresponding to the target natural language data based on the hidden Markov model and combined with the Viterbi algorithm and/or the maximum entropy algorithm. Annotate to get the part-of-speech tag of each phrase.
  • phrase structure analysis may include one or more of dependency syntax analysis or semantic dependency analysis.
  • dependency syntax analysis refers to the process of revealing the syntactic structure of the language unit by analyzing the dependencies between the components of the language unit. In other words, based on dependency syntax analysis, it can identify the grammatical components of "subject, predicate, object” and "fixed adverbial complement” in a sentence. , And analyze the semantic modification relationship between each component.
  • the relationship between the components can be one of the following relationships: subject-verb (SBV), verb-object (VOB), indirect-object (IOB), Fronting-object (FOB), double (DBL), definite-zhong (attribute, ATT), adverbial (ADV), verbal-complement (complement, CMP), and coordinate (coordinate) , COO), preposition-object (POB), left adjunct (LAD), right adjunct (RAD), independent structure (IS), punctuation (WP) , Core relationship (head, HED), quantity relationship (quantity, QUN), appositive relationship (appositive, APP), analogy relationship (similarity, SIM), temporal relationship (temporal, TMP), location relationship (locative, LOC), " " ⁇ " character structure (DE), “ ⁇ ” character structure (DI), " ⁇ ” character structure (DEI), “ ⁇ ” character structure (SUO) "Ba” character structure (BA), " ⁇ ” character structure (BEI) ), conjunction (CNJ), conjunctive structure
  • the central component can be the phrase in the phrase sequence whose part of speech is the verb, and the interdependence of each phrase in the phrase sequence can be determined.
  • the phrase sequence is "to”, “table”, “1”, “of”, “data”, “ ⁇ ”, “data source”, “generated”, “X-axis”, “ ⁇ ”, “month” , "And”, “Y-axis”, “for”, “sales”, “of”, and “histogram”, you can use "generation” as the central component to determine the interdependence of each phrase.
  • the interdependence relationship between "Yi” and “ ⁇ ” is the state structure
  • the interdependence relationship between "Yi” and “data” is the prepositional relationship
  • "data” and “ ⁇ ” The dependency relationship is the Dingzhong relationship
  • the dependency relationship between " ⁇ ” and “1” is the structure of " ⁇ ”
  • the dependency relationship between "1” and “ ⁇ ” is the Dingzhong relationship
  • the dependency relationship between " ⁇ " and "data source” Is the verb-object relationship
  • the dependence relationship between "data source” and "” is the punctuation
  • the dependence relationship between "take” and “generation” is the state structure
  • the dependence relationship between "generation” and “histogram” is the verb object relationship
  • the dependence relationship between "Histogram” and “ ⁇ ” is a fixed-Chinese relationship
  • the dependence relationship between “ ⁇ ” and “Sales” is a structure of " ⁇ ”
  • the dependence relationship between "Sales” and “ ⁇ ” is a prepositional relationship.
  • semantic dependence analysis refers to the analysis of the semantic relationship between each language unit of a sentence, and the semantic relationship between each language unit is presented in a dependent structure.
  • the process of semantic dependence analysis is to determine the semantics between each language unit in the sentence The process of relationship, where the language unit can be understood as a phrase.
  • the types of semantic relations between each language unit can include: agent relationship (agent, Agt), party relationship (experiencer, Exp), affection relationship (affection, Aft), consular relationship (possessor, Poss), acceptor relationship Matter relationship (patient, Pat), guest relationship (content, Cont), product relationship (product, Prod), source relationship (Origin, Orig), involved relationship (dative, Datv), comparative role (comitative, Comp) , Belongings (Belg), Classicfication (Classicfication, Class), According (Accd), Reason (Reas), Intention (Int), Consequence (Consequence, Cons) , Mode role (manner, Mann), tool role (tool, Tool), material role (material, Malt), time role (time, Time), space role (location, Loc), process role (process, Proc), trend Role (direction, Dir), scope role (scope, Sco), quantity role (quantity, Quan), quantity array (quantity-phrase, Qp), frequency role (frequency, Freq),
  • the phrase structure analysis method can be used to analyze the phrase structure of each phrase in the phrase sequence corresponding to the natural language data to determine the phrase structure relationship between each phrase.
  • the phrase structure analysis method may include a graph-based phrase structure analysis method, a transfer-based phrase structure analysis method, etc., and are not limited to the description here.
  • phrase structure tree with each phrase as a node.
  • the phrase structure tree includes the phrase structure relationship between each node and the parent-child node relationship between each node.
  • phrase structure tree is constructed with each phrase as a node, and two phrase sequences with phrase structure relationships are used as parent nodes and child nodes, respectively, and the phrase structure relationship between the phrases in the phrase sequence is expressed in a tree structure.
  • the constructed phrase structure tree can be called a syntactic structure tree, that is, the character structure relationship between each node (ie, each phrase) in the syntactic structure tree Can be a dependency relationship.
  • the parent-child node relationship between each node in the syntactic structure tree is determined by the dependency relationship between each node.
  • the constructed phrase structure tree can be called the semantic structure tree.
  • the nodes in the semantic structure tree are the same as the nodes in the syntactic structure tree, but the two types of phrase structure trees
  • the parent-child node relationship between nodes is different, and the character structure relationship between nodes is also different.
  • the character structure relationship between each node is a semantic relationship. Therefore, the parent-child node relationship between each node is determined by the semantic relationship between each node.
  • phrase sequence is "to”, “table”, “1”, “of”, “data”, “ ⁇ ”, “data source”, “generated”, “X-axis”, “ ⁇ ”, “month” , "And”, “Y axis”, “for”, “sales”, “of”, “histogram”.
  • phrase structure analysis is a dependency syntax analysis. The result of the dependency syntax analysis on the phrase sequence is as described above.
  • the constructed phrase structure tree can be as shown in Figure 3A, and each phrase is used as a node of the phrase structure tree, where,
  • the root node "root” is the parent node of the node “Generation”, and the dependency relationship between the node “Generation” and its parent node is HED (ie core relationship);
  • the node “Generation” is the parent of the node “Yi” and the node “Histogram” Node, and the dependency relationship between the node “Y” and its parent node is ADV (that is, the structure in state),
  • the dependency relationship between the node “Histogram” and its parent node is VOB (that is, the verb-object relationship);
  • the node “bes the node” as “And the parent node of the node "data”, and the dependency relationship between the node “ ⁇ ” and its parent node is ADV (that is, the state structure), and the dependency relationship between the node "data” and its parent node is POB (that is, the inter-
  • the node "data” is the parent node of node “ ⁇ ”, and the dependency relationship between node “ ⁇ ” and its parent node is ATT (ie fixed-center relationship);
  • node " ⁇ " It is the parent node of node “1”, and the dependency relationship between node “1” and its parent node is DE (ie the structure of " ⁇ ”);
  • node "1” is the parent node of node “table”, node “table” and its parent node
  • the dependency relationship of the node is ATT (fixed-center relationship);
  • the node “histogram” is the parent node of the node “the”, and the dependency relationship between the node “the” and its parent is ATT (that is, the fixed-center relationship);
  • the node “the” is the node
  • the parent node of "Sales", the dependency between the node “Sales” and its parent node is DE (that is, the structure of the word “of”);
  • the node “Sales” is the parent node of the no
  • the dependency relationship between node "Y axis” and its parent node is ATT (ie fixed-center relationship);
  • node "Y axis” is the node "X axis” and The parent node of the node “and”, the dependency relationship between the node “X axis” and its parent node is COO (that is, the parallel relationship), the dependency relationship between the node “and” and its parent node is LAD (that is, the left attachment relationship);
  • the node "X axis” “Is the parent node of the node “ ⁇ ” and the node “Month”, the dependency between the node “ ⁇ ” and its parent is POB (i.e. preposition relationship), and the dependency between the node “Mon” and its parent is ATT (i.e. fixed in relationship).
  • the phrase structure relationship between each node in the phrase structure tree, and the parent-child node relationship between each node construct the grammatical structure characteristics of the target natural language data.
  • the phrase structure tree can be traversed, starting from the root node of the phrase structure tree (the node belonging to the uppermost layer), and gradually traversing to the nodes of the lower layer, according to the order of traversal, you can set one part of speech tag corresponding to each node
  • the index number makes each node on the phrase structure tree unique.
  • the nodes of the phrase structure tree can be traversed based on a breadth-first approach, that is, starting from the root node, search and traverse along the width of the phrase structure tree, that is, first traverse the nodes of the first layer, and then Traverse the nodes of the second layer.
  • the breadth traversal of the phrase structure tree in Figure 3A after visiting the root node "root”, the second visited node is "Generate", and the part-of-speech tag of "Generate” is v (v means verb), set its index The number is 0, and the string "v_0" is used to characterize the node;
  • the third node visited is "Yi", the part-of-speech tag of "Yi” is prep (prep stands for preposition), set its index number to 0, and use the string "Prep_0” represents the node;
  • the fourth node visited is "histogram", the part-of-speech tag of "histogram” is n (n represents noun), set its index number to 0, and use the string "n_0” to represent the Node;
  • the fifth node visited is " ⁇ ", the part-of-speech tag of " ⁇ ” is v, the node " ⁇ ” is the verb visited the second time, so set the index number of the node " ⁇ ” to 1, and
  • the nodes of the phrase structure tree can also be traversed in a depth-first manner, that is, starting from the root node, and traversing along the depth search of the phrase structure tree, that is, along the parent of the first layer For nodes, first traverse the nodes of the left subtree, and then traverse the nodes of the right subtree.
  • a deep traversal of the phrase structure tree in Figure 3A after visiting the root node "root”, the second visited node is "Generate", and the part-of-speech tag of "Generate” is v (v means verb), set its index The number is 0, and the string "v_0" is used to characterize the node;
  • the third node visited is "Yi", the part-of-speech tag of "Yi” is prep (prep stands for preposition), set its index number to 0, and use the string "Prep_0” represents the node;
  • the fourth node visited is " ⁇ ", the part-of-speech tag of " ⁇ ” is v (v means verb), the node " ⁇ " is the verb visited the second time, so set the node "
  • the index number of "is” is 1, and the node is represented by the character string "v_1”;
  • the fifth node visited is "data source”, the part-of-speech tag of "data source” is n, set its index number to 0,
  • the grammatical structure feature of the grammatical structure where the grammatical structure feature can be composed of the character string used to represent the grammatical structure corresponding to each word character. Therefore, the grammatical structure feature is a string representation, and the grammatical structure feature of this representation Can improve the subsequent indexing speed.
  • the character strings used to represent the grammatical structure corresponding to each node can be combined according to the parent-child node relationship between each node in the phrase structure tree to obtain the grammatical structure feature corresponding to the target natural language data.
  • the brackets in the grammatical structure feature are used to indicate the parent-child node relationship of the phrase structure tree. For example, traverse the nodes of the phrase structure tree based on the breadth-first approach to obtain the index number and character string corresponding to each phrase, and then use the parent-child node relationship between each node in the phrase structure tree to use each node correspondingly.
  • v_HED_0 prep_ADV_0(v_ADV_1(n_VOB_3(wp_WP_0))n_ADV_1(a_ADV_1(q_DE_0(n_ATT_5)))n_VOB_0(a_ATT_0(V_PO6_n_4) )))).
  • the phrases whose part-of-speech tags are nouns and adjectives in the phrase sequence can be determined as the target phrase; the target phrase is matched with the preset template keyword words; if the target phrase is key to the preset template The relevance of the word is greater than the relevance threshold, and the target phrase is determined to be the key word of the data chart.
  • the preset template keywords are phrases used to describe the attributes of various aspects of the data chart.
  • the preset template keywords may include phrases used to describe the shape of the data chart.
  • the preset template keywords may include bar graphs, bar graphs, dot graphs, bar graphs, scatter graphs, and area graphs. And other phrases.
  • the preset template keywords may also include phrases used to describe the basic attributes of the data chart.
  • the preset template keywords may include phrases such as X axis, Y axis, data range, and data source.
  • the preset template keywords may also include phrases used to describe the style of the data chart.
  • the preset template keywords may include phrases such as color, color, and shape.
  • a preset template keyword corresponding to a data chart function template there can be a preset template keyword corresponding to a data chart function template, or multiple preset template keywords corresponding to a data chart function template.
  • Each data chart function template can be used to implement the corresponding data chart function template.
  • matching the target phrase with the preset template keyword may refer to comparing whether the target phrase is the same as the preset template keyword, and if the target phrase is the same as the preset template keyword, then determine The relevance degree between the target phrase and the preset template keyword is greater than the relevance degree threshold, and the target phrase is determined as the data chart keyword. For example, if the preset template keyword is the X-axis and the target phrase is the X-axis, it is determined that the target phrase is the data chart keyword.
  • matching the target phrase with the preset template keywords may refer to comparing the semantics of the target phrase with the semantics of the preset template keywords. If the semantics of the target phrase are the same as If the semantics of the preset template keywords are the same or close, it is determined that the degree of relevance between the target phrase and the preset template keyword is greater than the relevance threshold, and the target phrase is determined as the data chart keyword. For example, if the preset template keyword is the data range and the target phrase is the data source, it is determined that the degree of relevance between the target phrase and the preset template keyword is greater than the relevance threshold, and the target phrase is determined as the data chart keyword.
  • both the target phrase and the preset template keyword can be used, and it is determined that the target phrase is highly related to the preset template keyword.
  • the formed keyword sequence can be ⁇ X axis, Y axis, data source, histogram ⁇ .
  • S203 Determine at least one data chart function template corresponding to the keyword sequence.
  • the data chart function is a pre-designed function module, and different data chart function templates can realize different data charting functions.
  • the preset template keywords corresponding to each data chart keyword in the keyword sequence can be determined respectively, and the preset template keywords corresponding to each data chart keyword correspond to
  • the data chart function template is determined to be at least one data chart function template corresponding to the keyword sequence.
  • the data chart keywords in the keyword sequence are X-axis, Y-axis, and data source respectively, which correspond to the preset keywords X-axis, Y-axis, data source, and histogram respectively.
  • the X-axis and Y-axis correspond to the chart.
  • Function template 1 data source corresponds to chart function template 2
  • histogram corresponds to chart function template 3.
  • S204 Assemble at least one data chart function template according to the grammatical structure feature of the target natural language data to determine a data chart function template set corresponding to the target natural language data.
  • the neighboring nodes corresponding to the keywords of each data chart can be determined according to the grammatical structure characteristics, and then according to the phrase structure relationship between the keywords of each data chart and the neighboring nodes corresponding to each data chart, it is determined that each data chart keyword has a predetermined Set the phrase structure relationship; according to the corresponding relationship between the phrase and the parameter, respectively convert the phrase that has a preset phrase structure relationship with each data chart keyword into the parameter corresponding to the chart function template corresponding to each data chart keyword; respectively; Replace the default parameters in each chart function template with the parameters corresponding to each chart function template; assemble each chart function template in order to obtain the data chart function template set corresponding to the target natural language data.
  • determining the phrase that has a preset phrase structure relationship with each data graph keyword according to the phrase structure relationship of each data graph keyword and the neighboring node corresponding to each data graph refers to finding an association relationship with each data graph keyword Phrase.
  • each data graph keyword can be used as a starting point, and the nodes on a subtree of the data graph can be traversed to determine that the part of speech in the phrase structure tree is noun or adjective and the data graph keyword Adjacent one or more nodes that are not keywords of the data graph, combine the direct or indirect relationship between the one or more nodes and the keywords of the data graph, and determine the phrase that has an association relationship with each data graph.
  • the phrase that has an association relationship with each data chart keyword can be determined starting from the data chart keyword at the deepest level.
  • the keywords of the data chart are the X-axis, Y-axis, data source, and histogram.
  • the part of speech adjacent to the X-axis is The nodes of the noun are the month and the Y-axis.
  • the month is determined to be a phrase related to the X-axis; the node whose part of speech is the noun nearest to the Y-axis is sales, X-axis and histogram.
  • the sales are determined as the phrases related to the Y-axis; the node whose part of speech is the noun adjacent to the histogram is the sales, because the sales are related to the Y-axis Relational phrase, it is determined that there is no phrase in the phrase structure tree that has an association relationship with the histogram; the node whose part of speech is noun adjacent to the data source is data, 1, table, then determine data, 1, table as the data source For the phrase that has an association relationship, further analysis of the data, 1, and the data in Table 1 can determine that the phrase that has an association relationship with the data source is the data in Table 1.
  • the corresponding grammatical structure feature in the phrase structure tree can be traversed to determine the phrase that has an association relationship with the data chart keywords, and the grammatical structure feature can be traversed outward from the innermost layer.
  • other implementation manners can also be used to find phrases that have an association relationship with each data chart keyword, which is not limited in the embodiment of the present application. After determining the phrase that has an association relationship with each data chart keyword, the phrase that has an association relationship with each data chart keyword can be converted into a parameter according to a preset conversion rule.
  • each chart function template after parameter replacement can be determined according to the execution sequence between the chart function templates and the structural characteristics of the target natural language data, and the chart function templates after parameter replacement can be assembled in order to obtain the A collection of data chart functions corresponding to the target natural language data.
  • S205 Calling and executing the data chart function templates in the data chart function template set in turn to generate a data chart corresponding to the target natural language data.
  • the chart function module for drawing the chart matching the semantics is determined, and then the phrase structure relationship between the phrases in the natural language data is determined.
  • the parameters corresponding to each chart function module and the order of each chart function template, and the chart function modules are assembled in order to obtain the chart function module set corresponding to the user's target natural language data, and the charts in the chart function module set are executed in turn
  • the functional module can generate the chart corresponding to the target natural language data, eliminating the need for users to manually set the parameters of the chart, and improving the efficiency of chart production.
  • Figure 4 is a schematic flow diagram of another data chart generation method based on natural language processing provided by an embodiment of the present application. This method can be implemented in the above-mentioned communication system 100 or on an independent device that can generate data charts, such as As shown in the figure, the method includes the following steps:
  • S301 Obtain target natural language data input by a target user, where the target natural language data is natural language data related to generating a data chart.
  • S302 Perform word segmentation and semantic analysis on the target natural language data based on NLP to determine the grammatical structure feature of the target natural language data and the keyword sequence corresponding to the target natural language data, the keyword sequence includes at least one data chart keyword.
  • S303 Determine at least one data chart function template corresponding to the keyword sequence.
  • S304 Assemble at least one data chart function template according to the grammatical structure feature of the target natural language data to determine a data chart function template set corresponding to the target natural language data.
  • S305 Calling and executing the data chart function templates in the data chart function template set in turn to generate a data chart corresponding to the target natural language data.
  • steps S301 to S305 can refer to the description of steps S201 to S205 in the embodiment corresponding to FIG. 2, and details are not described herein again.
  • the chart generation situation corresponding to the target user includes the type of data chart that has been generated for the target user, the data source of the data chart that has been generated for the target user, or the data chart that has been generated for the target user. At least one of the number.
  • the data chart generation situation of the target user from a certain historical time to the current time can be counted; for example, the data chart generation situation in the past 5 days can be used. It is also possible to count all the data charts generated for the target user. For example, if the user first generated the data chart from December 31, 2018, it is possible to count all the data charts generated for the target user from December 31, 2018 to the current time.
  • a data chart storage space can be divided for each target user. The data chart storage space is used to store the relevant information of the data chart generated by a target user.
  • the relevant information stored in the data chart storage space corresponding to the target user determines the type of data chart that has been generated for the target user, the data source of the data chart that has been generated for the target user, or the number of data charts that have been generated for the target user. At least one kind of information.
  • S307 Generate a chart generation status report for the target user according to the chart generation status corresponding to the target user.
  • the chart generation status report can also be pushed to the target user.
  • pushing the chart generation report to the target user can be directed to the user to display the chart generation report, or the content of the chart generation report is played in the form of voice, or the chart generation report is pushed to the user terminal , So that the user terminal displays the chart generation report or plays the content in the chart generation report in the form of voice.
  • the chart label corresponding to the data chart can also be generated, and the chart label and the The data chart is saved to the chart storage space corresponding to the target user.
  • the chart label is label information used to describe various attributes of the data chart.
  • the chart label may include the name of the data chart, the function of the data chart, the general description information of the content corresponding to the data chart, One or more of label information such as the type of the data chart and the color information of the data chart.
  • FIG. 5 is a schematic diagram of the composition structure of a data chart generating device based on natural language processing provided by an embodiment of the present application.
  • the device 40 includes:
  • the data acquisition module 401 is configured to acquire target natural language data input by a target user, where the target natural language data is natural language data related to generating data charts;
  • the analysis module 402 is configured to perform word segmentation and semantic analysis on the target natural language data based on natural language processing to determine the grammatical structure features of the target natural language data and the keyword sequence corresponding to the target natural language data.
  • the keyword sequence includes at least one data chart keyword;
  • the function template determining module 403 is configured to determine at least one data chart function template corresponding to the keyword sequence
  • the assembling module 404 is configured to assemble the at least one data chart function template according to the grammatical structure feature to determine a data chart function template set corresponding to the target natural language data;
  • the chart generation module 405 is configured to sequentially call and execute the data chart function templates in the data chart function template set to generate a data chart corresponding to the target natural language data.
  • the analysis module 402 is specifically configured to:
  • phrase sequence corresponding to the target natural language data, where the phrase sequence includes multiple phrases
  • phrase structure tree Constructing a phrase structure tree with each phrase as a node, the phrase structure tree including the phrase structure relationship between each node and the parent-child node relationship between each node;
  • a keyword sequence corresponding to the target natural language data is formed according to the at least one data chart keyword.
  • the analysis module 402 is specifically configured to: determine the phrase in the phrase sequence whose part of speech tags are nouns and adjectives as the target phrase according to the part of speech tag of each phrase;
  • the target phrase is a data chart keyword.
  • the assembly module 404 is specifically used for:
  • the phrase that has a preset phrase structure relationship with each data chart keyword is converted into the parameter corresponding to the chart function template corresponding to each data chart keyword;
  • the device 40 further includes:
  • the statistics module 406 is used to count the chart generation status corresponding to the target user, the chart generation status including the type of the data chart generated for the target user, the data source of the data chart generated for the target user, or At least one of the numbers of data charts that have been generated for the target user;
  • the report generation module 407 is configured to generate a chart generation status report for the target user according to the chart generation status.
  • the data chart generation device based on natural language processing analyzes the semantics of the target natural language data input by the user, determines the chart function module that matches the semantics for drawing the chart, and then according to the natural language data The phrase structure relationship between phrases, determine the parameters corresponding to each chart function module and the order of each chart function template, and assemble the chart function modules in order to obtain a set of chart function modules corresponding to the user's target natural language data, in turn By executing the chart function module in the chart function module set, the chart corresponding to the target natural language data can be generated, eliminating the need for the user to manually set the chart parameters and other links, which improves the efficiency of chart production.
  • FIG. 6 is a schematic diagram of the composition structure of another data chart generation device based on natural language processing provided by an embodiment of the present application.
  • the device 50 includes a processor 501, a memory 502 and an input and output interface 503.
  • the processor 501 is connected to the memory 502 and the input/output interface 503.
  • the processor 501 may be connected to the memory 502 and the input/output interface 503 through a bus.
  • the processor 501 is configured to support the natural language processing-based data chart generating device to perform the corresponding functions in the natural language processing-based data chart generating method described in FIGS. 2 to 4.
  • the processor 501 may be a central processing unit (central processing unit, CPU), a network processor (network processor, NP), a hardware chip, or any combination thereof.
  • the aforementioned hardware chip may be an application specific integrated circuit (appldcatdon specdfdc dntegrated cdrcudt, ASDC), a programmable logic device (programmable logdc devdce, PLD) or a combination thereof.
  • the above-mentioned PLD may be a complex programmable logic device (complex programmable logic device (CPLD), a field programmable logic gate array (fdeld-programmable gate array, FPGA), a general logic array (generdc array, logdc, GAL) or any combination thereof.
  • CPLD complex programmable logic device
  • FPGA field programmable logic gate array
  • GAL general logic array
  • the memory 502 is used to store program codes and the like.
  • the memory 502 may include volatile memory (VM), such as random access memory (RAM); the memory 502 may also include non-volatile memory (NVM), such as read-only Memory (read-only memory, ROM), flash memory (flash memory), hard disk (hard ddsk drdve, HDD) or solid state hard disk (sold-state drdve, SSD); the memory 502 may also include a combination of the foregoing types of memories.
  • VM volatile memory
  • RAM random access memory
  • NVM non-volatile memory
  • read-only Memory read-only memory
  • ROM read-only Memory
  • flash memory flash memory
  • hard disk hard ddsk drdve
  • solid state hard disk solid state hard disk
  • SSD solid state hard disk
  • the memory 502 is used to store data chart function modules, data charts, data chart keywords, etc.
  • the input and output interface 503 is used to input or output data.
  • the processor 501 may call the program code to perform the following operations:
  • Target natural language data is natural language data related to generating a data chart
  • the keyword sequence includes at least one Data chart keywords
  • the data chart function templates in the data chart function template set are sequentially invoked and executed to generate a data chart corresponding to the target natural language data.
  • the processor 501 calls the program code to perform word segmentation and semantic analysis on the target natural language data based on natural language processing, so as to determine the grammatical structure characteristics and the semantic analysis of the target natural language data.
  • the keyword sequence corresponding to the target natural language data including:
  • phrase sequence corresponding to the target natural language data, where the phrase sequence includes multiple phrases
  • phrase structure tree Constructing a phrase structure tree with each phrase as a node, the phrase structure tree including the phrase structure relationship between each node and the parent-child node relationship between each node;
  • a keyword sequence corresponding to the target natural language data is formed according to the at least one data chart keyword.
  • the processor 501 calls the program code to execute the determination of at least one phrase in the phrase sequence that matches a preset template keyword as at least one data chart keyword, including:
  • the target phrase is a data chart keyword.
  • the processor 501 calls the program code to perform assembly of the at least one data chart function template according to the grammatical structure feature to determine the data chart function corresponding to the target natural language data Template set, including:
  • the phrase that has a preset phrase structure relationship with each data chart keyword is converted into the parameter corresponding to the chart function template corresponding to each data chart keyword;
  • the processor 501 may also call the program code to perform the following operations:
  • Statistics of the chart generation situation corresponding to the target user including the type of data chart that has been generated for the target user, the data source of the data chart that has been generated for the target user, or the data source of the data chart that has been generated for the target user At least one of the number of generated data charts;
  • each operation may correspond to the corresponding description of the method embodiment shown in FIGS. 2 to 4; the processor 501 may also cooperate with the input and output interface 503 to perform other operations in the above method embodiment.
  • An embodiment of the present application also provides a computer storage medium, the computer storage medium stores a computer program, the computer program includes program instructions, and the program instructions when executed by a computer cause the computer to execute as described in the foregoing embodiment
  • the computer may be a part of the above-mentioned data chart generation device based on natural language processing.
  • the aforementioned processor 501 For example, the aforementioned processor 501.
  • the program can be stored in a computer readable storage medium. During execution, it may include the procedures of the above-mentioned method embodiments.
  • the storage medium can be a magnetic disk, an optical disk, ROM or RAM, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Animal Behavior & Ethology (AREA)
  • Machine Translation (AREA)

Abstract

基于自然语言处理的数据图表生成方法和相关装置,其中,方法包括:获取目标用户输入的目标自然语言数据,目标自然语言数据为有关于生成数据图表的自然语言数据(S201);基于NLP对目标自然语言数据进行分词与语义分析,以确定目标自然语言数据的语法结构特征和目标自然语言数据对应的关键词序列,关键词序列包括至少一个数据图表关键词(S202);确定与关键词序列对应的至少一个数据图表功能模板(S203);根据目标自然语言数据的语法结构特征对至少一个数据图表功能模板进行组装,以确定目标自然语言数据对应的数据图表功能模板集(S204);依次调用并执行数据图表功能模板集中的数据图表功能模板,以生成目标自然语言数据对应的数据图表(S205)。该方案可提高数据图表的制作效率。

Description

基于自然语言处理的数据图表生成方法和相关装置
本申请要求于2019年5月21日提交中国专利局、申请号为201910426646.9,发明名称为“基于自然语言处理的数据图表生成方法和相关装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及自然语言处理领域,尤其涉及基于自然语言处理的数据图表生成方法和相关装置。
背景技术
随着计算机技术日新月异的发展,企业信息化成为企业进步的必然趋势,人们越来越多地使用计算机进行各种数据的分析与处理,从而为企业的决策提供数据支撑。图表的主要目的是将数据,利用系统化的整理,依据不同的需求,以便于理解的方式呈现出来。图表作为信息系统中的数据展现的最重要的途径,发挥着巨大的作用。
发明人意识到目前,对于制作图表的企业人员来说,需要企业人员根据数据来源,利用生成图表软件(如Excel等),手动选择需要用于制作图表的数据和设置各种图表的参数以生成图表,操作复杂。
发明内容
本申请实施例提供基于自然语言处理(natural language processing,NLP)的数据图表生成方法和相关装置,解决目手动生成图表操作复杂的问题。
第一方面,提供一种基于自然语言处理的数据图表生成方法,包括:
获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
确定与所述关键词序列对应的至少一个数据图表功能模板;
根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
在该技术方案中,通过分析用户输入的目标自然语言数据的语义,确定与该语义相匹配的用于绘制图表的图表功能模块,然后根据自然语言数据中的词组之间的词组结构关系,确定各个图表功能模块对应的参数和各个图表功能模板的顺序,并按顺序对图表功能模块进行组装,得到与用户的目标自然语言数据对应的图表功能模块集合,依次执行该图表功能模块集合中的图表功能模块,即可生成该目标自然语言数据对应的图表,省去用户手动设置图表的参数等环节,提高了图表的制作效率。
第二方面,提供一种基于自然语言处理的数据图表生成装置,包括:数据获取模块,用于获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;分析模块,用于基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;功能模板确定模块,用于确定与所述关键词序列对应的至少一个数据图表功能模板;组装模块,用于根据所述语法 结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;图表生成模块,用于依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
第三方面,提供另一种基于自然语言处理的数据图表生成装置,包括处理器、存储器以及输入输出接口,所述处理器、存储器和输入输出接口相互连接,其中,所述输入输出接口用于输入或输出数据,所述存储器用于存储基于自然语言处理的数据图表生成装置执行上述方法的应用程序代码,所述处理器被配置用于执行数据图表生成方法,包括:获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;确定与所述关键词序列对应的至少一个数据图表功能模板;根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
第四方面,提供一种计算机存储介质,所述计算机存储介质存储有计算机程序,所述计算机程序包括程序指令,所述程序指令当被处理器执行时使所述处理器执行数据图表生成方法,包括:获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;确定与所述关键词序列对应的至少一个数据图表功能模板;根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
上述方案存在如下有益效果:省去用户手动设置图表的参数等环节,提高图表的制作效率。
附图说明
图1是本申请实施例提供的一种通信系统的架构示意图;
图2是本申请实施例提供的一种基于自然语言处理的数据图表生成方法的流程示意图;
图3A是本申请实施例提供的一种词组结构树的示意图;
图3B是本申请实施例提供的另一种词组结构树的示意图;
图4是本申请实施例提供的另一种基于自然语言处理的数据图表生成方法的流程示意图;
图5是本申请实施例提供的一种基于自然语言处理的数据图表生成装置的组成结构示意图;
图6是本申请实施例提供的另一种基于自然语言处理的数据图表生成装置的组成结构示意图。
具体实施方式
本申请实施例的技术方案可应用于由终端设备和服务器组成的通信系统中。该通信系统可以如图1所示,通信系统100可包括一个或多个终端设备101和一个或多个服务器102。其中,终端设备101用于与用户交互,终端设备10可用于获取目标用户输入的自然语言数据,并将用户输入的自然语言数据提交给服务器102;终端设备101还可用于接收服务器 根据该自然语言数据生成的数据图表,并向用户显示该数据图表。具体地,该终端设备包括但不限于为个人电脑、平板电脑、手机、IPAD等。该一个或多个服务器102可组成数据处理后台系统,用于为终端设备提供后台业务支持,如用于为终端设备提供生成数据图表的业务支持,服务器可用于接收终端设备101获取到的用户输入的自然语言数据,根据该自然语言数据生成该自然语言数据对应的数据图表;服务器102还可用于将该数据图表发送给终端设备101。
在一种可能的实现方式中,该通信系统可以为基于浏览器与服务器(browser/server,B/S)模式或基于客户端与服务器模式的网站系统,该网站系统可包括网站客户端和网站服务端。其中,网站客户端可运行在终端设备101上,用于为用户提供服务,该网站客户端可以是通用型的客户端,通用型的客户端可以为多个网站服务器提供服务,通用型的客户端例如可以为浏览器;该网站客户端也为可以特定的客户端,该特定的客户端只用于为某个特定网站提供服务,特定的客户端例如可以为专为生成数据图表而设计的客户端。具体地,该特定客户端可以是指运行在电脑上的电脑客户端,也可以是指运行在手机、平板电脑等上的应用客户端(application,APP)。网站服务端由服务器102组成,用于管理并向网站客户端提供该网站系统的资源,网站服务端用于向网站客户端提供各种数据使得该网站客户端可以向用户显示各种页面。
可选地,本申请实施例的技术方案也可应用于可生成数据图表的独立设备上,该独立设备可以为上述提到的终端设备101或服务器102,该独立的设备也可以为其他用于生成数据图表的设备,本申请实施例不做限制。
以下介绍发明实施例的技术方案。
参见图2,图2是本申请实施例提供的一种基于自然语言处理的数据图表生成方法的流程示意图,该方法可实现在上述通信系统100中或可生成数据图表的独立设备上,如图所示,该方法包括如下步骤:
S201,获取目标用户输入的目标自然语言数据,目标自然语言数据为有关于生成数据图表的自然语言数据。
这里,目标自然语言数据可以为语音数据或文本数据。在一种可能的场景中,用户可以向终端设备或服务器等与用户交互的设备说出语音数据,该语音数据为自然语言,那么,该语音数据即为目标自然语言数据。例如,用户可以说出“以表1中的数据为数据来源,生成X轴为月份并且Y轴为销售额的柱状图”,那么该“以表1中的数据为数据来源,生成X轴为月份并且Y轴为销售额的柱状图”对应的语音数据即为有关于生成数据图表的自然语言数据,即目标自然语言数据。在另一种可能的场景中,用户也可以通过文字输入的方式向终端设备或服务器等与用户交互的设备输入文本数据,该文本数据即为目标语言数据。例如,用户在与用户交互的设备的显示界面上输入对图表的需求内容,该需求内容具体为“以表1中的数据为数据来源,生成X轴为月份并且Y轴为销售额的柱状图”,那么,该“以表1中的数据为数据来源,生成X轴为月份并且Y轴为销售额的柱状图”即为有关于生成数据图表的自然语言数据,即目标自然语言数据。
其中,当目标语言数据为语音数据时,可以基于语音识别技术生成该目标语言数据对应的目标文本信息,该目标文本信息可以为中文文本信息。
可选地,当获取到的目标用户的目标语言数据不为有关于生成数据图表的自然语言数据时,则结束当前的流程。进一步地,还可以向用户发出“输入错误”、“输入有误”、“请再次输入”、“请重写输入”等提示。
S202,基于NLP对目标自然语言数据进行分词与语义分析,以确定目标自然语言数据的语法结构特征和目标自然语言数据对应的关键词序列,关键词序列包括至少一个数据图表关键词。
本申请实施例中,基于NLP对目标自然语言数据进行分词与语义分析,以确定目标自然语言数据的语法结构特征和目标和目标自然语言数据对应的关键词序列包括如下步骤:
一、对目标自然语言数据进行分词(word segmentation,WS)处理,得到目标自然语言数据对应的词组序列,目标自然语言数据对应的词组序列包括多个词组。
这里,对目标自然语言数据进行分词处理,是指对目标自然语言数据对应的文本信息进行分词,分词可以是指将文本信息序列切分成一个或多个词序列,本申请实施例将对文本信息进行分词后得到的多个词序列称之为多个词组。
具体实现中,可以通过分词算法对目标自然语言数据对应的文本信息进行分词。其中,用于对目标自然语言数据对应的文本信息进行分词的分词算法可以包括基于字符串匹配的分词方法、基于理解的分词方法、基于统计的分词方法,等等,不限于这里的描述。
二、对目标自然语言数据对应的词组序列中的每个词组进行词性标注(part-of-speech tagging,POS tagging),以得到每个词组的词性标签。
这里,对每个词组进行词性标注是指为每个词组标注一个最为合适的词性的过程,也即确定每个词组为名词、动词、形容词或其他词性的过程。进行词性标注后,每个词组具备一个词性标签,其中,词性标签用于标识词组的词性。每个词组的词性标签可以为以下任意一种:名词、动词、形容词、数词、量词、代词、副词、介词、连词、助词、叹词以及拟声词。其中,名词、动词、形容词、数词、量词、代词为实词,副词、介词、连词、助词、叹词和拟声词为虚词。
例如,目标自然语言数据对应的文本信息为“以表1中的数据为数据来源,生成X轴为月份并且Y轴为销售额的柱状图”,对文本信息分词得到的词组分别为“以”、“表”、“1”、“的”、“数据”、“为”、“数据来源”、“生成”、“X轴”、“为”、“月份”、“并且”、“Y轴”、“为”、“销售额”、“的”、“柱状图”,分别对每个词组进行词性标注,得到每个词组的词性标签为:“以”的词性标签为介词;“表”的词性标签为名词;“1”的词性标签为量词;“的”的词性标签为助词;“数据”的词性标签为“名词”;“为”的词性标签为动词;“数据来源”的词性标签为“名词”;“生成”的词性标签为动词;“X轴”的词性标签为名词;“为”的词性标签为动词;“月份”的词性标签为名词;“并且”的词性标签为连词;“Y轴”的词性标签为名词;“为”的词性标签为动词;“销售额”的词性标签为名词;“的”的词性标签为助词;“柱状图”的词性标签为名词。
具体实现中,可以基于隐马尔可夫模型(hidden Markov model)并结合维特比(Viterbi)算法和/或最大熵(maximum entropy)算法对目标自然语言数据对应的词组序列中的每个词组进行词性标注,以得到每个词组的词性标签。
三、基于词组结构分析确定目标自然语言数据对应的词组序列中的各个词组相互之间的词组结构关系。
本申请实施例中,词组结构分析可以包括:依存句法分析或语义依存分析中的一种或多种。
这里,依存句法分析是指通过分析语言单位内成分之间的依存关系揭示其句法结构的过程,换言之,基于依存句法分析可以识别句子中的“主谓宾”、“定状补”这些语法成分,并分析各成分之间的语义修饰关系。其中,各成分之间的关系可以为如下关系中的一种:主谓关系(subject-verb,SBV)、动宾关系(verb-object,VOB)、间宾关系(indirect-object,IOB)、前置宾语(fronting-object,FOB)、兼语(double,DBL)、定中关系(attribute,ATT)、状中结构(adverbial,ADV)、动补结构(complement,CMP)、并列关系(coordinate,COO)、介宾关系(preposition-object,POB)、左附加关系(left adjunct,LAD)、右附加关系(right adjunct,RAD)、独立结构(independent structure,IS)、标点(punctuation, WP)、核心关系(head,HED)、数量关系(quantity,QUN)、同位关系(appositive,APP)、比拟关系(similarity,SIM)、时间关系(temporal,TMP)、处所关系(locative,LOC)、“的”字结构(DE)、“地”字结构(DI)、“得”字结构(DEI)、“所”字结构(SUO)“把”字结构(BA)、“被”字结构(BEI)、关联词(conjunction,CNJ)、关联结构(conjunctive structure,CS)、语态结构(mood-tense,MT)、连动结构(verb-verb,VV)、双宾语(double object,DOB)、主题(topic,TOP)、独立分句(independent clause,IC)、依存分句(dependent clause,DC)、叠词关系(verb-no-verb or verb-one-verb,VNV)、一个词(YGC)。
在进行依存句法分析分析的过程中,可以以词组序列中词性为动词的词组为中心成分,分别确定词组序列中的各个词组相互之间的依存关系。例如,词组序列为“以”、“表”、“1”、“的”、“数据”、“为”、“数据来源”、“生成”、“X轴”、“为”、“月份”、“并且”、“Y轴”、“为”、“销售额”、“的”、“柱状图”,则可以以“生成”为中心成分,确定各个词组相互之间的依存关系,这里,确定的各个词组相互之间的依存关系可以为:“以”与“为”的依存关系为状中结构,“以”与“数据”的依存关系为介宾关系,“数据”与“的”依存关系为定中关系,“的”与“1”的依存关系为“的”字结构,“1”与“表”的依存关系为定中关系,“为”与“数据来源”的依存关系为动宾关系,“数据来源”与“,”的依存关系为标点,“以”与“生成”的依存关系为状中结构,“生成”与“柱状图”的依存关系为动宾关系“柱状图”与“的”的依存关系为定中关系,“的”与“销售额”的依存关系为“的”字结构,“销售额”与“为”的依存关系为介宾关系,“销售额”与“Y轴”的依存关系为定中关系,“Y轴”与“并且”的依存关系为左附加关系,“Y轴”与“X轴”的依存关系为并列关系,“X轴”与“月份”的依存关系为定中关系,“月份”与“为”的依存关系为介宾关系。
这里,语义依存分析是指分析句子各个语言单位之间的语义关联,将各个语言单位之间的语义关联以依存结构呈现,语义依存分析的过程即为确定句子中的各个语言单位之间的语义关系的过程,其中,语言单位可以理解为词组。其中,各个语言单位之间的语义关系类型可以包括:施事关系(agent,Agt)、当事关系(experiencer,Exp)、感事关系(affection,Aft)、领事关系(possessor,Poss)、受事关系(patient,Pat)、客事关系(content,Cont)、成事关系(product,Prod)、源事关系(Origin,Orig)、涉事关系(dative,Datv)、比较角色(comitative,Comp)、属事角色(belongings,Belg)、类事角色(Classicfication,Class)、依据角色(according,Accd)、缘故角色(reason,Reas)、意图角色(intention,Int)、结局角色(Consequence,Cons)、方式角色(manner,Mann)、工具角色(tool,Tool)、材料角色(material,Malt)、时间角色(time,Time)、空间角色(location,Loc)、历程角色(process,Proc)、趋向角色(direction,Dir)、范围角色(scope,Sco)、数量角色(quantity,Quan)、数量数组(quantity-phrase,Qp)、频率角色(frequency,Freq)、顺序角色(sequence,Seq)、描写角色(description,Desc)、宿主角色(host,Host)、名字修饰角色(name-modifier,Nmod)、时间修饰角色(time-modifier,Tmod)、反角色、嵌套角色、并列关系(event coordination,eCoo)、选择关系(event seletion,eSelt)、等同关系(event equivalent,eEqu)、先行关系(event precedent,ePrec)、顺承关系(event successor,eSucc),等等。
具体实现中,可以通过词组结构分析方法对自然语言数据对应的词组序列中的各个词组进行词组结构分析,以确定各个词组相互之间的词组结构关系。其中,词组结构分析方法可以包括基于图的词组结构分析方法,基于转移的词组结构分析方法,等等,不限于这里的描述。
四、以每个词组为节点构建词组结构树,词组结构树包括每个节点之间的词组结构关 系以及每个节点之间的父子节点关系。
这里,以每个词组为节点构建词组结构树是以具备词组结构关系的两个词组序列分别作为父节点和子节点,以树形结构将词组序列中的词组之间的词组结构关系表示出来。
具体地,若词组结构分析为依存句法分析,则可以将所构建的词组结构树称之为句法结构树,即该句法结构树中的每个节点(即每个词组)之间的字符结构关系可以为依存关系。句法结构树中每个节点之间的父子节点关系是由每个节点之间的依存关系所确定的。若字符结构分析为语义依存分析,则可以将所构建的词组结构树称之为语义结构树,语义结构树中的节点与句法结构树中的节点是相同的,但两种词组结构树中的节点间的父子节点关系不同,且节点间的字符结构关系也不同。在语义结构树中,每个节点之间的字符结构关系为语义关系,因此,每个节点之间的父子节点关系是由每个节点之间的语义关系所确定的。
例如,词组序列为“以”、“表”、“1”、“的”、“数据”、“为”、“数据来源”、“生成”、“X轴”、“为”、“月份”、“并且”、“Y轴”、“为”、“销售额”、“的”、“柱状图”。词组结构分析为依存句法分析,对词组序列进行依存句法分析后得到的结果如前所述,则构建的词组结构树可以如图3A所示,每个词组均作为词组结构树的节点,其中,根节点“root”是节点“生成”的父节点,且节点“生成”与其父节点的依存关系为HED(即核心关系);节点“生成”是节点“以”和节点“柱状图”的父节点,且节点“以”与其父节点的依存关系为ADV(即状中结构),节点“柱状图”与其父节点的依存关系是VOB(即动宾关系);节点“以”是节点“为”和节点“数据”的父节点,且节点“为”与其父节点的依存关系为ADV(即状中结构),节点“数据”与其父节点的依存关系为POB(即介宾关系);节点“为”是节点“数据来源”的父节点,节点“数据来源”与其父节点的依存关系为VOB(即动宾关系);节点“数据来源”是节点“,”的父节点,节点“,”与其父节点的依存关系为WP(即标点);节点“数据”是节点“的”的父节点,节点“的”与其父节点的依存关系为ATT(即定中关系);节点“的”是节点“1”的父节点,节点“1”与其父节点的依存关系为DE(即“的”字结构);节点“1”是节点“表”的父节点,节点“表”与其父节点的依存关系为ATT(定中关系);节点“柱状图”是节点“的”的父节点,节点“的”与其父节点的依存关系为ATT(即定中关系);节点“的”是节点“销售额”父节点,节点“销售额”与其父节点的依存关系为DE(即“的”字结构);节点“销售额”是节点“为”和节点“Y轴”的父节点,节点“为”与其父节点的依存关系为POB(即介宾关系),节点“Y轴”与其父节点的依存关系为ATT(即定中关系);节点“Y轴”为节点“X轴”和节点“并且”的父节点,节点“X轴”与其父节点的依存关系为COO(即并列关系),节点“并且”与其父节点的依存关系为LAD(即左附加关系);节点“X轴”为节点“为”和节点“月份”的父节点,节点“为”与其父节点的依存关系为POB(即介宾关系),节点“月份”与其父节点的依存关系为ATT(即定中关系)。
通过构建词组结构树,可明确清楚地获知词组序列中的各个词组之间的关联关系。
五、根据每个词组的标签、词组结构树中每个节点之间的词组结构关系和每个节点之间的父子节点关系,构建目标自然语言数据的语法结构特征。
具体地,可以对词组结构树进行遍历,从词组结构树的根节点(属于最上层的节点)开始,逐渐往下层的节点遍历,按照遍历的顺序可以为每个节点分别对应的词性标签设置一个索引号,使得词组结构树上的每一个节点都是唯一的。
在一种可能的实现方式中,可以基于广度优先的方式对词组结构树的节点进行遍历,即从根节点开始,沿着词组结构树的宽度搜索遍历,即先遍历第一层的节点,再遍历第二层的节点。例如,对图3A的词组结构树进行广度遍历,访问根节点“root”后,第二个访 问到的节点为“生成”,“生成”的词性标签为v(v表示动词),设置其索引号为0,并用字符串“v_0”表征该节点;第三个访问到的节点为“以”,“以”的词性标签为prep(prep表示介词),设置其索引号为0,并用字符串“prep_0”表征该节点;第四个访问到的节点为“柱状图”,“柱状图”的词性标签为n(n表示名词),设置其索引号为0,并用字符串“n_0”表征该节点;第五个访问到的节点为“为”,“为”的词性标签为v,节点“为”是第二次访问到的动词,所以设置节点“为”的索引号为1,并用字符串“v_1”表征该节点;以此类推可得到词组结构树的所有节点对应的字符串,利用各个节点对应的字符串对图3A所示的词组结构树中的节点进行替换,可得到如图3B所示的用字符串表征节点的词组结构树。
在另一种可能的实现方式中,也可以基于深度优先的方式对词组结构树的节点进行遍历,即从根节点开始,沿着词组结构树的深度搜索遍历,即沿着第一层的父节点,先遍历左子树的节点,再遍历右子树的节点。例如,对图3A的词组结构树进行深度遍历,访问根节点“root”后,第二个访问到的节点为“生成”,“生成”的词性标签为v(v表示动词),设置其索引号为0,并用字符串“v_0”表征该节点;第三个访问到的节点为“以”,“以”的词性标签为prep(prep表示介词),设置其索引号为0,并用字符串“prep_0”表征该节点;第四个访问到的节点为“为”,“为”的词性标签为v(v表示动词),节点“为”是第二次访问到的动词,所以设置节点“为”的索引号为1,并用字符串“v_1”表征该节点;第五个访问到的节点为“数据来源”,“数据来源”的词性标签为n,设置其索引号为0,并用字符串“n_0”表征该节点;以此类推可得到词组结构树的所有节点对应的字符串。
在得到各个词组对应的索引号后,可以根据各个词组词性标签、词组结构树中的多个词组之间的依存关系和父子节点关系、每个词组分别对应的索引号,构建目标自然语言数据对应的语法结构特征,其中,该语法结构特征可以由每个词字符对应的用于表示语法结构的字符串组成,因此,语法结构特征是一种字符串表示形式,这种表示形式的语法结构特征可以提高后续的索引速度。具体地,可以根据该词组结构树中的每个节点之间的父子节点关系对每个节点对应的用于表示语法结构的字符串进行组合,得到目标自然语言数据对应的语法结构特征。其中,语法结构特征中的括号用于表示词组结构树的父子节点关系。例如,基于广度优先的方式对词组结构树的节点进行遍历得到各个词组对应的索引号和字符串,则根据该词组结构树中的每个节点之间的父子节点关系对每个节点对应的用于表示语法结构的字符串进行组合的语法结构特征为:v_HED_0(prep_ADV_0(v_ADV_1(n_VOB_3(wp_WP_0))n_ADV_1(a_ADV_1(q_DE_0(n_ATT_5))))n_VOB_0(a_ATT_0(n_DE_4(v_POB_2n_ATT_4(n_ADV_6(v_POB_3n_ATT_7)con_LAD_0)))))。
六、将词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词。
其中,可以根据每个词组的词性标签将词组序列中词性标签为名词和形容词的词组确定为目标词组;将目标词组与预设模板关键词词进行关联度匹配;如果目标词组与预设模板关键词的关联度大于关联度阈值,则确定目标词组为数据图表关键词。
这里,预设模板关键词为用于描述数据图表的各方面的属性的词组。预设模板关键词可以有多个。具体地,预设模板关键词可包括用于描述数据图表的形态的词组,例如,预设模板关键词可以包括柱状图、条形图、点状图、柱形图、散点图、面积图等词组。预设模板关键词还可包括用于描述数据图表的基本属性的词组,例如,预设模板关键词可包括X轴、Y轴、数据范围、数据来源等词组。预设模板关键词还可包括用于描述数据图表的样式的词组,例如,预设模板关键词可以包括颜色、色彩、形状等词组。其中,可以是一个预设模板关键词对应一个数据图表功能模板,也可以是多个预设模板关键词对应一个数 据图表功能模板,每个数据图表功能模板可用于实现该数据图表功能模板对应的预设模板关键词所对应的绘图功能。
在一种可能的实现方式中,将目标词组与预设模板关键词进行关联度匹配可以是指比较目标词组是否与预设模板关键词相同,如果目标词组与预设模板关键词相同,则确定目标词组与预设模板关键词的关联度大于关联度阈值,进而确定目标词组为数据图表关键词。例如,预设模板关键词为X轴,目标词组为X轴,则确定该目标词组为数据图表关键词。
在另一种可能的实现方式中,将目标词组与预设模板关键词进行关联度匹配可以是指比较目标词组的语义与预设模板关键词的语义是否相同或接近,如果目标词组的语义与预设模板关键词的语义相同或接近,则确定目标词组与预设模板关键词的关联度大于关联度阈值,进而确定目标词组为数据图表关键词。例如,预设模板关键词为数据范围,目标词组为数据来源,则确定该目标词组与预设模板关键词的关联度大于关联度阈值,进而确定目标词组为数据图表关键词。
在又一种可能的实现方式中,还可以联网查询目标词组与预设模板关键词在各种语境中的使用情况,根据目标词组与预设模板关键词在各种语境中的使用情况确定目标词组与预设模板关键词的关联度,进而确定目标词组与预设模板关键词的关联度是否大于关联度阈值。例如,在多个语境中,既可以使用该目标词组,又可以使用该预设模板关键词,则确定目标词组与预设模板关键词的关联度较高。
以下举例来对确定数据图表关键词的过程进行说明。例如,词组序列中分别包括的词组为“以”、“表”、“1”、“的”、“数据”、“为”、“数据来源”、“生成”、“X轴”、“为”、“月份”、“并且”、“Y轴”、“为”、“销售额”、“的”、“柱状图”,预设模板关键词包括数据来源、X轴、Y轴,则确定数据图表关键词的过程为:首先,根据各个词组的词性标签可确定词性为名词和形容词的词组为“数据”、“表”、“数据来源”、“X轴”、“月份”、“Y轴”、“销售额”以及“柱状图”,分别将这些词组与预设模板关键词进行关联度匹配,其中,“X轴”、“Y轴”以及“数据来源”与预设模板关键词相同,则确定“X轴”、“Y轴”、“数据来源”以及“柱状图”为数据图表关键字。
七、根据至少一个数据图表关键词形成目标自然语言数据对应的关键词序列。
例如,确定“X轴”、“Y轴”、“数据来源”以及“柱状图”为数据图表关键字,则形成的关键词序列可以为{X轴,Y轴,数据来源,柱状图}。
S203,确定与关键词序列对应的至少一个数据图表功能模板。
这里,数据图表功能是预先设计好的功能模块,不同的数据图表功能模板可实现不同的绘制数据图表的功能。在通过步骤S202确定了各个数据图表关键词后,可以分别确定关键词序列中的各个数据图表关键词对应的预设模板关键词,根据各个数据图表关键词对应的预设模板关键词所对应的数据图表功能模板确定为与关键词序列对应的至少一个数据图表功能模板。
例如,关键词序列中的数据图表关键词分别为X轴、Y轴和数据来源,其分别对应预设关键字X轴,Y轴、数据来源和柱状图,其中,X轴和Y轴对应图表功能模板1,数据来源对应图表功能模板2,柱状图对应图表功能模板3,则确定数据图表功能模板1、数据图表功能模板2以及数据图表功能模板3为关键词序列对应的至少一个数据图表功能模板。
S204,根据目标自然语言数据的语法结构特征对至少一个数据图表功能模板进行组装,以确定目标自然语言数据对应的数据图表功能模板集。
具体地,可以根据语法结构特征确定各个数据图表关键词对应的邻近节点,然后根据各个数据图表关键词与各个数据图表对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;根据词组与参数的对应关系分别将与所述 各个数据图表关键词具有预设词组结构关系的词组转换为各个数据图表关键词对应的图表功能模板所对应的参数;分别利用各个图表功能模板所对应的参数替换各个图表功能模板中的默认参数;按顺序组装各个图表功能模板,得到目标自然语言数据对应的数据图表功能模板集。
这里,根据各个数据图表关键词与各个数据图表对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组是指找到与各个数据图表关键词具备关联关系的词组。在一种可能的实现方式中,可以各个数据图表关键词为起点,遍历与数据图表在一个子树上的节点,确定在该词组结构树中的词性为名词或形容词的与该数据图表关键词邻近的并且不为数据图表关键词的一个或多个节点,结合该一个或多个节点与数据图表关键词之间的直接或间接的关系,确定与各个数据图表具备关联关系的词组。其中,在数据图表关键词有多个的情况下,可以以处于最深层的数据图表关键词开始确定与各个数据图表关键词具备关联关系的词组。以图3A的词组结构树为例,根据前述可知,数据图表关键词为X轴、Y轴、数据来源以及柱状图,则根据图3A所示的词组结构树可知,与X轴邻近的词性为名词的节点为月份和Y轴,由于Y轴为数据图表关键词,则确定月份为与X轴具备关联关系的词组;与Y轴最近的词性为名词的节点为销售、X轴和柱状图,由于X轴和柱状图为数据图表关键词,则将销售额确定为与Y轴具备关联关系的词组;与柱状图邻近的词性为名词的节点为销售额,由于销售额为与Y轴具备关联关系的词组,则确定该词组结构树中没有与该柱状图具备关联关系的词组;与数据来源邻近的词性为名词的节点为数据、1、表,则确定数据、1、表为与数据来源具备关联关系的词组,对数据、1、表1的数据进行进一步分析可确定与数据来源具备关联关系的词组为表1的数据。具体实现中,可通过遍历该词组结构树中对应语法结构特征以确定与数据图表关键词具备关联关系的词组,可以从该语法结构特征的最内层开始向外遍历。可选地,也可以通过其他的实现方式找到与各个数据图表关键词具备关联关系的词组,本申请实施例不做限制。在确定与各个数据图表关键词具备关联关系的词组后,可根据预设的转换规则将与各个数据图表关键词具备关联关系的词组转换为参数。
这里,可以根据各个图表功能模板之间的执行顺序以及目标自然语言数据的结构特征,确定进行参数替换后的各个图表功能模板的顺序,按顺序组装进行参数替换后的各个图表功能模板,得到该目标自然语言数据对应的数据图表功能集合。
S205,依次调用并执行数据图表功能模板集中的数据图表功能模板,以生成目标自然语言数据对应的数据图表。
本申请实施例中,通过分析用户输入的目标自然语言数据的语义,确定与该语义相匹配的用于绘制图表的图表功能模块,然后根据自然语言数据中的词组之间的词组结构关系,确定各个图表功能模块对应的参数和各个图表功能模板的顺序,并按顺序对图表功能模块进行组装,得到与用户的目标自然语言数据对应的图表功能模块集合,依次执行该图表功能模块集合中的图表功能模块,即可生成该目标自然语言数据对应的图表,省去用户手动设置图表的参数等环节,提高了图表的制作效率。
在一些可能的情况中,在根据用户的目标自然语言数据生成数据图表后,还可以统计当前已经为用户生成的图表的情况,并向用户展示。参见图4,图4是本申请实施例提供的另一种基于自然语言处理的数据图表生成方法的流程示意图,该方法可实现在上述通信系统100中或可生成数据图表的独立设备上,如图所示,该方法包括如下步骤:
S301,获取目标用户输入的目标自然语言数据,目标自然语言数据为有关于生成数据图表的自然语言数据。
S302,基于NLP对目标自然语言数据进行分词与语义分析,以确定目标自然语言数据的语法结构特征和目标自然语言数据对应的关键词序列,关键词序列包括至少一个数据 图表关键词。
S303,确定与关键词序列对应的至少一个数据图表功能模板。
S304,根据目标自然语言数据的语法结构特征对至少一个数据图表功能模板进行组装,以确定目标自然语言数据对应的数据图表功能模板集。
S305,依次调用并执行数据图表功能模板集中的数据图表功能模板,以生成目标自然语言数据对应的数据图表。
这里,步骤S301~S305的具体实现方式可参考前述图2对应的实施例中步骤S201~S205的描述,此处不再赘述。
S306,统计目标用户对应的图表生成情况,目标用户对应的图表生成情况包括已经为目标用户生成的数据图表的种类、已经为目标用户生成的数据图表的数据来源或已经为目标用户生成的数据图表的数量中的至少一种。
具体地,可以统计该目标用户从历史的某个时间至当前时间这一段时间内的数据图表的生成情况;例如,可以通过过去5天内的数据图表的生成情况。也可以统计为该目标用户生成的所有的数据图表的情况。例如,用户第一次生成数据图表的时间是从2018年12月31日,则可以统计从2018年12月31日至当前为该目标用户生成的所有数据图表的情况。具体实现中,可以为每个目标用户划分一个数据图表存储空间,该数据图表存储空间用于存储某一目标用户生成的数据图表的有关信息,在统计目标用户对应的图表生成情况时,可根据该目标用户对应的数据图表存储空间中存储的有关信息确定已经为目标用户生成的数据图表的种类、已经为目标用户生成的数据图表的数据来源或已经为目标用户生成的数据图表的数量中的至少一种信息。
S307,根据目标用户对应的图表生成情况为目标用户生成图表生成情况报表。
这里,在根据目标用户对应的图表生成情况为目标用户生成图表生成情况报表后,还可以将该图表生成情况报表推送给目标用户。其中,将图表生成情况报表推送给目标用户可以是指向用户显示该图表生成情况报表,或者,将该图表生成情况报表中的内容以语音的形式播放,或者,将该图表生成情况推送给用户终端,以使该用户终端显示该图表生成情况报表或以语音的形式播放该图表生成情况报表中的内容。
本申请实施例中,在根据用户输入的目标自然语言数据生成该目标自然语言数据对应的数据图表后,对用户生成图表的情况进行统计和分析并生成统计报表,可使用户了解自己的图表生成情况。
可选地,在依次调用并执行数据图表功能模板集中的数据图表功能模板,以生成该目标自然语言数据对应的数据图表之后,还可以生成该数据图表对应的图表标签,并将图表标签和该数据图表保存至目标用户对应的图表存储空间。其中,图表标签是用于对该数据图表的各种属性进行描述的标签信息,该图表标签可包括该数据图表的名称、该数据图表的作用、该数据图表对应的内容的概括性描述信息、该数据图表的类型、该数据图表的色彩信息等标签信息的一种或多种。通过为数据图表生成图表标签并保存,在后续查找时可直接利用图表标签查找数据图表,加快了查找的效率。
上面介绍了发明实施例的方法,下面介绍发明实施例的装置。
参见图5,图5是本申请实施例提供的一种基于自然语言处理的数据图表生成装置的组成结构示意图,该装置40包括:
数据获取模块401,用于获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
分析模块402,用于基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
功能模板确定模块403,用于确定与所述关键词序列对应的至少一个数据图表功能模板;
组装模块404,用于根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
图表生成模块405,用于依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
在一种可能的设计中,所述分析模块402具体用于:
对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
在一种可能的设计中,所述分析模块402具体用于:根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定为目标词组;
将所述目标词组与所述预设模板关键词进行关联度匹配;
如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
在一种可能的设计中,所述组装模块404具体用于:
根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能模板集。
在一种可能的设计中,所述装置40还包括:
统计模块406,用于统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
报表生成模块407,用于根据所述图表生成情况为所述目标用户生成图表生成情况报表。
需要说明的是,图5对应的实施例中未提及的内容可参见方法实施例的描述,这里不再赘述。
本申请实施例中,基于自然语言处理的数据图表生成装置通过分析用户输入的目标自然语言数据的语义,确定与该语义相匹配的用于绘制图表的图表功能模块,然后根据自然语言数据中的词组之间的词组结构关系,确定各个图表功能模块对应的参数和各个图表功 能模板的顺序,并按顺序对图表功能模块进行组装,得到与用户的目标自然语言数据对应的图表功能模块集合,依次执行该图表功能模块集合中的图表功能模块,即可生成该目标自然语言数据对应的图表,省去用户手动设置图表的参数等环节,提高了图表的制作效率。
参见图6,图6是本申请实施例提供的另一种基于自然语言处理的数据图表生成装置的组成结构示意图,该装置50包括处理器501、存储器502以及输入输出接口503。处理器501连接到存储器502和输入输出接口503,例如处理器501可以通过总线连接到存储器502和输入输出接口503。
处理器501被配置为支持基于自然语言处理的数据图表生成装置执行图2-图4所述的基于自然语言处理的数据图表生成方法中相应的功能。该处理器501可以是中央处理器(central processdng undt,CPU),网络处理器(network processor,NP),硬件芯片或者其任意组合。上述硬件芯片可以是专用集成电路(appldcatdon specdfdc dntegrated cdrcudt,ASDC),可编程逻辑器件(programmable logdc devdce,PLD)或其组合。上述PLD可以是复杂可编程逻辑器件(complex programmable logdc devdce,CPLD),现场可编程逻辑门阵列(fdeld-programmable gate array,FPGA),通用阵列逻辑(generdc array logdc,GAL)或其任意组合。
存储器502存储器用于存储程序代码等。存储器502可以包括易失性存储器(volatdle memory,VM),例如随机存取存储器(random access memory,RAM);存储器502也可以包括非易失性存储器(non-volatdle memory,NVM),例如只读存储器(read-only memory,ROM),快闪存储器(flash memory),硬盘(hard ddsk drdve,HDD)或固态硬盘(soldd-state drdve,SSD);存储器502还可以包括上述种类的存储器的组合。本申请实施例中,存储器502用于存储数据图表功能模块、数据图表、数据图表关键词等。
所述输入输出接口503用于输入或输出数据。
处理器501可以调用所述程序代码以执行以下操作:
获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
确定与所述关键词序列对应的至少一个数据图表功能模板;
根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
在一种可能的实施方式中,处理器501调用所述程序代码以执行基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,包括:
对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
在一种可能的实施方式中,处理器501调用所述程序代码以执行将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词,包括:
根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定为目标词组;
将所述目标词组与所述预设模板关键词进行关联度匹配;
如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
在一种可能的实现方式中,处理器501调用所述程序代码以执行根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集,包括:
根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能模板集。
在一种可能的实现方式中,处理器501还可以调用所述程序代码以执行以下操作:
统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
根据所述图表生成情况为所述目标用户生成图表生成情况报表。
需要说明的是,各个操作的实现可以对应参照图2-图4所示的方法实施例的相应描述;所述处理器501还可以与输入输出接口503配合执行上述方法实施例中的其他操作。
本申请实施例还提供一种计算机存储介质,所述计算机存储介质存储有计算机程序,所述计算机程序包括程序指令,所述程序指令当被计算机执行时使所述计算机执行如前述实施例所述的方法,所述计算机可以为上述提到的基于自然语言处理的数据图表生成装置的一部分。例如为上述的处理器501。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,所述的程序可存储于一计算机可读取存储介质中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,所述的存储介质可为磁碟、光盘、ROM或RAM等。

Claims (20)

  1. 一种基于自然语言处理的数据图表生成方法,其中,包括:
    获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
    基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
    确定与所述关键词序列对应的至少一个数据图表功能模板;
    根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
    依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
  2. 根据权利要求1所述的方法,其中,所述基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,包括:
    对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
    对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
    基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
    以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
    根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
    将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
    根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
  3. 根据权利要求2所述的方法,其中,所述将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词,包括:
    根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定为目标词组;
    将所述目标词组与所述预设模板关键词进行关联度匹配;
    如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
  4. 根据权利要求2所述的方法,其中,所述根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集,包括:
    根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
    根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
    根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
    分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
    按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能 模板集。
  5. 根据权利要求1-4任一项所述的方法,其中,所述依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表之后,还包括:
    统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
    根据所述图表生成情况为所述目标用户生成图表生成情况报表。
  6. 一种基于自然语言处理的数据图表生成装置,其中,包括:
    数据获取模块,用于获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
    分析模块,用于基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
    功能模板确定模块,用于确定与所述关键词序列对应的至少一个数据图表功能模板;
    组装模块,用于根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
    图表生成模块,用于依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
  7. 根据权利要求6所述的装置,其中,所述分析模块具体用于:
    对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
    对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
    基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
    以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
    根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
    将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
    根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
  8. 根据权利要求7所述的装置,其中,所述组装模块具体用于:
    根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
    根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
    根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
    分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
    按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能模板集。
  9. 根据权利要求7所述的装置,其中,所述分析模块在将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词时,具体用于:
    根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定 为目标词组;
    将所述目标词组与所述预设模板关键词进行关联度匹配;
    如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
  10. 根据权利要求6-9任一项所述的装置,其中,所述装置还包括:
    统计模块,用于统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
    报表生成模块,用于根据所述图表生成情况为所述目标用户生成图表生成情况报表。
  11. 一种基于自然语言处理的数据图表生成装置,包括处理器、存储器以及输入输出接口,所述处理器、存储器和输入输出接口相互连接,其中,所述输入输出接口用于输入或输出数据,所述存储器用于存储程序代码,所述处理器用于调用所述程序代码,执行一种数据图表生成的方法,其中,包括:
    获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
    基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
    确定与所述关键词序列对应的至少一个数据图表功能模板;
    根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
    依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
  12. 根据权利要求11所述的数据图表生成装置,其中,所述基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,包括:
    对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
    对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
    基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
    以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
    根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
    将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
    根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
  13. 根据权利要求12所述的数据图表生成装置,其中,所述将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词,包括:
    根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定为目标词组;
    将所述目标词组与所述预设模板关键词进行关联度匹配;
    如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
  14. 根据权利要求12所述的数据图表生成装置,其中,所述根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集,包括:
    根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
    根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
    根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
    分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
    按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能模板集。
  15. 根据权利要求11-14任一项所述的数据图表生成装置,其中,所述依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表之后,还包括:
    统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
    根据所述图表生成情况为所述目标用户生成图表生成情况报表。
  16. 一种计算机存储介质,其中,所述计算机存储介质存储有计算机程序,所述计算机程序包括程序指令,所述程序指令当被处理器执行时使所述处理器执行一种数据图表生成的方法,其中,包括:获取目标用户输入的目标自然语言数据,所述目标自然语言数据为有关于生成数据图表的自然语言数据;
    基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,所述关键词序列包括至少一个数据图表关键词;
    确定与所述关键词序列对应的至少一个数据图表功能模板;
    根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集;
    依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表。
  17. 根据权利要求16所述的存储介质,其中,所述基于自然语言处理对所述目标自然语言数据进行分词与语义分析,以确定所述目标自然语言数据的语法结构特征和所述目标自然语言数据对应的关键词序列,包括:
    对所述目标自然语言数据进行分词处理,得到所述目标自然语言数据对应的词组序列,所述词组序列包括多个词组;
    对所述词组序列中的每个词组进行词性标注,以得到所述每个词组的词性标签;
    基于词组结构分析确定所述词组序列中的各个词组相互之间的词组结构关系;
    以每个词组为节点构建词组结构树,所述词组结构树包括每个节点之间的词组结构关系以及每个节点之间的父子节点关系;
    根据所述每个词组的词性标签、所述词组结构树中每个节点之间的词组结构关系和所述每个节点之间的父子节点关系,构建所述目标自然语言数据的语法结构特征;
    将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词;
    根据所述至少一个数据图表关键词形成所述目标自然语言数据对应的关键词序列。
  18. 根据权利要求17所述的存储介质,其中,所述将所述词组序列中与预设模板关键词匹配的至少一个词组确定为至少一个数据图表关键词,包括:
    根据所述每个词组的词性标签将所述词组序列中词性标签为名词和形容词的词组确定为目标词组;
    将所述目标词组与所述预设模板关键词进行关联度匹配;
    如果所述目标词组与所述预设模板关键词的关联度大于关联度阈值,则确定所述目标词组为数据图表关键词。
  19. 根据权利要求17所述的存储介质,其中,所述根据所述语法结构特征对所述至少一个数据图表功能模板进行组装,以确定所述目标自然语言数据对应的数据图表功能模板集,包括:
    根据所述语法结构特征分别确定所述关键词序列中的各个数据图表关键词对应的邻近节点;
    根据所述各个数据图表关键词与所述各个数据图表关键词对应的邻近节点的词组结构关系分别确定与所述各个数据图表关键词具有预设词组结构关系的词组;
    根据词组与参数的对应关系分别将与所述各个数据图表关键词具有预设词组结构关系的词组转化为各个数据图表关键词对应的图表功能模板所对应的参数;
    分别利用各个图表功能模板所对应的参数替换所述图表功能模板中的默认参数;
    按顺序组装所述各个图表功能模板,得到所述目标自然语言数据对应的数据图表功能模板集。
  20. 根据权利要求16-19任一项所述的存储介质,其中,所述依次调用并执行所述数据图表功能模板集中的数据图表功能模板,以生成所述目标自然语言数据对应的数据图表之后,还包括:
    统计所述目标用户对应的图表生成情况,所述图表生成情况包括已经为所述目标用户生成的数据图表的种类、已经为所述目标用户生成的数据图表的数据来源或已经为所述目标用户生成的数据图表的数量中的至少一种;
    根据所述图表生成情况为所述目标用户生成图表生成情况报表。
PCT/CN2020/086680 2019-05-21 2020-04-24 基于自然语言处理的数据图表生成方法和相关装置 Ceased WO2020233345A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910426646.9 2019-05-21
CN201910426646.9A CN110222194B (zh) 2019-05-21 2019-05-21 基于自然语言处理的数据图表生成方法和相关装置

Publications (1)

Publication Number Publication Date
WO2020233345A1 true WO2020233345A1 (zh) 2020-11-26

Family

ID=67821724

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/086680 Ceased WO2020233345A1 (zh) 2019-05-21 2020-04-24 基于自然语言处理的数据图表生成方法和相关装置

Country Status (2)

Country Link
CN (1) CN110222194B (zh)
WO (1) WO2020233345A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117785986A (zh) * 2023-11-17 2024-03-29 钉钉(中国)信息技术有限公司 图表模板生成、可视化图表生成、服务组件更新方法

Families Citing this family (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110222194B (zh) * 2019-05-21 2022-10-04 深圳壹账通智能科技有限公司 基于自然语言处理的数据图表生成方法和相关装置
CN112579066A (zh) * 2019-09-30 2021-03-30 北京国双科技有限公司 图表展示方法、装置、存储介质及设备
CN110837545A (zh) * 2019-11-13 2020-02-25 贵州医渡云技术有限公司 交互式数据分析方法、装置、介质及电子设备
CN114327418B (zh) * 2020-09-30 2025-06-24 微软技术许可有限责任公司 电子设备和计算机实现的方法
CN113486230A (zh) * 2021-07-28 2021-10-08 黄泽恒 标签化报文模板生成方法
US12045825B2 (en) 2021-10-01 2024-07-23 International Business Machines Corporation Linguistic transformation based relationship discovery for transaction validation
CN114625754A (zh) * 2022-03-08 2022-06-14 亚信科技(南京)有限公司 语句查询方法、装置、电子设备及计算机可读存储介质
CN114742032A (zh) * 2022-04-24 2022-07-12 广州亚信技术有限公司 交互式数据分析方法、装置、设备、介质及程序产品
CN114579111B (zh) * 2022-05-09 2022-07-29 中国联合重型燃气轮机技术有限公司 燃气轮机保护系统的代码生成方法、装置及电子设备
CN116126312B (zh) * 2022-12-30 2026-03-17 北京大学 一种基于自然语言构建可视化图表的方法及系统

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20070106499A1 (en) * 2005-08-09 2007-05-10 Kathleen Dahlgren Natural language search system
CN103631882A (zh) * 2013-11-14 2014-03-12 北京邮电大学 基于图挖掘技术的语义化业务生成系统和方法
CN106155999A (zh) * 2015-04-09 2016-11-23 科大讯飞股份有限公司 自然语言语义理解方法及系统
CN106649223A (zh) * 2016-12-23 2017-05-10 北京文因互联科技有限公司 基于自然语言处理的金融报告自动生成方法
CN107861933A (zh) * 2017-11-29 2018-03-30 北京百度网讯科技有限公司 生成运维报表的方法和装置
CN110222194A (zh) * 2019-05-21 2019-09-10 深圳壹账通智能科技有限公司 基于自然语言处理的数据图表生成方法和相关装置

Family Cites Families (17)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104915340B (zh) * 2014-03-10 2019-09-10 北京大学 自然语言问答方法及装置
CN104484353A (zh) * 2014-11-28 2015-04-01 华为技术有限公司 数据图形化方法、装置及数据库服务器
US11030406B2 (en) * 2015-01-27 2021-06-08 Verint Systems Ltd. Ontology expansion using entity-association rules and abstract relations
US9754051B2 (en) * 2015-02-25 2017-09-05 International Business Machines Corporation Suggesting a message to user to post on a social network based on prior posts directed to same topic in a different tense
US20160335251A1 (en) * 2015-05-11 2016-11-17 Hristo Georgiev NEWINFO, A Computer System for Automated Reasoning to find new information in Natural Language Sentences
GB2540534A (en) * 2015-06-15 2017-01-25 Erevalue Ltd A method and system for processing data using an augmented natural language processing engine
CN105930362B (zh) * 2016-04-12 2019-03-12 晶赞广告(上海)有限公司 搜索目标识别方法、装置及终端
US11093703B2 (en) * 2016-09-29 2021-08-17 Google Llc Generating charts from data in a data table
CN106844335A (zh) * 2016-12-21 2017-06-13 海航生态科技集团有限公司 自然语言处理方法及装置
CN107122398B (zh) * 2017-03-17 2021-01-01 武汉斗鱼网络科技有限公司 一种数据展示图表生成方法及系统
CN107273474A (zh) * 2017-06-08 2017-10-20 成都数联铭品科技有限公司 基于潜在语义分析的自动摘要抽取方法及系统
US20190108276A1 (en) * 2017-10-10 2019-04-11 NEGENTROPICS Mesterséges Intelligencia Kutató és Fejlesztõ Kft Methods and system for semantic search in large databases
CN107797991B (zh) * 2017-10-23 2020-11-24 南京云问网络技术有限公司 一种基于依存句法树的知识图谱扩充方法及系统
CN109285030A (zh) * 2018-08-29 2019-01-29 深圳壹账通智能科技有限公司 产品推荐方法、装置、终端及计算机可读存储介质
CN109145102B (zh) * 2018-09-06 2021-02-09 杭州安恒信息技术股份有限公司 智能问答方法及其知识图谱系统构建方法、装置、设备
CN109710733A (zh) * 2018-11-28 2019-05-03 北京永洪商智科技有限公司 一种基于智能语音识别的数据交互方法和系统
CN109684638B (zh) * 2018-12-24 2023-08-11 北京金山安全软件有限公司 分句方法及其装置、电子设备、计算机可读存储介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20070106499A1 (en) * 2005-08-09 2007-05-10 Kathleen Dahlgren Natural language search system
CN103631882A (zh) * 2013-11-14 2014-03-12 北京邮电大学 基于图挖掘技术的语义化业务生成系统和方法
CN106155999A (zh) * 2015-04-09 2016-11-23 科大讯飞股份有限公司 自然语言语义理解方法及系统
CN106649223A (zh) * 2016-12-23 2017-05-10 北京文因互联科技有限公司 基于自然语言处理的金融报告自动生成方法
CN107861933A (zh) * 2017-11-29 2018-03-30 北京百度网讯科技有限公司 生成运维报表的方法和装置
CN110222194A (zh) * 2019-05-21 2019-09-10 深圳壹账通智能科技有限公司 基于自然语言处理的数据图表生成方法和相关装置

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117785986A (zh) * 2023-11-17 2024-03-29 钉钉(中国)信息技术有限公司 图表模板生成、可视化图表生成、服务组件更新方法

Also Published As

Publication number Publication date
CN110222194A (zh) 2019-09-10
CN110222194B (zh) 2022-10-04

Similar Documents

Publication Publication Date Title
WO2020233345A1 (zh) 基于自然语言处理的数据图表生成方法和相关装置
US11017178B2 (en) Methods, devices, and systems for constructing intelligent knowledge base
US9971967B2 (en) Generating a superset of question/answer action paths based on dynamically generated type sets
US8108367B2 (en) Constraints with hidden rows in a database
US7809552B2 (en) Instance-based sentence boundary determination by optimization
WO2021017721A1 (zh) 智能问答方法、装置、介质及电子设备
KR20220123187A (ko) 다중 시스템 기반 지능형 질의 응답 방법, 장치와 기기
CN111553556A (zh) 业务数据分析方法、装置、计算机设备及存储介质
US11055353B2 (en) Typeahead and autocomplete for natural language queries
US20250094460A1 (en) Query answering method based on large model, electronic device, storage medium, and intelligent agent
CN106446122B (zh) 信息检索的方法、装置与计算设备
US20200242490A1 (en) Method and device for acquiring data model in knowledge graph, and medium
US11940996B2 (en) Unsupervised discriminative facet generation for dynamic faceted search
WO2023103914A1 (zh) 文本情感分析方法、装置及计算机可读存储介质
US20250307777A1 (en) Multi-party cross-platform query and content creation service and interface for collaboration platforms
US20140379753A1 (en) Ambiguous queries in configuration management databases
CN113220838A (zh) 确定关键信息的方法、装置、电子设备和存储介质
WO2025194913A1 (zh) 数据查询方法、装置、计算机设备和存储介质
CN117076661A (zh) 面向预训练大语言模型调优的立法规划意图识别方法
US12282513B2 (en) Optimistic facet set selection for dynamic faceted search
CN109033082A (zh) 语义模型的学习训练方法、装置及计算机可读存储介质
CN119719424A (zh) 数据血缘分析方法、装置、终端设备及计算机程序产品
CN114880351B (zh) 慢查询语句的识别方法及装置、存储介质、电子设备
CN118689988A (zh) 一种基于企业内部管理手册的问答方法
CN116340483A (zh) 信息检索方法、装置、计算机设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20809201

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 020322)

122 Ep: pct application non-entry in european phase

Ref document number: 20809201

Country of ref document: EP

Kind code of ref document: A1