WO2020155749A1 - 构建个人知识图谱的方法、装置、计算机设备和存储介质 - Google Patents

构建个人知识图谱的方法、装置、计算机设备和存储介质 Download PDF

Info

Publication number
WO2020155749A1
WO2020155749A1 PCT/CN2019/117212 CN2019117212W WO2020155749A1 WO 2020155749 A1 WO2020155749 A1 WO 2020155749A1 CN 2019117212 W CN2019117212 W CN 2019117212W WO 2020155749 A1 WO2020155749 A1 WO 2020155749A1
Authority
WO
WIPO (PCT)
Prior art keywords
content
voice
user
knowledge
folder
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/117212
Other languages
English (en)
French (fr)
Inventor
吴壮伟
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2020155749A1 publication Critical patent/WO2020155749A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/36Creation of semantic tools, e.g. ontology or thesauri

Definitions

  • This application relates to the field of knowledge graphs, in particular to a method, device, computer equipment and storage medium for constructing a personal knowledge graph.
  • Knowledge Graph also known as scientific knowledge graph, is called knowledge domain visualization or knowledge domain mapping map in the library and information industry. It is a series of various graphs showing the relationship between the development process and structure of knowledge, and the visualization technology is used to describe knowledge. Resources and their carriers, mining, analyzing, constructing, drawing and displaying knowledge and their interrelationships.
  • the main purpose of the application is to provide a method, device, computer equipment and storage medium for constructing a personal knowledge graph, which aims to solve the problem of the lack of knowledge of the user's voice output in the personal knowledge graph in the prior art.
  • this application proposes a method for constructing a personal knowledge graph, including the steps:
  • the knowledge content is added to the chain of the corresponding category in the linked list of the user to update the knowledge graph of the user.
  • the step of receiving a voice command input by a user, searching for a folder corresponding to the voice command, and storing the content text file in the folder includes:
  • the content text file is stored in the folder.
  • the step of receiving a voice command input by a user, searching for a folder corresponding to the voice command, and storing the content text file in the folder includes:
  • the content text file is stored in the folder.
  • the knowledge content is added to the chain of the corresponding category in the linked list of the user according to the timestamp of the content voice converted into the content text file and the category of the folder to update the After the steps of the user's knowledge graph, include:
  • the search chain of the linked list is determined according to the search key, and the search content is searched in the search chain.
  • the knowledge content is added to the chain of the corresponding category in the linked list of the user according to the timestamp of the content voice converted into the content text file and the category of the folder to update the After the steps of the user's knowledge graph, it also includes:
  • the knowledge content is added to the chain of the corresponding category in the linked list of the user according to the timestamp of the content voice converted into the content text file and the category of the folder to update the After the steps of the user's knowledge graph, it also includes:
  • the knowledge content is added to the chain of the corresponding category in the linked list of the user according to the timestamp of the content voice converted into the content text file and the category of the folder to update the After the steps of the user's knowledge graph, it also includes:
  • This application also provides a device for constructing a personal knowledge graph, including:
  • the receiving unit is used to receive the content voice input by the user
  • the first conversion unit is used to convert the content voice into a content text file
  • the receiving and storing unit is configured to receive a voice command input by a user, find a folder corresponding to the voice command, and store the content text file in the folder, where there are multiple folders, Different folders correspond to different voice commands;
  • a processing unit configured to perform key content extraction and similar content clustering processing on the content text files in the folder to obtain organized knowledge content
  • the update unit is configured to add the knowledge content to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update the User’s knowledge graph.
  • the present application also provides a computer device, including a memory and a processor, the memory stores computer readable instructions, and the processor implements the steps of any one of the above methods when the computer readable instructions are executed.
  • the present application also provides a computer-readable storage medium on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the steps of any one of the methods described above are implemented.
  • the method, device, computer equipment and storage medium of the present application for constructing a personal knowledge graph obtain the user’s voice information, convert it into a content text file, and then obtain a voice command to perform preliminary file classification on the content text file to improve the classification speed, And it is convenient for key content extraction and similar content clustering processing to obtain the sorted knowledge content, and finally add it to the chain of knowledge graph.
  • the user of this application can output the knowledge through voice broadcast, and it is more convenient to build the user's knowledge graph without manual typing by the user, which improves the efficiency of the knowledge graph establishment.
  • FIG. 1 is a schematic flowchart of a method for constructing a personal knowledge graph according to an embodiment of the application
  • FIG. 2 is a schematic block diagram of the structure of an apparatus for constructing a personal knowledge graph according to an embodiment of the application;
  • FIG. 3 is a schematic block diagram of the structure of a computer device according to an embodiment of the application.
  • this application provides a method for constructing a personal knowledge graph, including the steps:
  • devices that receive user input voice include various smart electronic devices that contain voice input modules, such as smart phones, tablet computers, and computers.
  • the content voice input by the user may be the voice input by the user alone, or may be the voice generated by the user and other people when the user interacts with other people in voice.
  • the content voice is the user's voice separated from the integrated voice generated when the user interacts with other people through voice separation technology. Specifically, the voice with the same characteristics as the preset user's voiceprint is separated from the above-mentioned integrated voice, and the separated voice is the content voice of the user.
  • the received content voice is converted into text content through the technology of voice-to-text, and the text content forms the above-mentioned content text file.
  • the above-mentioned speech-to-text technology can use any disclosed technology, which will not be repeated here.
  • the voice command is used to give instructions to the smart electronic device.
  • the above-mentioned voice command is used to instruct the above-mentioned content file to be stored in the corresponding folder, which plays a role of preliminary classification.
  • these folders are used to store content text files, but because the content recorded in the content text file is not easy to be classified by the computer, if you use the above smart electronic device to The content in the content text file is recognized and classified by keywords such as keywords.
  • keywords such as keywords.
  • the classification is faster by the user's active input of voice commands. , The classification result is more accurate, and the calculation amount of computer classification is reduced.
  • the above-mentioned linked list is a non-contiguous, non-sequential storage structure on a physical storage unit, and the logical sequence of data elements is realized by the link order of pointers in the linked list.
  • the content of each node of each chain on the entire linked list forms the user's knowledge graph.
  • each category of folder corresponds to a chain in the linked list.
  • Each node on the chain does not directly store all the content text files, but first preprocesses the content text files.
  • the process of preprocessing is to extract the main content of the content text files through keywords, etc., and Similar content is merged (clustered), etc., in order to streamline the content text file, clean out irrelevant content, obtain knowledge content, and then add the knowledge content to the chain node of the linked list, and the chain node will be marked
  • the above-mentioned timestamp is used to understand the completion time of the content recorded by the node, further reflect the establishment time of each knowledge point in the user's knowledge graph, and help the user sort out the knowledge points and conduct corresponding reviews.
  • the step S3 of receiving a voice command input by a user, searching for a folder corresponding to the voice command, and storing the content text file in the folder includes:
  • the content text file is stored in the folder.
  • the voice command is first parsed into text, and then the command keyword is extracted, and the corresponding folder is searched in the command list according to the command keyword.
  • the above command list is a one-to-one correspondence between command keywords and folder names, and the folder name and the corresponding folder have a one-to-one mapping relationship.
  • the aforementioned folder name is generally a category of knowledge.
  • the folder is searched by extracting the command keywords to improve the recognition accuracy of the voice command.
  • the step S3 of receiving a voice command input by a user, searching for a folder corresponding to the voice command, and storing the content text file in the folder includes:
  • the content text file is stored in the folder.
  • the method of voice similarity is used to find a standard voice command similar to the voice command.
  • Each category of standard voice commands corresponds to a category folder.
  • the calculation of the voice similarity can be performed by using the existing technology, which will not be repeated here.
  • the category standard voice commands in this application can be input by the user, which can improve the accuracy of the voice commands input by the user, that is, the category standard voice commands are entered by the user into the above-mentioned smart electronic device using their own pronunciation standards .
  • the above-mentioned knowledge content is added to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update After step S5 of the user's knowledge graph, it includes:
  • the search chain of the linked list is determined according to the search key, and the search content is searched in the search chain.
  • it is the process of retrieving the required knowledge in the above-mentioned personal knowledge graph. First convert the search voice into a text file, then extract the search keywords, determine the category to be searched according to the search keywords, and then find the chain corresponding to the linked list, search for the knowledge corresponding to the search voice at each node on the chain, and search It is fast and saves the computing resources of the computer.
  • the above-mentioned knowledge content is added to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update After step S5 of the user's knowledge graph, it further includes:
  • the above-mentioned knowledge list is the knowledge list corresponding to the knowledge graph before the update, and the knowledge list records the information of each node on each chain in the knowledge graph before the update, the time stamp of the content text on the node, and Content summary, etc.
  • Forming and displaying knowledge reports can enable users to better understand their own knowledge graph.
  • the above-mentioned knowledge content is added to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update After step S5 of the user's knowledge graph, it further includes:
  • Marking methods can include highlighting colors, highlighting text bubbles, and text flashing.
  • the above-mentioned knowledge content is added to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update After step S5 of the user's knowledge graph, it further includes:
  • the above process is the process of removing duplicate knowledge points. Because at the beginning, various types of storage are performed according to the voice commands input by the user, so there is a problem of user classification errors, such as input at different times The same content voice, and corresponding input of different voice commands, there will be different chain nodes in the knowledge graph that store the same knowledge content, so the processing of removing duplicate knowledge content in this embodiment is required.
  • the foregoing removal of duplicate knowledge content can be performed at a preset frequency, for example, once every 7 days.
  • the method for constructing a personal knowledge graph in the embodiment of the application obtains the user's voice information, converts it into a content text file, and then obtains voice commands to perform preliminary file classification on the content text file, which improves the classification speed and facilitates key content extraction and integration Similar content is clustered, and the sorted knowledge content is obtained, and finally added to the chain of knowledge graph.
  • the user of this application can output the knowledge through voice broadcast, and it is more convenient to build the user's knowledge graph without manual typing by the user, which improves the efficiency of the knowledge graph establishment.
  • the present application provides a device for constructing a personal knowledge graph, including the steps:
  • the receiving unit 10 is configured to receive content voice input by a user
  • the first conversion unit 20 is configured to convert the content voice into a content text file
  • the receiving and storing unit 30 is configured to receive a voice command input by a user, find a folder corresponding to the voice command, and store the content text file in the folder, wherein the folder is provided with multiple , Different folders correspond to different voice commands;
  • the processing unit 40 is configured to perform key content extraction and similar content clustering processing on the content text files in the folder to obtain organized knowledge content;
  • the update unit 50 is configured to add the knowledge content to the chain of the corresponding category in the user's linked list according to the timestamp of the content voice converted into the content text file and the category of the folder to update all Describe the user’s knowledge graph.
  • devices that receive user input voice include various smart electronic devices containing voice input modules, such as smart phones, tablet computers, and computers.
  • the content voice input by the user may be the voice input by the user alone, or may be the voice generated by the user and other people when the user interacts with other people in voice.
  • the content voice is the user's voice separated from the integrated voice generated when the user interacts with other people through voice separation technology. Specifically, the voice with the same characteristics as the preset user's voiceprint is separated from the above-mentioned integrated voice, and the separated voice is the content voice of the user.
  • the first conversion unit 20 is a unit that converts the received content voice into text content through a voice-to-text technology, and forms the text content into the above-mentioned content text file.
  • the above-mentioned speech-to-text technology can use any disclosed technology, which will not be repeated here.
  • the received voice command is used to give instructions to the smart electronic device.
  • the above-mentioned voice command is used to instruct the above-mentioned content file to be stored in the corresponding folder, which plays a role of preliminary classification.
  • the content recorded in content text files is not easy to be classified by a computer, if you use the above smart electronic device to The content in the content text file is recognized and classified by keywords such as keywords.
  • keywords such as keywords.
  • the classification is faster by the user's active input of voice commands. , The classification result is more accurate, and the calculation amount of computer classification is reduced.
  • the above-mentioned linked list is a non-contiguous and non-sequential storage structure on a physical storage unit, and the logical sequence of data elements is realized by the link order of pointers in the linked list.
  • the content of each node of each chain on the entire linked list forms the user's knowledge graph.
  • each category of folder corresponds to a chain in the linked list.
  • Each node on the chain does not directly store all the content text files, but first preprocesses the content text files.
  • the process of preprocessing is to extract the main content of the content text files through keywords, etc., and Similar content is merged (clustered), etc., in order to streamline the content text file, clean out irrelevant content, obtain knowledge content, and then add the knowledge content to the chain node of the linked list, and the chain node will be marked
  • the above-mentioned timestamp is used to understand the completion time of the content recorded by the node, further reflect the establishment time of each knowledge point in the user's knowledge graph, and help the user sort out the knowledge points and conduct corresponding reviews.
  • the foregoing receiving and storing unit 30 includes:
  • the first receiving module is configured to receive the voice command input by the user
  • the search module is used to search for the folder corresponding to the command keyword in the preset command list
  • the first storage module is configured to store the content text file in the folder.
  • the voice command is first parsed into text, and then the command keyword is extracted, and the corresponding folder is searched in the command list according to the command keyword.
  • the above command list is a one-to-one correspondence between command keywords and folder names, and the folder name and the corresponding folder have a one-to-one mapping relationship.
  • the aforementioned folder name is generally a category of knowledge.
  • the folder is searched by extracting the command keywords to improve the recognition accuracy of the voice command.
  • the foregoing receiving and storing unit 30 includes:
  • the second receiving module is configured to receive the voice command input by the user
  • the comparison module is used to compare the similarity between the voice command and the standard voice commands of each preset category
  • An obtaining module configured to obtain a folder corresponding to the standard voice command of the category with the greatest similarity to the voice command
  • the second storage module is used to store the content text file in the folder.
  • the method of voice similarity is used to find a standard voice command similar to the voice command.
  • Each category of standard voice commands corresponds to a category folder.
  • the calculation of the voice similarity can be performed by using the existing technology, which will not be repeated here.
  • the category standard voice commands in this application can be input by the user, which can improve the accuracy of the voice commands input by the user, that is, the category standard voice commands are entered by the user into the above-mentioned smart electronic device using their own pronunciation standards .
  • the above-mentioned apparatus for constructing a personal knowledge graph includes:
  • Receiving retrieval unit for receiving retrieval voice input by the user
  • the second conversion unit is used to convert the search voice into a search text file
  • the search unit is used to determine the search chain of the linked list according to the search key, and search for the search content in the search chain.
  • the knowledge corresponding to the voice can be retrieved quickly and save the computing resources of the computer.
  • the above-mentioned apparatus for constructing a personal knowledge graph further includes:
  • the inserting display unit is used to insert the time stamp of the content text file, the node information of the stored content text file and the knowledge content summary into the preset knowledge list to form and display the knowledge report.
  • the above-mentioned knowledge list is the knowledge list corresponding to the knowledge graph before the update, and the knowledge list records the information of each node on each chain in the knowledge graph before the update, the time stamp of the content text on the node, and Content summary, etc.
  • Forming and displaying knowledge reports can enable users to better understand their own knowledge graph.
  • the above-mentioned apparatus for constructing a personal knowledge graph further includes:
  • the marking unit is used to mark the chain corresponding to the knowledge content.
  • Marking methods can include highlighting colors, highlighting text bubbles, and text flashing.
  • the above-mentioned apparatus for constructing a personal knowledge graph further includes:
  • the traversal unit is used to traverse the nodes on each chain of the knowledge graph and determine whether the same knowledge content exists on each node
  • the similarity calculation unit is used to extract the knowledge keywords of the same knowledge content if the same knowledge content exists on each node, and calculate the similarity between the knowledge keywords and the category of each chain;
  • the retention and removal unit is used to retain the same knowledge content on the chain corresponding to the category with the highest similarity of the knowledge keyword, and remove other same knowledge content.
  • the apparatus for constructing a personal knowledge graph in the embodiment of the application obtains the user's voice information, converts it into a content text file, and then obtains voice commands to perform preliminary file classification on the content text file, improves the classification speed, and facilitates key content extraction and integration Similar content is clustered, and the sorted knowledge content is obtained, and finally added to the chain of knowledge graph.
  • the user of this application can output the knowledge through voice broadcast, and it is more convenient to build the user's knowledge graph without manual typing by the user, which improves the efficiency of the knowledge graph establishment.
  • an embodiment of the present application also provides a computer device.
  • the computer device may be the above-mentioned management server or a server corresponding to the management node, and its internal structure may be as shown in FIG. 3.
  • the computer equipment includes a processor, a memory, a network interface and a database connected by a system bus. Among them, the computer designed processor is used to provide calculation and control capabilities.
  • the memory of the computer device includes a non-volatile storage medium and an internal memory.
  • the non-volatile storage medium stores an operating system, computer readable instructions, and a database.
  • the memory provides an environment for the operation of the operating system and computer readable instructions in the non-volatile storage medium.
  • the database of the computer equipment is used to store data such as knowledge graphs.
  • the network interface of the computer device is used to communicate with an external terminal through a network connection.
  • FIG. 3 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
  • the computer device of the embodiment of the present application obtains the user's voice information, converts it into a content text file, and then obtains voice commands to perform preliminary file classification on the content text file, improve classification speed, and facilitate key content extraction and similar content clustering Process, get the organized knowledge content, and finally add it to the chain of knowledge graph.
  • the user of this application can output the knowledge through voice broadcast, and it is more convenient to build the user's knowledge graph without manual typing by the user, which improves the efficiency of the knowledge graph establishment.
  • An embodiment of the present application also provides a computer-readable storage medium.
  • the computer-readable storage medium may be a non-volatile readable storage medium or a volatile readable storage medium on which computer-readable instructions are stored.
  • the computer-readable instructions are executed by the processor, the method for constructing a personal knowledge graph as described in any of the foregoing embodiments is implemented.
  • Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • Volatile memory may include random access memory (RAM) or external cache memory.
  • RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous Link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

本申请揭示了一种构建个人知识图谱的方法、装置、计算机设备和存储介质,其中方法包括:接收用户输入的内容语音;将内容语音转换成内容文本文件;接收用户输入的语音命令,查找与语音命令对应的文件夹,并将内容文本文件存储到文件夹中,其中,文件夹设置有多个,不同的文件夹对应不同的语音命令;对文件夹中的内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;根据内容语音转换成内容文本文件的时间戳、以及文件夹的类别,将知识内容加入到用户的链表中对应类别的链条中,以更新用户的知识图谱。本申请用户可以通过语音播出的方式,对知识进行输出,无需用户手动打字,提高知识图谱建立的效率。

Description

构建个人知识图谱的方法、装置、计算机设备和存储介质
本申请要求于2019年1月31日提交中国专利局、申请号为2019101004144,申请名称为“构建个人知识图谱的方法、装置、计算机设备和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及到知识图谱领域,特别是涉及到一种构建个人知识图谱的方法、装置、计算机设备和存储介质。
背景技术
知识图谱(Knowledge Graph)又称为科学知识图谱,在图书情报界称为知识域可视化或知识领域映射地图,是显示知识发展进程与结构关系的一系列各种不同的图形,用可视化技术描述知识资源及其载体,挖掘、分析、构建、绘制和显示知识及它们之间的相互联系。
构建个人的知识图谱的时候,一般需要获取用户个人的各种数据,如抓取用户网络浏览数据、日常撰写文件的输入数据等,将这些内容进行抽取和聚类,已形成用户的知识图谱,但是这样的知识图谱并不全面,用户的知识好多是存储在大脑中,然后通过语音交互的方式进行输出,这些知识并没有很好的接入到用户的个人知识图谱中,所以,提供一种基于用户语音内容构建知识图谱的方法,是有必要的。
技术问题
申请的主要目的为提供一种构建个人知识图谱的方法、装置、计算机设备和存储介质,旨在解决现有技术中的个人知识图谱缺少用户语音输出的知识的问题。
技术解决方案
为了实现上述发明目的,本申请提出一种构建个人知识图谱的方法,括步骤:
接收用户输入的内容语音;
将所述内容语音转换成内容文本文件;
接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
进一步地,所述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
接收所述用户输入的语音命令;
将所述语音命令转换成语音文本;
提取所述语音文本中的命令关键字;
在预设的命令列表中查找与所述命令关键字对应的文件夹;
将所述内容文本文件存储到所述文件夹中。
进一步地,所述上接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
接收所述用户输入的语音命令;
将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
将所述内容文本文件存储到所述文件夹中。
进一步地,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,包括:
接收用户输入的检索语音;
将所述检索语音转换成检索文本文件;
提取所述检索文本文件的检索关键字;
根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
进一步地,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
生成所述内容文本文件的知识内容摘要;
将所述述内容文本文件的时间戳、存储内容文本文件的节点信息和知识内容摘要插入到预设的知识列表中,形成知识报表并展示。
进一步地,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
对所述知识内容对应的链条进行标记。
进一步地,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
遍历所述知识图谱的各链条上的节点,判断各节点上是否存在相同的知识内容;
若存在,则提取相同的知识内容的知识关键词,并将所述知识关键词与各所述链条的类别进行相似度计算;
保留与所述知识关键词相似度最高的类别对应的链条上的相同的知识内容,将他的相同的知识内容清除。
本申请还提供一种构建个人知识图谱的装置,包括:
接收单元,用于接收用户输入的内容语音;
第一转换单元,用于将所述内容语音转换成内容文本文件;
接收存储单元,用于接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
处理单元,用于对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
更新单元,用于根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
本申请还提供一种计算机设备,包括存储器和处理器,所述存储器存储有计算机可读指令,所述处理器执行所述计算机可读指令时实现上述任一项所述方法的步骤。
本申请还提供一种计算机可读存储介质,其上存储有计算机可读指令,所述计算机可读指令被处理器执行时实现上述任一项所述的方法的步骤。
有益效果
本申请的构建个人知识图谱的方法、装置、计算机设备和存储介质,获取用户的语音信息,将其转换成内容文本文件,然后获取语音命令对内容文本文件进行初步的文件分类,提高分类速度,并方便关键内容抽取和相似内容聚类处理,得到整理后的知识内容,最后加入到知识图谱的链条中。本申请用户可以通过语音播出的方式,对知识进行输出,建立用户的知识图谱更加方便,无需用户手动打字,提高知识图谱建立的效率。
附图说明
图1 为本申请一实施例的构建个人知识图谱的方法的流程示意图;
图2 为本申请一实施例的构建个人知识图谱的装置的结构示意框图;
图3 为本申请一实施例的计算机设备的结构示意框图。
本申请目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
本发明的最佳实施方式
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
参照图1,本申请提供一种构建个人知识图谱的方法,包括步骤:
S1、接收用户输入的内容语音;
S2、将所述内容语音转换成内容文本文件;
S3、接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
S4、对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
S5、根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
如上述步骤S1所述,接收用户输入语音的设备包括各种含有语音输入模块的智能电子设备,如智能手机、平板电脑、计算机等。上述用户输入的内容语音可以为用户独自输入的语音,也可以是用户与其他人进行语音交互时,用户与其他人共同产生的语音。在一个实施例中,内容语音是从用户与其他人进行语音交互时产生的综合语音中,通过声音分离技术,分离出的用户的语音。具体的,在上述综合语音中分离出与预设的用户声纹特征相同的语音,分离出的语音即为用户的内容语音。
如上述步骤S2所述,即为通过语音转文字的技术将接收到的内容语音转换成文字内容,文字内容形成上述的内容文本文件。上述语音转文字技术,可以使用任何一种已经公开的技术,在此不在赘述。
如上述步骤S3所述,上述语音命令用于给上述智能电子设备下达指令。本申请中,上述语音命令用于指导上述内容文件文件存储到对应的文件夹中,起到初步分类的作用。比如,预设有多个不同类别的文件夹,这些文件夹都是用于存储内容文本文件的,但是由于内容文本文件中记载的内容不容易通过计算机进行分类,若果使用上述智能电子设备对内容文本文件中的内容进行关键字等识别分类,当内容文本文件中记载的内容较多时,会消耗大量的智能电子设备的计算资源,而通过用户主动的输入语音命令进行分类,分类速度更快、分类结果更加准确,而且减少计算机分类的计算量。
如上述步骤S4和S5所述,上述链表是一种物理存储单元上的非连续、非顺序的存储结构,数据元素的逻辑顺序是通过链表中的指针链接次序实现的。整个链表上各链条的各节点的内容即形成了用户的知识图谱。本申请中,每一个类别的文件夹对应链表中的一个链条。链条上的各节点,并不是直接存储全部的内容文本文件,而是先对内容文本文件进行预处理,预处理的过程即为通过关键字等对内容文本文件中的主要内容进行提取,以及将相似的内容进行合并(聚类)等,以起到精简内容文本文件的目的,将无关的内容清洗掉,得到知识内容,然后将知识内容添加到链表的链条节点上,该链条节点上会标记上述时间戳,以便于了解该节点记录的内容的完成时间,进一步地体现用户的知识图谱中各知识点的建立时间,有助于用户梳理知识点,进行相应的复习等。
在一个实施例中,上述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤S3,包括:
接收所述用户输入的语音命令;
将所述语音命令转换成语音文本;
提取所述语音文本中的命令关键字;
在预设的命令列表中查找与所述命令关键字对应的文件夹;
将所述内容文本文件存储到所述文件夹中。
在本实施例中,先将语音命令解析成文字,然后提取出命令关键字,根据命令关键字在命令列表中查找对应的文件夹。上述命令列表是命令关键字和文件夹名一一对应的列表,文件夹名与对应的文件夹成一对一映射关系。上述文件夹名一般为知识的类别。本申请中,当用户输入的语音命令中存在多余的语句时,通过提取命令关键字查找文件夹,提高语音命令的识别准确率。
在另一个实施例中,上述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤S3,包括:
接收所述用户输入的语音命令;
将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
将所述内容文本文件存储到所述文件夹中。
在本实施例中,使用语音相似度的方法查找与所述语音命令近似的类别标准语音命令。每一种类别标准语音命令对应一个类别的文件夹。本申请中,语音相似度的计算,可以利用现有技术进行计算,在此不在赘述。需要注意的,本申请中的类别标准语音命令可以是用户输入的,可以提高对用户输入的语音命令的准确度,即,类别标准语音命令是用户使用自己的发音标准录入到上述智能电子设备中。
在一个实施例中,上述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤S5之后,包括:
接收用户输入的检索语音;
将所述检索语音转换成检索文本文件;
提取所述检索文本文件的检索关键字;
根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
在本实施例中,即为在上述的个人知识图谱中检索需要的知识的过程。先将检索语音转换成文本文件,然后提取出检索关键字,根据检索关键字确定需要检索的类别,进而查找到链表对应的链条,在该链条上的各节点查找与检索语音对应的知识,检索速度快,节约计算机的计算资源。
在一个实施例中,上述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤S5之后,还包括:
生成所述内容文本文件的知识内容摘要;
将所述述内容文本文件的时间戳、存储内容文本文件的节点信息和知识内容摘要插入到预设的知识列表中,形成知识报表并展示。
在本实施例中,上述知识列表是未更新之前的知识图谱对应的知识列表,该知识列表中记录有未更新之前的知识图谱中各链条上的各节点信息、节点上内容文本的时间戳和内容摘要等。形成知识报表并展示,可以使用户更好的了解自己的知识图谱。
在一个实施例中,上述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤S5之后,还包括:
对所述知识内容对应的链条进行标记。
在本实施例中,对所述知识内容对应的链条进行标记,可以使用户知道该链条上的实施内容是最新更新的。标记的方式可以包括突出颜色,突出文字气泡、文字闪烁等。
在一个实施例中,上述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤S5之后,还包括:
遍历所述知识图谱的各链条上的节点,判断各节点上是否存在相同的知识内容;
若存在,则提取相同的知识内容的知识关键词,并将所述知识关键词与各所述链条的类别进行相似度计算;
保留与所述知识关键词相似度最高的类别对应的链条上的相同的知识内容,将其他的相同的知识内容清除。
在本实施例中,上述过程即为去除重复知识点的过程,因为在开始的时候,是根据用户输入的语音命令进行各类存储的,所以存在用户分类错误的问题,比如在不同的时间输入相同的内容语音,而且对应输入不同的语音命令,则会出现知识图谱中存在不同的链条节点上存储有相同的知识内容,所以需要本实施例的清除重复的知识内容的处理。上述清除重复的知识内容,可以按照预设的频率进行,比如,每经过7天的时间进行一次等。
本申请实施例的构建个人知识图谱的方法,获取用户的语音信息,将其转换成内容文本文件,然后获取语音命令对内容文本文件进行初步的文件分类,提高分类速度,并方便关键内容抽取和相似内容聚类处理,得到整理后的知识内容,最后加入到知识图谱的链条中。本申请用户可以通过语音播出的方式,对知识进行输出,建立用户的知识图谱更加方便,无需用户手动打字,提高知识图谱建立的效率。
参照图2,本申请提供一种构建个人知识图谱的装置,包括步骤:
接收单元10,用于接收用户输入的内容语音;
第一转换单元20,用于将所述内容语音转换成内容文本文件;
接收存储单元30,用于接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
处理单元40,用于对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
更新单元50,用于根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
如上述接收单元10,接收用户输入语音的设备包括各种含有语音输入模块的智能电子设备,如智能手机、平板电脑、计算机等。上述用户输入的内容语音可以为用户独自输入的语音,也可以是用户与其他人进行语音交互时,用户与其他人共同产生的语音。在一个实施例中,内容语音是从用户与其他人进行语音交互时产生的综合语音中,通过声音分离技术,分离出的用户的语音。具体的,在上述综合语音中分离出与预设的用户声纹特征相同的语音,分离出的语音即为用户的内容语音。
如上述第一转换单元20,即为通过语音转文字的技术将接收到的内容语音转换成文字内容,将文字内容形成上述的内容文本文件的单元。上述语音转文字技术,可以使用任何一种已经公开的技术,在此不在赘述。
如上述接收存储单元30,接收到的语音命令用于给上述智能电子设备下达指令。本申请中,上述语音命令用于指导上述内容文件文件存储到对应的文件夹中,起到初步分类的作用。比如,预设有多个不同类别的文件夹,这些文件夹都是用于存储内容文本文件的,但是由于内容文本文件中记载的内容不容易通过计算机进行分类,若果使用上述智能电子设备对内容文本文件中的内容进行关键字等识别分类,当内容文本文件中记载的内容较多时,会消耗大量的智能电子设备的计算资源,而通过用户主动的输入语音命令进行分类,分类速度更快、分类结果更加准确,而且减少计算机分类的计算量。
如上述处理单元40更新单元50,上述链表是一种物理存储单元上的非连续、非顺序的存储结构,数据元素的逻辑顺序是通过链表中的指针链接次序实现的。整个链表上各链条的各节点的内容即形成了用户的知识图谱。本申请中,每一个类别的文件夹对应链表中的一个链条。链条上的各节点,并不是直接存储全部的内容文本文件,而是先对内容文本文件进行预处理,预处理的过程即为通过关键字等对内容文本文件中的主要内容进行提取,以及将相似的内容进行合并(聚类)等,以起到精简内容文本文件的目的,将无关的内容清洗掉,得到知识内容,然后将知识内容添加到链表的链条节点上,该链条节点上会标记上述时间戳,以便于了解该节点记录的内容的完成时间,进一步地体现用户的知识图谱中各知识点的建立时间,有助于用户梳理知识点,进行相应的复习等。
在一个实施例中,上述接收存储单元30,包括:
第一接收模块,用于接收所述用户输入的语音命令;
转换模块,用于将所述语音命令转换成语音文本;
提取模块,用于提取所述语音文本中的命令关键字;
查找模块,用于在预设的命令列表中查找与所述命令关键字对应的文件夹;
第一存储模块,用于将所述内容文本文件存储到所述文件夹中。
在本实施例中,先将语音命令解析成文字,然后提取出命令关键字,根据命令关键字在命令列表中查找对应的文件夹。上述命令列表是命令关键字和文件夹名一一对应的列表,文件夹名与对应的文件夹成一对一映射关系。上述文件夹名一般为知识的类别。本申请中,当用户输入的语音命令中存在多余的语句时,通过提取命令关键字查找文件夹,提高语音命令的识别准确率。
在另一个实施例中,上述接收存储单元30,包括:
第二接收模块,用于接收所述用户输入的语音命令;
比较模块,用于将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
获取模块,用于获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
第二存储模块,用于将所述内容文本文件存储到所述文件夹中。
在本实施例中,使用语音相似度的方法查找与所述语音命令近似的类别标准语音命令。每一种类别标准语音命令对应一个类别的文件夹。本申请中,语音相似度的计算,可以利用现有技术进行计算,在此不在赘述。需要注意的,本申请中的类别标准语音命令可以是用户输入的,可以提高对用户输入的语音命令的准确度,即,类别标准语音命令是用户使用自己的发音标准录入到上述智能电子设备中。
在一个实施例中,上述构建个人知识图谱的装置,包括:
接收检索单元,用于接收用户输入的检索语音;
第二转换单元,用于将所述检索语音转换成检索文本文件;
提取单元,用于提取所述检索文本文件的检索关键字;
查找单元,用于根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
在本实施例中,先将检索语音转换成文本文件,然后提取出检索关键字,根据检索关键字确定需要检索的类别,进而查找到链表对应的链条,在该链条上的各节点查找与检索语音对应的知识,检索速度快,节约计算机的计算资源。
在一个实施例中,上述构建个人知识图谱的装置,还包括:
生成单元,用于生成所述内容文本文件的知识内容摘要;
插入展示单元,用于将所述述内容文本文件的时间戳、存储内容文本文件的节点信息和知识内容摘要插入到预设的知识列表中,形成知识报表并展示。
在本实施例中,上述知识列表是未更新之前的知识图谱对应的知识列表,该知识列表中记录有未更新之前的知识图谱中各链条上的各节点信息、节点上内容文本的时间戳和内容摘要等。形成知识报表并展示,可以使用户更好的了解自己的知识图谱。
在一个实施例中,上述构建个人知识图谱的装置,还包括:
标记单元,用于对所述知识内容对应的链条进行标记。
在本实施例中,对所述知识内容对应的链条进行标记,可以使用户知道该链条上的实施内容是最新更新的。标记的方式可以包括突出颜色,突出文字气泡、文字闪烁等。
在一个实施例中,上述构建个人知识图谱的装置,还包括:
遍历单元,用于遍历所述知识图谱的各链条上的节点,判断各节点上是否存在相同的知识内容
相似度计算单元,用于若各节点上存在相同的知识内容在,则提取相同的知识内容的知识关键词,并将所述知识关键词与各所述链条的类别进行相似度计算;
保留清除单元,用于保留与所述知识关键词相似度最高的类别对应的链条上的相同的知识内容,将其他的相同的知识内容清除。
在本实施例中,因为在开始的时候,是根据用户输入的语音命令进行各类存储的,所以存在用户分类错误的问题,比如在不同的时间输入相同的内容语音,而且对应输入不同的语音命令,则会出现知识图谱中存在不同的链条节点上存储有相同的知识内容,所以需要本实施例的清除重复的知识内容的处理。上述清除重复的知识内容,可以按照预设的频率进行,比如,每经过7天的时间进行一次等。
本申请实施例的构建个人知识图谱的装置,获取用户的语音信息,将其转换成内容文本文件,然后获取语音命令对内容文本文件进行初步的文件分类,提高分类速度,并方便关键内容抽取和相似内容聚类处理,得到整理后的知识内容,最后加入到知识图谱的链条中。本申请用户可以通过语音播出的方式,对知识进行输出,建立用户的知识图谱更加方便,无需用户手动打字,提高知识图谱建立的效率。
参照图3,本申请实施例中还提供一种计算机设备,该计算机设备可以是上述的管理服务器,或者管理节点对应的服务器,其内部结构可以如图3所示。该计算机设备包括通过系统总线连接的处理器、存储器、网络接口和数据库。其中,该计算机设计的处理器用于提供计算和控制能力。该计算机设备的存储器包括非易失性存储介质、内存储器。该非易失性存储介质存储有操作系统、计算机可读指令和数据库。该内存器为非易失性存储介质中的操作系统和计算机可读指令的运行提供环境。该计算机设备的数据库用于存储知识图谱等数据。该计算机设备的网络接口用于与外部的终端通过网络连接通信。该计算机可读指令被处理器执行时以实现如上述任意实施例中所述构建个人知识图谱的方法。
本领域技术人员可以理解,图3中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定。
本申请实施例的计算机设备,获取用户的语音信息,将其转换成内容文本文件,然后获取语音命令对内容文本文件进行初步的文件分类,提高分类速度,并方便关键内容抽取和相似内容聚类处理,得到整理后的知识内容,最后加入到知识图谱的链条中。本申请用户可以通过语音播出的方式,对知识进行输出,建立用户的知识图谱更加方便,无需用户手动打字,提高知识图谱建立的效率。
本申请一实施例还提供一种计算机可读存储介质,该计算机可读存储介质可以是非易失性可读存储介质,也可以是易失性可读存储介质,其上存储有计算机可读指令,计算机可读指令被处理器执行时实现如上述任意实施例中所述的构建个人知识图谱的方法。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,所述的计算机可读指令可存储于一非易失性计算机可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的和实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可以包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双速据率SDRAM(SSRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
以上所述仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。

Claims (20)

  1. 一种构建个人知识图谱的方法,其特征在于,包括步骤:
    接收用户输入的内容语音;
    将所述内容语音转换成内容文本文件;
    接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
    对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
    根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
  2. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
    接收所述用户输入的语音命令;
    将所述语音命令转换成语音文本;
    提取所述语音文本中的命令关键字;
    在预设的命令列表中查找与所述命令关键字对应的文件夹;
    将所述内容文本文件存储到所述文件夹中。
  3. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述上接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
    接收所述用户输入的语音命令;
    将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
    获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
    将所述内容文本文件存储到所述文件夹中。
  4. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,包括:
    接收用户输入的检索语音;
    将所述检索语音转换成检索文本文件;
    提取所述检索文本文件的检索关键字;
    根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
  5. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
    生成所述内容文本文件的知识内容摘要;
    将所述述内容文本文件的时间戳、存储内容文本文件的节点信息和知识内容摘要插入到预设的知识列表中,形成知识报表并展示。
  6. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
    对所述知识内容对应的链条进行标记。
  7. 根据权利要求1所述的构建个人知识图谱的方法,其特征在于,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,还包括:
    遍历所述知识图谱的各链条上的节点,判断各节点上是否存在相同的知识内容;
    若存在,则提取相同的知识内容的知识关键词,并将所述知识关键词与各所述链条的类别进行相似度计算;
    保留与所述知识关键词相似度最高的类别对应的链条上的相同的知识内容,将其他的相同的知识内容清除。
  8. 一种构建个人知识图谱的装置,其特征在于,包括:
    接收单元,用于接收用户输入的内容语音;
    第一转换单元,用于将所述内容语音转换成内容文本文件;
    接收存储单元,用于接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
    处理单元,用于对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
    更新单元,用于根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
  9. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,所述接收存储单元,包括:
    第一接收模块,用于接收所述用户输入的语音命令;
    转换模块,用于将所述语音命令转换成语音文本;
    提取模块,用于提取所述语音文本中的命令关键字;
    查找模块,用于在预设的命令列表中查找与所述命令关键字对应的文件夹;
    第一存储模块,用于将所述内容文本文件存储到所述文件夹中。
  10. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,所述接收存储单元,包括:
    第二接收模块,用于接收所述用户输入的语音命令;
    比较模块,用于将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
    获取模块,用于获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
    第二存储模块,用于将所述内容文本文件存储到所述文件夹中。
  11. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,还包括:
    接收检索单元,用于接收用户输入的检索语音;
    第二转换单元,用于将所述检索语音转换成检索文本文件;
    提取单元,用于提取所述检索文本文件的检索关键字;
    查找单元,用于根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
  12. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,还包括:
    生成单元,用于生成所述内容文本文件的知识内容摘要;
    插入展示单元,用于将所述述内容文本文件的时间戳、存储内容文本文件的节点信息和知识内容摘要插入到预设的知识列表中,形成知识报表并展示。
  13. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,还包括:
    标记单元,用于对所述知识内容对应的链条进行标记。
  14. 根据权利要求8所述的构建个人知识图谱的装置,其特征在于,还包括:
    遍历单元,用于遍历所述知识图谱的各链条上的节点,判断各节点上是否存在相同的知识内容;
    相似度计算单元,用于若各节点上存在相同的知识内容在,则提取相同的知识内容的知识关键词,并将所述知识关键词与各所述链条的类别进行相似度计算;
    保留清除单元,用于保留与所述知识关键词相似度最高的类别对应的链条上的相同的知识内容,将其他的相同的知识内容清除。
  15. 一种计算机设备,包括存储器和处理器,所述存储器存储有计算机可读指令,其特征在于,所述处理器执行所述计算机可读指令时实现一种构建个人知识图谱的方法,包括步骤:
    接收用户输入的内容语音;
    将所述内容语音转换成内容文本文件;
    接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
    对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
    根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
  16. 根据权利要求15所述的计算机设备,其特征在于,所述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
    接收所述用户输入的语音命令;
    将所述语音命令转换成语音文本;
    提取所述语音文本中的命令关键字;
    在预设的命令列表中查找与所述命令关键字对应的文件夹;
    将所述内容文本文件存储到所述文件夹中。
  17. 根据权利要求15所述的计算机设备,其特征在于,所述上接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
    接收所述用户输入的语音命令;
    将所述语音命令与预设各类别的类别标准语音命令进行相似度比较;
    获取与所述语音命令相似度最大的类别标准语音命令对应的文件夹;
    将所述内容文本文件存储到所述文件夹中。
  18. 根据权利要求15所述的计算机设备,其特征在于,所述根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱的步骤之后,包括:
    接收用户输入的检索语音;
    将所述检索语音转换成检索文本文件;
    提取所述检索文本文件的检索关键字;
    根据所述检索关键字确定所述链表的检索链条,在所述检索链条中查找检索内容。
  19. 一种计算机可读存储介质,其上存储有计算机可读指令,其特征在于,所述计算机可读指令被处理器执行时实现一种构建个人知识图谱的方法,包括步骤:
    接收用户输入的内容语音;
    将所述内容语音转换成内容文本文件;
    接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中,其中,所述文件夹设置有多个,不同的文件夹对应不同的语音命令;
    对所述文件夹中的所述内容文本文件进行关键内容抽取和相似内容聚类处理,得到整理后的知识内容;
    根据所述内容语音转换成内容文本文件的时间戳、以及所述文件夹的类别,将所述知识内容加入到所述用户的链表中对应类别的链条中,以更新所述用户的知识图谱。
  20. 根据权利要求19所述的计算机可读存储介质,其特征在于,所述接收用户输入的语音命令,查找与所述语音命令对应的文件夹,并将所述内容文本文件存储到所述文件夹中的步骤,包括:
    接收所述用户输入的语音命令;
    将所述语音命令转换成语音文本;
    提取所述语音文本中的命令关键字;
    在预设的命令列表中查找与所述命令关键字对应的文件夹;
    将所述内容文本文件存储到所述文件夹中。
PCT/CN2019/117212 2019-01-31 2019-11-11 构建个人知识图谱的方法、装置、计算机设备和存储介质 Ceased WO2020155749A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910100414.4A CN109933671A (zh) 2019-01-31 2019-01-31 构建个人知识图谱的方法、装置、计算机设备和存储介质
CN201910100414.4 2019-01-31

Publications (1)

Publication Number Publication Date
WO2020155749A1 true WO2020155749A1 (zh) 2020-08-06

Family

ID=66985387

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/117212 Ceased WO2020155749A1 (zh) 2019-01-31 2019-11-11 构建个人知识图谱的方法、装置、计算机设备和存储介质

Country Status (2)

Country Link
CN (1) CN109933671A (zh)
WO (1) WO2020155749A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112163077A (zh) * 2020-09-28 2021-01-01 华南理工大学 一种面向领域问答的知识图谱构建方法
CN118245600A (zh) * 2024-03-20 2024-06-25 佛山职业技术学院 一种基于数字化的思政课程知识图谱构建方法及相关装置

Families Citing this family (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109933671A (zh) * 2019-01-31 2019-06-25 平安科技(深圳)有限公司 构建个人知识图谱的方法、装置、计算机设备和存储介质
CN110609905A (zh) * 2019-09-12 2019-12-24 深圳众赢维融科技有限公司 超点类型识别和图数据处理方法及装置
CN111368099B (zh) * 2020-03-31 2024-01-19 中国建设银行股份有限公司 核心信息语义图谱生成方法及装置
CN111563170A (zh) * 2020-04-30 2020-08-21 北京明略软件系统有限公司 一种知识图谱的生成方法、装置、计算机存储介质及终端
CN113539253B (zh) * 2020-09-18 2024-05-14 厦门市和家健脑智能科技有限公司 一种基于认知评估的音频数据处理方法和装置
CN112905805B (zh) * 2021-03-05 2023-09-15 北京中经惠众科技有限公司 知识图谱构建方法及装置、计算机设备和存储介质
CN116304068A (zh) * 2021-12-21 2023-06-23 国网上海市电力公司 一种基于知识图谱的数字员工供应链报表管理方法
CN115455243A (zh) * 2022-09-14 2022-12-09 北京神舟航天软件技术股份有限公司 一种数字资源产品管理方法及系统
CN121009965B (zh) * 2025-10-27 2026-01-27 江苏电力信息技术有限公司 一种面向电网智能规划领域的语料生成方法和装置

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106156365A (zh) * 2016-08-03 2016-11-23 北京智能管家科技有限公司 一种知识图谱的生成方法及装置
CN107633005A (zh) * 2017-08-09 2018-01-26 广州思涵信息科技有限公司 一种基于课堂教学内容的知识图谱构建、对比系统及方法
CN107644062A (zh) * 2017-08-29 2018-01-30 广州思涵信息科技有限公司 一种基于知识图谱的知识内容权重分析系统及方法
CN107967267A (zh) * 2016-10-18 2018-04-27 中兴通讯股份有限公司 一种知识图谱构建方法、装置及系统
CN108885626A (zh) * 2017-02-22 2018-11-23 谷歌有限责任公司 优化图形遍历
US20180349755A1 (en) * 2017-06-02 2018-12-06 Microsoft Technology Licensing, Llc Modeling an action completion conversation using a knowledge graph
CN109145123A (zh) * 2018-09-30 2019-01-04 国信优易数据有限公司 知识图谱模型的构建方法、智能交互方法、系统及电子设备
CN109933671A (zh) * 2019-01-31 2019-06-25 平安科技(深圳)有限公司 构建个人知识图谱的方法、装置、计算机设备和存储介质

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105095195B (zh) * 2015-07-03 2018-09-18 北京京东尚科信息技术有限公司 基于知识图谱的人机问答方法和系统

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106156365A (zh) * 2016-08-03 2016-11-23 北京智能管家科技有限公司 一种知识图谱的生成方法及装置
CN107967267A (zh) * 2016-10-18 2018-04-27 中兴通讯股份有限公司 一种知识图谱构建方法、装置及系统
CN108885626A (zh) * 2017-02-22 2018-11-23 谷歌有限责任公司 优化图形遍历
US20180349755A1 (en) * 2017-06-02 2018-12-06 Microsoft Technology Licensing, Llc Modeling an action completion conversation using a knowledge graph
CN107633005A (zh) * 2017-08-09 2018-01-26 广州思涵信息科技有限公司 一种基于课堂教学内容的知识图谱构建、对比系统及方法
CN107644062A (zh) * 2017-08-29 2018-01-30 广州思涵信息科技有限公司 一种基于知识图谱的知识内容权重分析系统及方法
CN109145123A (zh) * 2018-09-30 2019-01-04 国信优易数据有限公司 知识图谱模型的构建方法、智能交互方法、系统及电子设备
CN109933671A (zh) * 2019-01-31 2019-06-25 平安科技(深圳)有限公司 构建个人知识图谱的方法、装置、计算机设备和存储介质

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112163077A (zh) * 2020-09-28 2021-01-01 华南理工大学 一种面向领域问答的知识图谱构建方法
CN112163077B (zh) * 2020-09-28 2024-06-04 华南理工大学 一种面向领域问答的知识图谱构建方法
CN118245600A (zh) * 2024-03-20 2024-06-25 佛山职业技术学院 一种基于数字化的思政课程知识图谱构建方法及相关装置

Also Published As

Publication number Publication date
CN109933671A (zh) 2019-06-25

Similar Documents

Publication Publication Date Title
WO2020155749A1 (zh) 构建个人知识图谱的方法、装置、计算机设备和存储介质
CN111753099B (zh) 一种基于知识图谱增强档案实体关联度的方法及系统
CN108932294B (zh) 基于索引的简历数据处理方法、装置、设备及存储介质
WO2021000555A1 (zh) 基于知识图谱的问答方法、装置、计算机设备和存储介质
CN113220782A (zh) 多元测试数据源生成方法、装置、设备及介质
CN109947952B (zh) 基于英语知识图谱的检索方法、装置、设备及存储介质
WO2019227584A1 (zh) 简历数据信息解析处理方法、装置、设备及存储介质
CN103186639B (zh) 数据生成方法及系统
US9317608B2 (en) Systems and methods for parsing search queries
CN107885844A (zh) 基于分类检索的自动问答方法及系统
CN109101551B (zh) 一种问答知识库的构建方法及装置
CN105045852A (zh) 一种教学资源的全文搜索引擎系统
CN113343108A (zh) 推荐信息处理方法、装置、设备及存储介质
CN102156712A (zh) 一种基于云存储的电力信息检索方法及系统
CN116303923A (zh) 一种知识图谱问答方法、装置、计算机设备和存储介质
CN111309773A (zh) 一种车辆信息的查询方法、装置、系统及存储介质
CN117271700A (zh) 集成智能学习功能的设备使用与维修知识库
CN112380848A (zh) 文本生成方法、装置、设备及存储介质
WO2020133186A1 (zh) 一种文档信息提取方法、存储介质及终端
US20100185438A1 (en) Method of creating a dictionary
CN111177401A (zh) 一种电网自由文本知识抽取方法
CN117725182A (zh) 基于大语言模型的数据检索方法、装置、设备和存储介质
CN108399157B (zh) 实体与属性关系的动态抽取方法、服务器及可读存储介质
CN114997167A (zh) 简历内容提取方法及装置
CN109522396B (zh) 一种面向国防科技领域的知识处理方法及系统

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19913799

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19913799

Country of ref document: EP

Kind code of ref document: A1