WO2022237013A1 - 基于实体关系联合抽取的法律知识图谱构建方法及设备 - Google Patents

基于实体关系联合抽取的法律知识图谱构建方法及设备 Download PDF

Info

Publication number
WO2022237013A1
WO2022237013A1 PCT/CN2021/116053 CN2021116053W WO2022237013A1 WO 2022237013 A1 WO2022237013 A1 WO 2022237013A1 CN 2021116053 W CN2021116053 W CN 2021116053W WO 2022237013 A1 WO2022237013 A1 WO 2022237013A1
Authority
WO
WIPO (PCT)
Prior art keywords
entity
relationship
model
extraction
text
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/116053
Other languages
English (en)
French (fr)
Inventor
郑庆华
马昆明
刘均
李星熠
马黛露丝
王佳欣
朱海萍
麻珂欣
李鸿轩
魏笔凡
张玲玲
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Xian Jiaotong University
Original Assignee
Xian Jiaotong University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Xian Jiaotong University filed Critical Xian Jiaotong University
Priority to US17/956,864 priority Critical patent/US12530597B2/en
Publication of WO2022237013A1 publication Critical patent/WO2022237013A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/36Creation of semantic tools, e.g. ontology or thesauri
    • G06F16/367Ontology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/02Knowledge representation; Symbolic representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques
    • G06F18/241Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/205Parsing
    • G06F40/211Syntactic parsing, e.g. based on context-free grammar [CFG] or unification grammars
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/289Phrasal analysis, e.g. finite state techniques or chunking
    • G06F40/295Named entity recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/044Recurrent networks, e.g. Hopfield networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/19Recognition using electronic means
    • G06V30/191Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
    • G06V30/19127Extracting features by transforming the feature space, e.g. multidimensional scaling; Mappings, e.g. subspace methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/19Recognition using electronic means
    • G06V30/191Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
    • G06V30/19187Graphical models, e.g. Bayesian networks or Markov models
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/29Graphical models, e.g. Bayesian networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/042Knowledge-based neural networks; Logical representations of neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/02Knowledge representation; Symbolic representation
    • G06N5/022Knowledge engineering; Knowledge acquisition

Definitions

  • the invention belongs to the field of electronic information, and in particular relates to a method and equipment for constructing a legal knowledge graph based on joint extraction of entity relationships.
  • Knowledge Graph known as knowledge domain visualization or knowledge domain mapping map in the library and information industry, is a series of different graphics showing the relationship between knowledge development process and structure, using visualization technology to describe knowledge resources and their carriers, mining , analyze, construct, map and display knowledge and their interconnections.
  • the knowledge map is a combination of theories and methods of applied mathematics, graphics, information visualization technology, information science and other disciplines with metrology citation analysis, co-occurrence analysis and other methods, and uses the visual map to vividly display the core structure and development of the subject. History, frontier fields, and modern theories of multidisciplinary integration can provide practical and valuable references for disciplinary research.
  • Knowledge graphs are divided into general knowledge graphs and domain knowledge graphs. These two knowledge graphs mainly differ in coverage and usage.
  • the general knowledge graph is oriented to the general field and mainly contains a large amount of common sense knowledge in the real world, covering a wide range.
  • Domain knowledge graph also known as industry knowledge graph or vertical knowledge graph, is oriented to a specific field and is an industry knowledge base composed of professional data in this field. Because it is constructed based on industry data, it has strict and rich data patterns. Therefore, there are higher requirements for the depth and accuracy of knowledge in this field.
  • Domain knowledge graphs are knowledge graphs for specific domains, such as e-commerce, finance, and medical care. In comparison, domain knowledge graphs have more knowledge sources, faster scale expansion requirements, more complex knowledge structures, higher knowledge quality requirements, and wider application forms of knowledge.
  • Entity extraction technology also known as named entity recognition technology, refers to the identification of entities with specific meaning in the extracted text, mainly including names of people, places, institutions, proper nouns, etc., as well as time, quantity, currency, and proportional values.
  • the main task of relation extraction is to extract the relation between entities in the text.
  • the relationship between entities and entities is formally described as the relationship of triples ⁇ h, r, t>, where h, t represent the head entity and tail entity, and r represents the relationship between entities.
  • the purpose of the present invention is to solve the problem that the accuracy of knowledge map construction in the legal field cannot be guaranteed in the prior art, and provide a method and device for constructing a legal knowledge map based on joint extraction of entity relationships, so as to obtain a knowledge map with high accuracy.
  • the present invention has the following technical solutions:
  • a method for constructing a legal knowledge map based on entity-relationship joint extraction comprising the following steps:
  • the model architecture includes model encoding layer, head entity extraction layer and relationship-tail entity extraction layer;
  • model encoding layer uses the bert pre-training model
  • the head entity extraction layer uses two BiLSTMs as binary classifiers, and uses text encoding as the input of the classifier.
  • the output of the entity starting position corresponding to the first BiLSTM binary classifier is 1, and the output of the other positions is 0.
  • the output of the end position of the entity corresponding to the second BiLSTM binary classifier is 1, and the output of other positions is 0;
  • the relationship-tail entity extraction layer combines the encoding information of the head entity with the encoding of the sentence as input, and for each head entity, finds the tail entities that may exist under each relationship, and finally obtains a complete triple;
  • the subject and the object are used as the head entity and the tail entity respectively, and the predicate is used as the relationship.
  • the relationship set is determined according to the marked relationships, and the relationships with the same or similar semantics are merged.
  • the head entity extraction layer uses the feature vector x i output by the bert coding layer as input, and outputs the start and end signs extracted to the entity;
  • xi is the feature vector of each word
  • W s , W e are the weight matrices that two classifiers can train
  • b s , be e are the respective bias vectors; It is the sign of the starting position of the entity. When its value is close to 1, it means that the position is the starting position of the entity; It is the flag of the end position of the entity. When its value is close to 1, it means that the position is the end position of the entity.
  • the relationship-tail entity extraction layer is composed of two BiLSTMs with the same structure.
  • the input of this layer model includes the feature vector h s of the sentence, and incorporates The header entity code extracted by the previous layer where k represents the kth head entity;
  • the present invention also provides a legal knowledge map construction system based on entity relationship joint extraction, including:
  • the triplet data set building module is used to split legal text sentences into short sentences, and can complete the default subject in short sentences, and finally extract triplets from short sentences to construct triplet data set;
  • the model building and training module is used to separately construct the model encoding layer, the head entity extraction layer and the relationship-tail entity extraction layer in the model architecture, and obtain a model capable of extracting triples through training;
  • the inter-sentence relationship judgment module is used to judge the relationship between each short sentence for the legal text sentence that has not been split into short sentences;
  • the knowledge map visualization module is used to combine the triples extracted from the model with the relationship between sentences in the text to obtain the compound triples corresponding to the legal text, and to realize the visualization of the legal knowledge map.
  • the present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and operable on the processor.
  • a terminal device including a memory, a processor, and a computer program stored in the memory and operable on the processor.
  • the processor executes the computer program, the The steps of the legal knowledge graph construction method based on entity-relationship joint extraction.
  • the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a legal knowledge map based on joint entity-relationship extraction are implemented .
  • the present invention has the following beneficial effects:
  • the present invention uses a model based on the joint extraction of entity relationships to extract triples from unstructured text in the field of contract law, and finally builds a knowledge map in the legal field.
  • the present invention can avoid the wrong transmission caused by the pipeline method, and the accuracy rate high.
  • the design of the model framework of the present invention adopts the Chinese bert pre-training model as an encoder, and has a good effect on encoding Chinese text.
  • the entity extraction part uses two BiLSTM binary classifiers to identify the start position and end position of entities, which can effectively extract entities in the form of phrases in the text.
  • the present invention first extracts the head entity, and then extracts the tail entity corresponding to the entity relationship from the extracted head entity.
  • entity relations and tail entities not only the encoding information of the sentence is used, but also the encoding information of the head entity is incorporated.
  • the present invention can obtain a legal knowledge map with high accuracy, and the constructed knowledge map can be combined with deep learning technology to realize functions such as question-and-answer reasoning and related recommendations in the field of contract law.
  • Fig. 1 is a flowchart of a method for constructing a contract law knowledge graph according to an embodiment of the present invention
  • Fig. 2 is a sample diagram of a data set constructed by model training in an embodiment of the present invention
  • Fig. 3 is a model architecture diagram of extracting triples from text in an embodiment of the present invention.
  • Fig. 4 is a schematic diagram of the training and use of the extracted triplet model in the embodiment of the present invention.
  • Fig. 5 is a schematic diagram of a visualized knowledge map constructed by an embodiment of the present invention.
  • the present invention proposes a legal knowledge map construction method based on the joint extraction of entity relationships.
  • the embodiment is described by taking the contract law as an example.
  • the present invention can use the given contract law text to perform entity extraction and relationship extraction at the same time, and finally obtain a complete triplet information.
  • the contract law knowledge map can be formed by connecting the extracted triples end to end.
  • the completed knowledge graph can be combined with deep learning technology to implement functions such as question-and-answer reasoning and related recommendations in the field of contract law.
  • Entity extraction For any complete contract law text statement, it can be decomposed into the form of (h, r, t), where h represents the head entity, r represents the entity relationship, and t represents the tail entity. Entity extraction means extracting the head entity and tail entity in the text.
  • Relationship extraction The relationship here refers to the relationship between entities or the attributes of entities. This step usually extracts the corresponding entity relationship after the head entity and tail entity are extracted.
  • Joint extraction Unlike the previous entity extraction and relationship extraction, which are performed independently, the entities and relationships extracted by joint extraction affect each other. Using joint extraction can reduce the error transmission problem caused by entity extraction.
  • the embodiment of the present invention is based on the method for constructing the contract law map based on the joint extraction of entity relationships, including the following steps:
  • Step 1 Construction of contract law triple data set, including:
  • the split text sentences obtained in the first step will have the phenomenon of missing subjects in some short sentences, which will affect the subsequent triplet extraction work, so the default subject needs to be completed.
  • This method uses the open source tool pyltp combined with the method of dependency parsing to parse the default part and complete the default subject.
  • the example in the first step after the completion of the subject results in short sentences "the parties conclude the contract” and "(the parties) should have the corresponding capacity for civil rights and capacity for civil conduct”.
  • a complete sentence usually consists of three parts: subject, predicate, and object. Therefore, when labeling data, the subject and object of the sentence are used as the head entity and tail entity respectively, and the predicate is used as the initial labeling of the relationship.
  • After labeling some triples determine the relationship set according to the labeled relationship, and merge the relationships with the same or similar semantics. For example, the relationship "concept” and the relationship “definition” have similar semantics, and the relationship "definition” is unified into the relationship "concept”.
  • Figure 2 is an example of a partially labeled triplet. In the end, 798 artificially calibrated triples were obtained, and the entity-relationship set contained 25 relationships.
  • Model architecture design and model training including:
  • the design of the experimental model mainly considers the following aspects: First, the bert pre-training model uses a bidirectional Transformer, and at the same time uses the Masked Language Model in the pre-training process ( MLM) captures the representation at the word level, which makes the word vector change from only containing the previous information to the information that can learn the context. In the pre-training process, Next Sentence Prediction (NSP) is used to capture the sentence-level representation. Therefore, using the bert pre-training model at the encoding layer can better represent the deep meaning of the sentence. Second, the entities in the text of the contract law are different from those in the general field.
  • MLM Masked Language Model in the pre-training process
  • NSP Next Sentence Prediction
  • the input of the entity relationship and tail entity extraction part of the model is not only the encoding information of the entire sentence, but the encoding information of the head entity and the sentence The combination of encoding, which has a good effect on the extraction of entity relations and tail entities.
  • model framework for the joint extraction of entity relations in the field of contract law is designed, and the model framework is shown in the figure.
  • the model is divided into three parts, namely the model coding layer, the head entity extraction layer, and the relationship-tail entity extraction layer.
  • the model architecture diagram refers to Figure 3. The specific content of each part is as follows:
  • the model coding layer of the present invention adopts the Chinese pre-training model BERT-wwm-ext based on the full-word Mask on a larger-scale corpus by the Harbin Institute of Technology Xunfei Joint Laboratory. This model has obtained further performance improvements in multiple benchmark tests. . Using this model, the input text can be converted into the form of feature vectors.
  • This layer is mainly composed of two BiLSTMs with the same structure, with the feature vector x i output by the bert encoding layer as input, and the output is extracted to the start and end marks of the entity.
  • W s , W e are the weight matrices that two classifiers can train
  • b s be are their respective bias vectors. It is the sign of the starting position of the entity. When its value is close to 1, it means that the position is the starting position of the entity; It is the sign of the end position of the entity. When its value is close to 1, it means that the position is the end position of the entity; in the example in the figure, for the text "the parties should have the corresponding capacity for civil rights and capacity for civil conduct", the head entity is extracted
  • the entities that can be extracted by the layer are "parties" and "with corresponding capacity for civil rights and capacity for civil conduct”. The start and end positions of the entity.
  • This layer is similar to the head entity extraction layer, and is also composed of two BiLSTMs with the same structure.
  • the input of this layer model is no longer just the sentence feature vector h s , it also incorporates the head entity code extracted by the previous layer.
  • k represents the kth head entity. Will as the input vector for this layer.
  • the specific formula is:
  • the middle part is the training process of the model.
  • the input of the model is the text of the contract law.
  • the Bert encoding layer After passing through the Bert encoding layer, the head entity extraction layer, and the relationship-tail entity extraction layer to obtain the triplet output by the model, and then use the given
  • the fixed loss function is continuously iteratively optimized, and when the value of the loss function tends to be stable, the iteration is stopped, the training of the model is completed, and the trained model is saved.
  • Step 3 Judging the relationship between text sentences.
  • Inter-sentence relations include four kinds of relations: "condition”, “turning”, “parallel” and “cause and effect”. Among them, 85 cases of causal relations, 194 cases of conditional relations, 34 cases of turning relations and 8 cases of parallel relations were obtained.
  • the parties shall have the corresponding capacity for civil rights and capacity for civil conduct" when entering into a contract, the "conditional relationship" of the inter-sentence relationship extracted from the two short sentences.
  • Step 4 Triple composition and map visualization, including:
  • the compound triples corresponding to the legal text of the contract can be obtained.
  • the final triplet form that can be obtained is ((parties, conclude, contract), condition, (parties, should, have the corresponding capacity for civil rights and capacity for civil conduct)).
  • all the extracted triples are integrated and spliced to obtain a complete knowledge map of contract law.
  • FIG. 5 is a partial schematic diagram of the visualized contract law knowledge map.
  • the completed contract law knowledge graph can be combined with deep learning technology to realize the functions of question-and-answer reasoning and related recommendations in the field of contract law.
  • a legal knowledge graph construction system based on entity-relationship joint extraction including:
  • the triplet data set building module is used to split legal text sentences into short sentences, and can complete the default subject in short sentences, and finally extract triplets from short sentences to construct triplet data set;
  • the model building and training module is used to separately construct the model encoding layer, the head entity extraction layer and the relationship-tail entity extraction layer in the model architecture, and obtain a model capable of extracting triples through training;
  • the inter-sentence relationship judgment module is used to judge the relationship between each short sentence for the legal text sentence that has not been split into short sentences;
  • the knowledge map visualization module is used to combine the triples extracted from the model with the relationship between sentences in the text to obtain the compound triples corresponding to the legal text, and to realize the visualization of the legal knowledge map.
  • a terminal device comprising a memory, a processor, and a computer program stored in the memory and operable on the processor, when the processor executes the computer program, the joint extraction based on entity relationship is realized.
  • a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a legal knowledge map based on entity-relationship joint extraction are realized.
  • the computer program can be divided into one or more modules/units, and the one or more modules/units are stored in the memory and executed by the processor to complete the construction of the knowledge map of the present invention method.
  • the processor can be a central processing unit (Central Processing Unit, CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), off-the-shelf programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
  • the memory can be used to store computer programs and/or modules, and the processor implements the knowledge graph construction system of the present invention by running or executing the computer programs and/or modules stored in the memory, and calling the data stored in the memory. various functions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Mathematical Physics (AREA)
  • Computing Systems (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Multimedia (AREA)
  • Animal Behavior & Ethology (AREA)
  • Databases & Information Systems (AREA)
  • Machine Translation (AREA)

Abstract

一种基于实体关系联合抽取的法律知识图谱构建方法及设备,构建方法包括:三元组数据集的构建;模型架构的设计和模型的训练;模型架构包括模型编码层、头实体抽取层以及关系-尾实体抽取层;文本句间关系判断;以及三元组复合与图谱可视化;本发明模型架构的设计采用了中文bert预训练模型作为编码器,对中文的文本编码效果好。实体抽取部分采用两个BiLSTM二分类器来判别实体的起始位置和结束位置,可以有效地抽取出文本中短语形式的实体。本发明先抽取头实体,再由抽取到的头实体抽取对应实体关系的尾实体,抽取实体关系和尾实体时不仅用到了句子的编码信息,还融入了头实体的编码信息。本发明能够得到准确率较高的法律知识图谱。

Description

基于实体关系联合抽取的法律知识图谱构建方法及设备 技术领域
本发明属于电子信息领域,具体涉及一种基于实体关系联合抽取的法律知识图谱构建方法及设备。
背景技术
知识图谱(Knowledge Graph),在图书情报界称为知识域可视化或知识领域映射地图,是显示知识发展进程与结构关系的一系列各种不同的图形,用可视化技术描述知识资源及其载体,挖掘、分析、构建、绘制和显示知识及它们之间的相互联系。知识图谱是通过将应用数学、图形学、信息可视化技术、信息科学等学科的理论与方法与计量学引文分析、共现分析等方法结合,并利用可视化的图谱形象地展示学科的核心结构、发展历史、前沿领域以及整体知识架构达到多学科融合目的的现代理论,能为学科研究提供切实的、有价值的参考。
知识图谱分为通用知识图谱与领域知识图谱两类。这两种知识图谱主要存在覆盖范围和使用方式上的差异。通用知识图谱面向通用领域,主要包含了大量现实世界中的常识性知识,覆盖面广。领域知识图谱又称为行业知识图谱或垂直知识图谱,是面向某一特定领域的,是由该领域的专业数据构成的行业知识库,因其基于行业数据构建,有着严格而丰富的数据模式,所以对该领域知识的深度、知识准确性有着更高的要求。领域知识图谱是面向特定领域的知识图谱,如电商、金融、医疗等。相比较而言,领域知识图谱的知识来源更多、规模化扩展要求更迅速、知识结构更加复杂、知识质量要求更高、知识的应用形式也更加广泛。
面向法律领域的知识图谱构建与研究现今仍然较为匮乏,在法律发展较为快速全面的如今,法律领域知识图谱的需求也渐渐浮现。领域知识图谱的构建需要大量该领域的信息,如何从海量的无结构或半结构中抽取出有价值的信息,引起了众多学者的关注,信息抽取技术 应运而生。其中构建知识图谱主要用到信息抽取中的实体抽取和关系抽取子任务。实体抽取技术,又称命名实体识别技术,是指识别抽取文本中具有特定意义的实体,主要包括人名、地名、机构名、专有名词等,以及时间、数量、货币、比例数值等文字。关系抽取的主要任务是抽取文本中实体之间的关系。通常将实体和实体之间的关系形式化地描述为三元组的关系<h,r,t>,其中h,t表示头实体和尾实体,r表示实体之间的关系。例如,在“《霸王别姬》的导演是陈凯歌”这句话中,“《霸王别姬》”和“陈凯歌”都是实体,两个实体之间的关系是“导演”关系,用三元组就可以表示为<《霸王别姬》,导演,陈凯歌>,信息抽取的主要目的就是在大量无结构或半结构文本中抽取出三元组形式的数据,最终由大量的三元组构成知识图谱。然而,传统的信息抽取技术采用的是一种pipeline的方式,对于非结构化的文本,先进行实体抽取,然后在实体抽取结果的基础上进行关系抽取,这样做有一个很大的弊端,一旦实体抽取的结果出错,将会很大程度上影响关系抽取的准确率,这样就会导致错误的传递。
发明内容
本发明目的在于针对现有技术中法律领域知识图谱构建准确率不能保证的问题,提供一种基于实体关系联合抽取的法律知识图谱构建方法及设备,得到准确率较高的知识图谱。
为了实现上述目的,本发明有以下的技术方案:
一种基于实体关系联合抽取的法律知识图谱构建方法,包括以下步骤:
-三元组数据集的构建;
将法律文本语句拆分成短句的形式;
将短句中缺省的主语补全;
从短句中抽取三元组,构建三元组数据集;
-模型架构的设计和模型的训练;
模型架构包括模型编码层、头实体抽取层以及关系-尾实体抽取层;
具体的,模型编码层使用bert预训练模型;
头实体抽取层使用两个BiLSTM作为二分类器,将文本的编码作为分类器的输入,输出信息中,第一个BiLSTM二分类器对应的实体起始位置输出为1,其余位置输出都为0,第二个BiLSTM二分类器对应的实体结束位置输出为1,其余位置输出都为0;
关系-尾实体抽取层,将头实体的编码信息与句子的编码相结合作为输入,对于每个头实体,找到每个关系下可能存在的尾实体,最终得到完整的三元组;
-文本句间关系判断;
对于未进行短句拆分的法律文本语句,判断各个短句之间的关系;
-三元组复合与图谱可视化;
根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组;
法律知识图谱的可视化。
作为本发明基于实体关系联合抽取的法律知识图谱构建方法的一种优选方案,构建三元组数据集时将主语和宾语分别作为头实体和尾实体,将谓语作为关系。
作为本发明基于实体关系联合抽取的法律知识图谱构建方法的一种优选方案,根据标注的关系确定关系集合,合并语义相同或相似的关系。
作为本发明基于实体关系联合抽取的法律知识图谱构建方法的一种优选方案,头实体抽取层以bert编码层输出的特征向量x i作为输入,输出抽取到实体的起始和结束标志;
Figure PCTCN2021116053-appb-000001
Figure PCTCN2021116053-appb-000002
其中x i为每个词的特征向量,W s,W e为两个二分类器能够训练的权重矩阵,b s,b e为各自的偏置向量;
Figure PCTCN2021116053-appb-000003
为实体的起始位置的标志,当其值趋近于1时,表示该位置是实体的起始位置;
Figure PCTCN2021116053-appb-000004
为实体的结束位置的标志,当其值趋近于1时,表示该位置是实体的结束位置。
作为本发明基于实体关系联合抽取的法律知识图谱构建方法的一种优选方案,关系-尾实体抽取层采用两个结构相同的BiLSTM组成,该层模型的输入包括句子的特征向量h s,并融入了上一层抽取的头实体编码
Figure PCTCN2021116053-appb-000005
其中k表示第k个头实体;
Figure PCTCN2021116053-appb-000006
作为该层的输入向量x i,具体的计算公式如下:
Figure PCTCN2021116053-appb-000007
Figure PCTCN2021116053-appb-000008
其中,向量h s
Figure PCTCN2021116053-appb-000009
是直接向量相加的关系,维度相同;对于第k个头实体,取其开始位置到结束位置词向量的平均值作为向量
Figure PCTCN2021116053-appb-000010
的表示;
Figure PCTCN2021116053-appb-000011
Figure PCTCN2021116053-appb-000012
表示起始位置和结束位置能够训练的参数矩阵;对于每个头实体,遍历关系集合中所有的关系,重复上面的计算公式,找到每个关系下可能存在的尾实体,最终得到完整的三元组。
本发明还提供一种基于实体关系联合抽取的法律知识图谱构建系统,包括:
三元组数据集构建模块,用于将法律文本语句拆分成短句的形式,并能够将短句中缺省的主语补全,最终从短句中抽取三元组,构建三元组数据集;
模型建立和训练模块,用于对模型架构中的模型编码层、头实体抽取层以及关系-尾实体抽取层分别构建,经过训练得到能够抽取出三元组的模型;
句间关系判断模块,用于对未进行短句拆分的法律文本语句,判断各个短句之间的关系;
知识图谱可视化模块,用于根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组,并实现法律知识图谱的可视化。
本发明还提供一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机程序,所述的处理器执行所述的计算机程序时实现所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
本发明还提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序, 所述的计算机程序被处理器执行时实现所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
相较于现有技术,本发明有如下的有益效果:
现有的知识图谱构建方法往往采用的是pipeline的思想,先进行实体抽取,并以实体抽取的结果进行关系抽取,这样会导致错误的累计。本发明利用一种基于实体关系联合抽取的模型,对合同法领域的非结构化文本进行三元组抽取,最终构建出法律领域的知识图谱,本发明能够避免pipeline方法导致的错误传递,准确率高。本发明模型架构的设计采用了中文bert预训练模型作为编码器,对中文的文本编码效果好。由于法律领域的实体是短语形式,因此,实体抽取部分采用两个BiLSTM二分类器来判别实体的起始位置和结束位置,可以有效地抽取出文本中短语形式的实体。本发明给定一段文本,先抽取头实体,再由抽取到的头实体抽取对应实体关系的尾实体。抽取实体关系和尾实体时不仅用到了句子的编码信息,还融入了头实体的编码信息。本发明能够得到准确率较高的法律知识图谱,构建完成的知识图谱可以结合深度学习技术实现合同法领域的问答推理及相关推荐等功能。
附图说明
图1是本发明实施例合同法知识图谱构建方法流程图;
图2是本发明实施例模型训练构建的数据集样例图;
图3是本发明实施例从文本抽取三元组的模型架构图;
图4是本发明实施例抽取三元组模型的训练和使用示意图;
图5是本发明实施例构建的知识图谱可视化示意图。
具体实施方式
下面结合附图及实施例对本发明做进一步的详细说明。
本发明提出一种基于实体关系联合抽取的法律知识图谱构建方法,实施例以合同法为例 进行说明,本发明可以利用给定的合同法文本,同时进行实体抽取和关系抽取,最终得到完整的三元组信息。由抽取到的三元组首尾连接即可构成合同法知识图谱。构建完成的知识图谱可以结合深度学习技术实现合同法领域的问答推理及相关推荐等功能。
实体抽取:对于任何一个完整的合同法文本语句,都可以将其分解为(h,r,t)的形式,h表示头实体,r表示实体关系,t表示尾实体。实体抽取表示抽取出文本中的头实体和尾实体。
关系抽取:这里的关系指的是实体之间的关系或实体的属性,这一步通常在头实体和尾实体抽取完成后抽取对应的实体关系。
联合抽取:不同于以往的实体抽取和关系抽取各自独立分别进行,联合抽取所抽取到的实体和关系互相影响,采用联合抽取可以减少由实体抽取导致的错误传递问题。
参见图1,本发明实施例基于实体关系联合抽取的合同法图谱构建方法,包括以下步骤:
步骤一、合同法三元组数据集的构建,包括:
1.1)将复杂的合同法文本语句拆分成简单短句的形式。
根据合同法文本语句本身特点,绝大部分合同法文本都是由两个及以上短句构成,且短句之间都存在一定的逻辑关系。如合同法第九条“当事人订立合同,应当具有相应的民事权利能力和民事行为能力”。为了准确抽取出文本中的三元组,需要将其拆分成多个短句的形式。本例可以得到两个短句“当事人订立合同”和“应当具有相应的民事权利能力和民事行为能力”。
1.2)利用零指代消解方面的技术解决因短句拆分导致的主语缺失问题。
经过第1步得到的拆分后的文本短句会存在部分短句中主语缺失的现象,这会影响后续三元组的抽取工作,因此需要将缺省的主语补全。本方法采用开源工具pyltp结合依存句法分析的方法对缺省部分进行句法分析,并补全缺省的主语。第1步中的例子经过主语补全后的结果为短句“当事人订立合同”和“(当事人)应当具有相应的民事权利能力和民事行为能力”。
1.3)合同法三元组数据集的构建;
对于进行主语补全的短句,可以从中抽取出需要的三元组。为了保证三元组抽取模型的性能,需要人工标注三元组数据以训练模型。一个完整的句子通常由主语、谓语和宾语三部分构成,因此在标注数据时,将句子的主语和宾语分别作为头实体和尾实体,将谓语作为关系初步标注。标注部分三元组后,根据标注的关系确定关系集合,合并语义相同或相似的关系,如关系“概念”和关系“定义”语义相似,将关系“定义”统一化为关系“概念”。图2是部分标注的三元组示例。最终得到人工标定的三元组798个,实体关系集合中包含25个关系。
步骤二、模型架构的设计和模型的训练,包括:
2.1)模型架构的设计;
三元组训练数据标注完成后,下面进行实验模型的设计,本模型的设计主要考虑以下几个方面:第一,bert预训练模型使用了双向Transformer,同时在预训练过程中使用Masked Language Model(MLM)捕获词语级别的表示,这使得词向量从先前只包含前文信息变成了可以学习上下文的信息,在预训练过程中使用Next Sentence Prediction(NSP)捕获句子级别的表示。因此在编码层使用bert预训练模型可以更好地表征句子的深层含义。第二,合同法法条文本中的实体与通用领域的实体有所不同,不仅包含词语实体,还包含短语实体,这样再用传统的NER方法无法准确抽取出短语实体。因此考虑使用两个BiLSTM作为二分类器,将文本的编码作为分类器的输入,输出信息中,第一个BiLSTM二分类器对应的实体起始位置输出为1其余位置输出都为0,第二个BiLSTM二分类器对应的实体结束位置输出为1,其余位置输出都为0。分别提取实体的起始位置和结束位置的位置编码,这样可以根据需要很好地抽取出短语实体。第三,为了使得实体关系和尾实体的抽取充分利用头实体的编码信息,对于模型的实体关系和尾实体抽取部分的输入不只是整个句子的编码信息,而是将头实体的编码信息与句子的编码相结合,这对于实体关系和尾实体的抽取有很好的效果提升。
通过以上的分析,设计了合同法领域实体关系联合抽取的算法模型,模型框架图如图所 示。该模型共分为三个部分,分别是模型编码层,头实体抽取层,关系-尾实体抽取层,模型架构图参考图3,各个部分的具体内容如下:
(a)模型编码层
本发明的模型编码层采用的是哈工大讯飞联合实验室在更大规模语料上基于全词Mask的中文预训练模型BERT-wwm-ext,该模型在多项基准测试上获得了进一步的性能提升。利用该模型可以将输入的文本转化为特征向量的形式。
(b)头实体抽取层
该层主要由两个结构相同的BiLSTM组成,以bert编码层输出的特征向量x i作为输入,输出抽取到实体的起始和结束标志。
Figure PCTCN2021116053-appb-000013
Figure PCTCN2021116053-appb-000014
其中x i为每个词的特征向量,W s,W e为两个二分类器可以训练的权重矩阵,b s,b e为各自的偏置向量。
Figure PCTCN2021116053-appb-000015
为实体的起始位置的标志,当其值趋近于1时,表示该位置是实体的起始位置;
Figure PCTCN2021116053-appb-000016
为实体的结束位置的标志,当其值趋近于1时,表示该位置是实体的结束位置;图中示例,对于文本“当事人应当具有相应的民事权利能力和民事行为能力”,经过头实体抽取层可以抽取到的实体分别为“当事人”和“具有相应的民事权利能力和民事行为能力”,图3中黑色标记分别为“当事人”实体的起始位置和结束位置,浅灰色标记则为另一个实体的起始位置和结束位置。
(c)关系-尾实体抽取层
该层的与头实体抽取层相似,也是由两个结构相同的BiLSTM组成,该层模型的输入不再只是句子的特征向量h s,它还融入了上一层抽取的头实体编码
Figure PCTCN2021116053-appb-000017
其中k表示第k个头实体。将
Figure PCTCN2021116053-appb-000018
作为该层的输入向量。具体公式为:
Figure PCTCN2021116053-appb-000019
Figure PCTCN2021116053-appb-000020
其中,向量h s
Figure PCTCN2021116053-appb-000021
是直接向量相加的关系,因此必须要保证维度相同,所以对于第k个头实体,取其开始位置到结束位置词向量的平均值作为向量
Figure PCTCN2021116053-appb-000022
的表示;
Figure PCTCN2021116053-appb-000023
Figure PCTCN2021116053-appb-000024
表示起始位置和结束位置可训练的参数矩阵,与头实体抽取层不同,这里表示对于每个头实体,遍历关系集合中所有的关系,重复上面的计算公式,从而找到每个关系下可能存在的尾实体,最终可以得到完整的三元组。对于图3中的示例,头实体“当事人”在关系为“应当”时对应的尾实体为“具有相应的民事权利能力和民事行为能力”,因此得到三元组(当事人,应当,具有相应的民事权利能力和民事行为能力)。
2.2)模型的使用;
参见图4,中间部分为模型的训练过程,模型的输入为合同法法条文本,分别经过bert编码层,头实体抽取层,关系-尾实体抽取层得到模型输出的三元组,然后利用给定的损失函数不断地进行迭代优化,当损失函数值趋于稳定时停止迭代,完成模型的训练,保存训练完成的模型。
对于未包含在测试集中的合同法文本三元组的抽取,利用训练好的模型,将其作为模型的输入,模型的输出即为文本对应的三元组。
步骤三、文本句间关系的判断。
对于未进行短句拆分的合同法文本,利用开源工具pyltp结合规则匹配的方法判断各个短句之间的关系。句间关系包含“条件”、“转折”、“并列”、“因果”四种关系,其中共得到因果关系85例,条件关系194例,转折关系34例,并列关系8例。例如,对于合同法文本“当事人订立合同,应当具有相应的民事权利能力和民事行为能力”,两个短句中抽取到的句间关系的“条件关系”。
步骤四、三元组复合与图谱可视化,包括:
4.1)三元组的整合;
将模型抽取出的三元组,结合过程3得到的句间关系,可以得到合同法文本对应的复合三元组。对于例子“当事人订立合同,应当具有相应的民事权利能力和民事行为能力”,最终可以得到的三元组形式为((当事人,订立,合同),条件,(当事人,应当,具有相应的民事权利能力和民事行为能力))。由此将抽取到的所有三元组进行整合拼接,可以得到完整的合同法知识图谱。
4.2)合同法知识图谱的可视化;
参见图5,是合同法知识图谱可视化后的部分示意图。构建完成的合同法知识图谱可以结合深度学习技术实现合同法领域的问答推理及相关推荐等功能。
一种基于实体关系联合抽取的法律知识图谱构建系统,包括:
三元组数据集构建模块,用于将法律文本语句拆分成短句的形式,并能够将短句中缺省的主语补全,最终从短句中抽取三元组,构建三元组数据集;
模型建立和训练模块,用于对模型架构中的模型编码层、头实体抽取层以及关系-尾实体抽取层分别构建,经过训练得到能够抽取出三元组的模型;
句间关系判断模块,用于对未进行短句拆分的法律文本语句,判断各个短句之间的关系;
知识图谱可视化模块,用于根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组,并实现法律知识图谱的可视化。
一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机程序,所述的处理器执行所述的计算机程序时实现所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述的计算机 程序被处理器执行时实现所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
所述的计算机程序可以被分割成一个或多个模块/单元,所述一个或者多个模块/单元被存储在所述存储器中,并由所述处理器执行,以完成本发明的知识图谱构建方法。
处理器可以是中央处理单元(CentralProcessingUnit,CPU),还可以是其他通用处理器、数字信号处理器(DigitalSignalProcessor,DSP)、专用集成电路(ApplicationSpecificIntegratedCircuit,ASIC)、现成可编程门阵列(Field-ProgrammableGateArray,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。存储器可用于存储计算机程序和/或模块,所述处理器通过运行或执行存储在所述存储器内的计算机程序和/或模块,以及调用存储在存储器内的数据,实现本发明知识图谱构建系统的各种功能。
以上所述的仅仅是本发明的较佳实施例,并不用以对本发明的技术方案进行任何限制,本领域技术人员应当理解的是,在不脱离本发明精神和原则的前提下,该技术方案还可以进行若干简单的修改和替换,这些修改和替换也均属于权利要求书所涵盖的保护范围之内。

Claims (8)

  1. 一种基于实体关系联合抽取的法律知识图谱构建方法,其特征在于,包括以下步骤:
    -三元组数据集的构建;
    将法律文本语句拆分成短句的形式;
    将短句中缺省的主语补全;
    从短句中抽取三元组,构建三元组数据集;
    -模型架构的设计和模型的训练;
    模型架构包括模型编码层、头实体抽取层以及关系-尾实体抽取层;
    具体的,模型编码层使用bert预训练模型;
    头实体抽取层使用两个BiLSTM作为二分类器,将文本的编码作为分类器的输入,输出信息中,第一个BiLSTM二分类器对应的实体起始位置输出为1,其余位置输出都为0,第二个BiLSTM二分类器对应的实体结束位置输出为1,其余位置输出都为0;
    关系-尾实体抽取层,将头实体的编码信息与句子的编码相结合作为输入,对于每个头实体,找到每个关系下可能存在的尾实体,最终得到完整的三元组;
    -文本句间关系判断;
    对于未进行短句拆分的法律文本语句,判断各个短句之间的关系;
    -三元组复合与图谱可视化;
    根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组;
    法律知识图谱的可视化。
  2. 根据权利要求1所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:构建三元组数据集时将主语和宾语分别作为头实体和尾实体,将谓语作为关系。
  3. 根据权利要求2所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:根据标注的关系确定关系集合,合并语义相同或相似的关系。
  4. 根据权利要求1所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:头实体抽取层以bert编码层输出的特征向量x i作为输入,输出抽取到实体的起始和结束标志;
    Figure PCTCN2021116053-appb-100001
    Figure PCTCN2021116053-appb-100002
    其中x i为每个词的特征向量,W s,W e为两个二分类器能够训练的权重矩阵,b s,b e为各自的偏置向量;
    Figure PCTCN2021116053-appb-100003
    为实体的起始位置的标志,当其值趋近于1时,表示该位置是实体的起始位置;
    Figure PCTCN2021116053-appb-100004
    为实体的结束位置的标志,当其值趋近于1时,表示该位置是实体的结束位置。
  5. 根据权利要求1所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:关系-尾实体抽取层采用两个结构相同的BiLSTM组成,该层模型的输入包括句子的特征向量h s,并融入了上一层抽取的头实体编码
    Figure PCTCN2021116053-appb-100005
    其中k表示第k个头实体;
    Figure PCTCN2021116053-appb-100006
    作为该层的输入向量x i,具体的计算公式如下:
    Figure PCTCN2021116053-appb-100007
    Figure PCTCN2021116053-appb-100008
    其中,向量h s
    Figure PCTCN2021116053-appb-100009
    是直接向量相加的关系,维度相同;对于第k个头实体,取其开始位置到结束位置词向量的平均值作为向量
    Figure PCTCN2021116053-appb-100010
    的表示;
    Figure PCTCN2021116053-appb-100011
    Figure PCTCN2021116053-appb-100012
    表示起始位置和结束位置能够训练的参数矩阵;对于每个头实体,遍历关系集合中所有的关系,重复上面的计算公式,找到每个关系下可能存在的尾实体,最终得到完整的三元组。
  6. 一种基于实体关系联合抽取的法律知识图谱构建系统,其特征在于,包括:
    三元组数据集构建模块,用于将法律文本语句拆分成短句的形式,并能够将短句中缺省的主语补全,最终从短句中抽取三元组,构建三元组数据集;
    模型建立和训练模块,用于对模型架构中的模型编码层、头实体抽取层以及关系-尾实体抽取层分别构建,经过训练得到能够抽取出三元组的模型;
    句间关系判断模块,用于对未进行短句拆分的法律文本语句,判断各个短句之间的关系;
    知识图谱可视化模块,用于根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组,并实现法律知识图谱的可视化。
  7. 一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机程序,其特征在于:所述的处理器执行所述的计算机程序时实现如权利要求1至5中任意一项所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
  8. 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,其特征在于:所述的计算机程序被处理器执行时实现如权利要求1至5中任意一项所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
PCT/CN2021/116053 2021-05-11 2021-09-01 基于实体关系联合抽取的法律知识图谱构建方法及设备 Ceased WO2022237013A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/956,864 US12530597B2 (en) 2021-05-11 2022-09-30 Method and device for constructing legal knowledge graph based on joint entity and relation extraction

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110513432.2A CN113204649A (zh) 2021-05-11 2021-05-11 基于实体关系联合抽取的法律知识图谱构建方法及设备
CN202110513432.2 2021-05-11

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/956,864 Continuation US12530597B2 (en) 2021-05-11 2022-09-30 Method and device for constructing legal knowledge graph based on joint entity and relation extraction

Publications (1)

Publication Number Publication Date
WO2022237013A1 true WO2022237013A1 (zh) 2022-11-17

Family

ID=77030936

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/116053 Ceased WO2022237013A1 (zh) 2021-05-11 2021-09-01 基于实体关系联合抽取的法律知识图谱构建方法及设备

Country Status (3)

Country Link
US (1) US12530597B2 (zh)
CN (1) CN113204649A (zh)
WO (1) WO2022237013A1 (zh)

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116361477A (zh) * 2022-11-28 2023-06-30 中国电力科学研究院有限公司 信息驱动的电网运行态势知识图谱智能构建系统及方法
CN116680392A (zh) * 2023-06-05 2023-09-01 北京沃东天骏信息技术有限公司 一种关系三元组的抽取方法和装置
CN116955560A (zh) * 2023-07-21 2023-10-27 广州拓尔思大数据有限公司 基于思考链和知识图谱的数据处理方法及系统
CN116992042A (zh) * 2023-07-14 2023-11-03 珠海中科先进技术研究院有限公司 基于新型研发机构科技创新服务知识图谱系统的构建方法
CN117033666A (zh) * 2023-10-07 2023-11-10 之江实验室 一种多模态知识图谱的构建方法、装置、存储介质及设备
CN117236432A (zh) * 2023-09-26 2023-12-15 中国科学院沈阳自动化研究所 一种面向多模态数据的制造工艺知识图谱构建方法及系统
CN118297065A (zh) * 2024-03-01 2024-07-05 华中科技大学 一种低资源场景下基于提示学习的关系抽取方法

Families Citing this family (35)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111291185B (zh) * 2020-01-21 2023-09-22 京东方科技集团股份有限公司 信息抽取方法、装置、电子设备及存储介质
CN113204649A (zh) * 2021-05-11 2021-08-03 西安交通大学 基于实体关系联合抽取的法律知识图谱构建方法及设备
CN113779260B (zh) * 2021-08-12 2023-07-18 华东师范大学 一种基于预训练模型的领域图谱实体和关系联合抽取方法及系统
CN114004223B (zh) * 2021-10-12 2022-05-24 北京理工大学 一种基于行为基的事件知识表示方法
CN114118056A (zh) * 2021-10-13 2022-03-01 中国人民解放军军事科学院国防工程研究院工程防护研究所 一种战争类研究报告的信息抽取方法
CN113962214B (zh) * 2021-10-25 2024-07-16 东南大学 基于eletric-bert的实体抽取方法
CN114138979B (zh) * 2021-10-29 2022-09-16 中南民族大学 基于词拓展无监督文本分类的文物安全知识图谱创建方法
CN114357163B (zh) * 2021-11-30 2025-09-02 腾讯科技(深圳)有限公司 文本类型识别方法、装置、计算机可读介质及电子设备
CN115358227A (zh) * 2022-04-13 2022-11-18 苏州空天信息研究院 一种基于短语增强的开放域关系联合抽取方法及系统
CN114898384A (zh) * 2022-06-01 2022-08-12 北京金山数字娱乐科技有限公司 合同识别方法及装置
CN115168599B (zh) * 2022-06-20 2023-06-20 北京百度网讯科技有限公司 多三元组抽取方法、装置、设备、介质及产品
CN115114382B (zh) * 2022-07-01 2024-09-06 沈阳航空航天大学 基于预训练模型与规则结合的武器装备实体关系抽取方法
CN115495585B (zh) * 2022-08-31 2026-01-09 上海海洋大学 一种基于知识图谱的花卉病虫害的本体建模方法和建模系统
CN115757815A (zh) * 2022-11-04 2023-03-07 北京中科凡语科技有限公司 知识图谱的构建方法、装置及存储介质
CN115730078A (zh) * 2022-11-04 2023-03-03 南京擎盾信息科技有限公司 用于类案检索的事件知识图谱构建方法、装置及电子设备
CN116049422B (zh) * 2022-12-07 2026-03-06 安徽大学 基于联合抽取模型的包虫病知识图谱构建方法及其应用
CN116150404A (zh) * 2023-03-03 2023-05-23 成都康赛信息技术有限公司 一种基于联合学习的教育资源多模态知识图谱构建方法
CN116226408B (zh) * 2023-03-27 2023-12-19 中国科学院空天信息创新研究院 农产品生长环境知识图谱构建方法及装置、存储介质
CN116049148B (zh) * 2023-04-03 2023-07-18 中国科学院成都文献情报中心 一种元出版环境下领域元知识引擎的构建方法
CN117009478A (zh) * 2023-06-25 2023-11-07 武汉光庭信息技术股份有限公司 一种基于软件知识图谱问答问句解析过程的算法融合方法
CN116501830B (zh) * 2023-06-29 2023-09-05 中南大学 一种生物医学文本的重叠关系联合抽取方法及相关设备
CN116562303B (zh) * 2023-07-04 2023-11-21 之江实验室 一种参考外部知识的指代消解方法及装置
CN117252258B (zh) * 2023-07-05 2025-11-11 浙江点创信息科技有限公司 基于IE-Triple的知识图谱构建方法
CN116932660B (zh) * 2023-07-11 2026-04-07 中国电信股份有限公司技术创新中心 元数据关系提取的建模方法、提取方法及相关设备
CN117033420B (zh) * 2023-10-09 2024-01-09 之江实验室 一种知识图谱同概念下实体数据可视化展示方法及装置
CN117093728B (zh) * 2023-10-19 2024-02-02 杭州同花顺数据开发有限公司 一种金融领域事理图谱构建方法、装置、设备及存储介质
CN117371534B (zh) * 2023-12-07 2024-02-27 同方赛威讯信息技术有限公司 一种基于bert的知识图谱构建方法及系统
CN117874212B (zh) * 2024-01-02 2024-07-09 中国司法大数据研究院有限公司 一种基于案情标签提取的法条推荐方法和装置
CN117540035B (zh) * 2024-01-09 2024-05-14 安徽思高智能科技有限公司 一种基于实体类型信息融合的rpa知识图谱构建方法
CN118070886B (zh) * 2024-04-19 2024-07-30 南昌工程学院 一种水库防洪应急预案知识图谱构建方法及系统
CN118504551B (zh) * 2024-05-10 2024-11-22 中国传媒大学 一种新闻人物的言论抽取方法、设备及介质
CN118396118B (zh) * 2024-05-11 2025-06-03 上海云阙智能科技有限公司 大语言模型输出幻觉矫正方法、系统、介质、电子设备
CN118428369B (zh) * 2024-05-17 2025-02-07 南京邮电大学 一种实体识别和关系抽取方法
CN119149829B (zh) * 2024-11-15 2025-03-07 江西跃山科技有限责任公司 一种基于知识图谱的启蒙图书推荐系统
CN119719381B (zh) * 2024-11-23 2025-10-31 北京计算机技术及应用研究所 一种借助知识图谱生成关系型数据的方法

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108874878A (zh) * 2018-05-03 2018-11-23 众安信息技术服务有限公司 一种知识图谱的构建系统及方法
CN109977237A (zh) * 2019-05-27 2019-07-05 南京擎盾信息科技有限公司 一种面向法律领域的动态法律事件图谱构建方法
CN110781254A (zh) * 2020-01-02 2020-02-11 四川大学 一种案情知识图谱自动构建方法及系统及设备及介质
US20200073933A1 (en) * 2018-08-29 2020-03-05 National University Of Defense Technology Multi-triplet extraction method based on entity-relation joint extraction model
CN111858940A (zh) * 2020-07-27 2020-10-30 湘潭大学 一种基于多头注意力的法律案例相似度计算方法及系统
CN113204649A (zh) * 2021-05-11 2021-08-03 西安交通大学 基于实体关系联合抽取的法律知识图谱构建方法及设备

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8805861B2 (en) * 2008-12-09 2014-08-12 Google Inc. Methods and systems to train models to extract and integrate information from data sources
CN109902171B (zh) * 2019-01-30 2020-12-25 中国地质大学(武汉) 基于分层知识图谱注意力模型的文本关系抽取方法及系统

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108874878A (zh) * 2018-05-03 2018-11-23 众安信息技术服务有限公司 一种知识图谱的构建系统及方法
US20200073933A1 (en) * 2018-08-29 2020-03-05 National University Of Defense Technology Multi-triplet extraction method based on entity-relation joint extraction model
CN109977237A (zh) * 2019-05-27 2019-07-05 南京擎盾信息科技有限公司 一种面向法律领域的动态法律事件图谱构建方法
CN110781254A (zh) * 2020-01-02 2020-02-11 四川大学 一种案情知识图谱自动构建方法及系统及设备及介质
CN111858940A (zh) * 2020-07-27 2020-10-30 湘潭大学 一种基于多头注意力的法律案例相似度计算方法及系统
CN113204649A (zh) * 2021-05-11 2021-08-03 西安交通大学 基于实体关系联合抽取的法律知识图谱构建方法及设备

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
ZHEPEI WEI; JIANLIN SU; YUE WANG; YUAN TIAN; YI CHANG: "A Novel Cascade Binary Tagging Framework for Relational Triple Extraction", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 22 June 2020 (2020-06-22), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081679963 *

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116361477A (zh) * 2022-11-28 2023-06-30 中国电力科学研究院有限公司 信息驱动的电网运行态势知识图谱智能构建系统及方法
CN116680392A (zh) * 2023-06-05 2023-09-01 北京沃东天骏信息技术有限公司 一种关系三元组的抽取方法和装置
CN116992042A (zh) * 2023-07-14 2023-11-03 珠海中科先进技术研究院有限公司 基于新型研发机构科技创新服务知识图谱系统的构建方法
CN116955560A (zh) * 2023-07-21 2023-10-27 广州拓尔思大数据有限公司 基于思考链和知识图谱的数据处理方法及系统
CN116955560B (zh) * 2023-07-21 2024-01-05 广州拓尔思大数据有限公司 基于思考链和知识图谱的数据处理方法及系统
CN117236432A (zh) * 2023-09-26 2023-12-15 中国科学院沈阳自动化研究所 一种面向多模态数据的制造工艺知识图谱构建方法及系统
CN117033666A (zh) * 2023-10-07 2023-11-10 之江实验室 一种多模态知识图谱的构建方法、装置、存储介质及设备
CN117033666B (zh) * 2023-10-07 2024-01-26 之江实验室 一种多模态知识图谱的构建方法、装置、存储介质及设备
CN118297065A (zh) * 2024-03-01 2024-07-05 华中科技大学 一种低资源场景下基于提示学习的关系抽取方法

Also Published As

Publication number Publication date
CN113204649A (zh) 2021-08-03
US20230196127A1 (en) 2023-06-22
US12530597B2 (en) 2026-01-20

Similar Documents

Publication Publication Date Title
WO2022237013A1 (zh) 基于实体关系联合抽取的法律知识图谱构建方法及设备
CN111125331A (zh) 语义识别方法、装置、电子设备及计算机可读存储介质
CN110309511B (zh) 基于共享表示的多任务语言分析系统及方法
CN112100348A (zh) 一种多粒度注意力机制的知识库问答关系检测方法及系统
WO2023184633A1 (zh) 一种中文拼写纠错方法及系统、存储介质及终端
CN110321563A (zh) 基于混合监督模型的文本情感分析方法
TW201403354A (zh) 以資料降維法及非線性算則建構中文文本可讀性數學模型之系統及其方法
CN117577254A (zh) 医疗领域语言模型构建及电子病历文本结构化方法、系统
US20220129768A1 (en) Method and apparatus for training model, and method and apparatus for predicting text
CN116304748B (zh) 一种文本相似度计算方法、系统、设备及介质
CN113886601A (zh) 电子文本事件抽取方法、装置、设备及存储介质
CN115359799A (zh) 语音识别方法、训练方法、装置、电子设备及存储介质
CN117273012A (zh) 电力知识语义分析系统及方法
CN113221564B (zh) 训练实体识别模型的方法、装置、电子设备和存储介质
WO2024021334A1 (zh) 关系抽取方法、计算机设备及程序产品
TW201905734A (zh) 語意分析裝置、方法及其電腦程式產品
CN114117189B (zh) 一种问题解析方法、装置、电子设备及存储介质
WO2024087297A1 (zh) 文本情感分析方法、装置、电子设备及存储介质
CN116756267A (zh) 事件时序关系识别方法、装置、设备及介质
CN115062146A (zh) 基于BiLSTM结合多头注意力的中文重叠事件抽取系统
CN103530280A (zh) 以数据降维法及非线性算则建构中文文本可读性模型的系统及其方法
CN119226513A (zh) 基于大模型的文本处理方法、装置、设备、介质、程序产品及智能体
CN119295790A (zh) 一种零样本图像分类系统及方法
CN119090010B (zh) 一种人工智能交互的法律文本生成方法、系统及设备
Ruan et al. DISCERN: Chain-of-Thought-Augmented Syntactic-based In-Context Learning for Chinese Semantic Error Detection

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21941580

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21941580

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 21941580

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 24/06/2024)

122 Ep: pct application non-entry in european phase

Ref document number: 21941580

Country of ref document: EP

Kind code of ref document: A1