WO2022237013A1 - 基于实体关系联合抽取的法律知识图谱构建方法及设备 - Google Patents
基于实体关系联合抽取的法律知识图谱构建方法及设备 Download PDFInfo
- Publication number
- WO2022237013A1 WO2022237013A1 PCT/CN2021/116053 CN2021116053W WO2022237013A1 WO 2022237013 A1 WO2022237013 A1 WO 2022237013A1 CN 2021116053 W CN2021116053 W CN 2021116053W WO 2022237013 A1 WO2022237013 A1 WO 2022237013A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- entity
- relationship
- model
- extraction
- text
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/36—Creation of semantic tools, e.g. ontology or thesauri
- G06F16/367—Ontology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/02—Knowledge representation; Symbolic representation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/211—Syntactic parsing, e.g. based on context-free grammar [CFG] or unification grammars
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
- G06F40/295—Named entity recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/191—Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06V30/19127—Extracting features by transforming the feature space, e.g. multidimensional scaling; Mappings, e.g. subspace methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/191—Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06V30/19187—Graphical models, e.g. Bayesian networks or Markov models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/29—Graphical models, e.g. Bayesian networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/042—Knowledge-based neural networks; Logical representations of neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/02—Knowledge representation; Symbolic representation
- G06N5/022—Knowledge engineering; Knowledge acquisition
Definitions
- the invention belongs to the field of electronic information, and in particular relates to a method and equipment for constructing a legal knowledge graph based on joint extraction of entity relationships.
- Knowledge Graph known as knowledge domain visualization or knowledge domain mapping map in the library and information industry, is a series of different graphics showing the relationship between knowledge development process and structure, using visualization technology to describe knowledge resources and their carriers, mining , analyze, construct, map and display knowledge and their interconnections.
- the knowledge map is a combination of theories and methods of applied mathematics, graphics, information visualization technology, information science and other disciplines with metrology citation analysis, co-occurrence analysis and other methods, and uses the visual map to vividly display the core structure and development of the subject. History, frontier fields, and modern theories of multidisciplinary integration can provide practical and valuable references for disciplinary research.
- Knowledge graphs are divided into general knowledge graphs and domain knowledge graphs. These two knowledge graphs mainly differ in coverage and usage.
- the general knowledge graph is oriented to the general field and mainly contains a large amount of common sense knowledge in the real world, covering a wide range.
- Domain knowledge graph also known as industry knowledge graph or vertical knowledge graph, is oriented to a specific field and is an industry knowledge base composed of professional data in this field. Because it is constructed based on industry data, it has strict and rich data patterns. Therefore, there are higher requirements for the depth and accuracy of knowledge in this field.
- Domain knowledge graphs are knowledge graphs for specific domains, such as e-commerce, finance, and medical care. In comparison, domain knowledge graphs have more knowledge sources, faster scale expansion requirements, more complex knowledge structures, higher knowledge quality requirements, and wider application forms of knowledge.
- Entity extraction technology also known as named entity recognition technology, refers to the identification of entities with specific meaning in the extracted text, mainly including names of people, places, institutions, proper nouns, etc., as well as time, quantity, currency, and proportional values.
- the main task of relation extraction is to extract the relation between entities in the text.
- the relationship between entities and entities is formally described as the relationship of triples ⁇ h, r, t>, where h, t represent the head entity and tail entity, and r represents the relationship between entities.
- the purpose of the present invention is to solve the problem that the accuracy of knowledge map construction in the legal field cannot be guaranteed in the prior art, and provide a method and device for constructing a legal knowledge map based on joint extraction of entity relationships, so as to obtain a knowledge map with high accuracy.
- the present invention has the following technical solutions:
- a method for constructing a legal knowledge map based on entity-relationship joint extraction comprising the following steps:
- the model architecture includes model encoding layer, head entity extraction layer and relationship-tail entity extraction layer;
- model encoding layer uses the bert pre-training model
- the head entity extraction layer uses two BiLSTMs as binary classifiers, and uses text encoding as the input of the classifier.
- the output of the entity starting position corresponding to the first BiLSTM binary classifier is 1, and the output of the other positions is 0.
- the output of the end position of the entity corresponding to the second BiLSTM binary classifier is 1, and the output of other positions is 0;
- the relationship-tail entity extraction layer combines the encoding information of the head entity with the encoding of the sentence as input, and for each head entity, finds the tail entities that may exist under each relationship, and finally obtains a complete triple;
- the subject and the object are used as the head entity and the tail entity respectively, and the predicate is used as the relationship.
- the relationship set is determined according to the marked relationships, and the relationships with the same or similar semantics are merged.
- the head entity extraction layer uses the feature vector x i output by the bert coding layer as input, and outputs the start and end signs extracted to the entity;
- xi is the feature vector of each word
- W s , W e are the weight matrices that two classifiers can train
- b s , be e are the respective bias vectors; It is the sign of the starting position of the entity. When its value is close to 1, it means that the position is the starting position of the entity; It is the flag of the end position of the entity. When its value is close to 1, it means that the position is the end position of the entity.
- the relationship-tail entity extraction layer is composed of two BiLSTMs with the same structure.
- the input of this layer model includes the feature vector h s of the sentence, and incorporates The header entity code extracted by the previous layer where k represents the kth head entity;
- the present invention also provides a legal knowledge map construction system based on entity relationship joint extraction, including:
- the triplet data set building module is used to split legal text sentences into short sentences, and can complete the default subject in short sentences, and finally extract triplets from short sentences to construct triplet data set;
- the model building and training module is used to separately construct the model encoding layer, the head entity extraction layer and the relationship-tail entity extraction layer in the model architecture, and obtain a model capable of extracting triples through training;
- the inter-sentence relationship judgment module is used to judge the relationship between each short sentence for the legal text sentence that has not been split into short sentences;
- the knowledge map visualization module is used to combine the triples extracted from the model with the relationship between sentences in the text to obtain the compound triples corresponding to the legal text, and to realize the visualization of the legal knowledge map.
- the present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and operable on the processor.
- a terminal device including a memory, a processor, and a computer program stored in the memory and operable on the processor.
- the processor executes the computer program, the The steps of the legal knowledge graph construction method based on entity-relationship joint extraction.
- the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a legal knowledge map based on joint entity-relationship extraction are implemented .
- the present invention has the following beneficial effects:
- the present invention uses a model based on the joint extraction of entity relationships to extract triples from unstructured text in the field of contract law, and finally builds a knowledge map in the legal field.
- the present invention can avoid the wrong transmission caused by the pipeline method, and the accuracy rate high.
- the design of the model framework of the present invention adopts the Chinese bert pre-training model as an encoder, and has a good effect on encoding Chinese text.
- the entity extraction part uses two BiLSTM binary classifiers to identify the start position and end position of entities, which can effectively extract entities in the form of phrases in the text.
- the present invention first extracts the head entity, and then extracts the tail entity corresponding to the entity relationship from the extracted head entity.
- entity relations and tail entities not only the encoding information of the sentence is used, but also the encoding information of the head entity is incorporated.
- the present invention can obtain a legal knowledge map with high accuracy, and the constructed knowledge map can be combined with deep learning technology to realize functions such as question-and-answer reasoning and related recommendations in the field of contract law.
- Fig. 1 is a flowchart of a method for constructing a contract law knowledge graph according to an embodiment of the present invention
- Fig. 2 is a sample diagram of a data set constructed by model training in an embodiment of the present invention
- Fig. 3 is a model architecture diagram of extracting triples from text in an embodiment of the present invention.
- Fig. 4 is a schematic diagram of the training and use of the extracted triplet model in the embodiment of the present invention.
- Fig. 5 is a schematic diagram of a visualized knowledge map constructed by an embodiment of the present invention.
- the present invention proposes a legal knowledge map construction method based on the joint extraction of entity relationships.
- the embodiment is described by taking the contract law as an example.
- the present invention can use the given contract law text to perform entity extraction and relationship extraction at the same time, and finally obtain a complete triplet information.
- the contract law knowledge map can be formed by connecting the extracted triples end to end.
- the completed knowledge graph can be combined with deep learning technology to implement functions such as question-and-answer reasoning and related recommendations in the field of contract law.
- Entity extraction For any complete contract law text statement, it can be decomposed into the form of (h, r, t), where h represents the head entity, r represents the entity relationship, and t represents the tail entity. Entity extraction means extracting the head entity and tail entity in the text.
- Relationship extraction The relationship here refers to the relationship between entities or the attributes of entities. This step usually extracts the corresponding entity relationship after the head entity and tail entity are extracted.
- Joint extraction Unlike the previous entity extraction and relationship extraction, which are performed independently, the entities and relationships extracted by joint extraction affect each other. Using joint extraction can reduce the error transmission problem caused by entity extraction.
- the embodiment of the present invention is based on the method for constructing the contract law map based on the joint extraction of entity relationships, including the following steps:
- Step 1 Construction of contract law triple data set, including:
- the split text sentences obtained in the first step will have the phenomenon of missing subjects in some short sentences, which will affect the subsequent triplet extraction work, so the default subject needs to be completed.
- This method uses the open source tool pyltp combined with the method of dependency parsing to parse the default part and complete the default subject.
- the example in the first step after the completion of the subject results in short sentences "the parties conclude the contract” and "(the parties) should have the corresponding capacity for civil rights and capacity for civil conduct”.
- a complete sentence usually consists of three parts: subject, predicate, and object. Therefore, when labeling data, the subject and object of the sentence are used as the head entity and tail entity respectively, and the predicate is used as the initial labeling of the relationship.
- After labeling some triples determine the relationship set according to the labeled relationship, and merge the relationships with the same or similar semantics. For example, the relationship "concept” and the relationship “definition” have similar semantics, and the relationship "definition” is unified into the relationship "concept”.
- Figure 2 is an example of a partially labeled triplet. In the end, 798 artificially calibrated triples were obtained, and the entity-relationship set contained 25 relationships.
- Model architecture design and model training including:
- the design of the experimental model mainly considers the following aspects: First, the bert pre-training model uses a bidirectional Transformer, and at the same time uses the Masked Language Model in the pre-training process ( MLM) captures the representation at the word level, which makes the word vector change from only containing the previous information to the information that can learn the context. In the pre-training process, Next Sentence Prediction (NSP) is used to capture the sentence-level representation. Therefore, using the bert pre-training model at the encoding layer can better represent the deep meaning of the sentence. Second, the entities in the text of the contract law are different from those in the general field.
- MLM Masked Language Model in the pre-training process
- NSP Next Sentence Prediction
- the input of the entity relationship and tail entity extraction part of the model is not only the encoding information of the entire sentence, but the encoding information of the head entity and the sentence The combination of encoding, which has a good effect on the extraction of entity relations and tail entities.
- model framework for the joint extraction of entity relations in the field of contract law is designed, and the model framework is shown in the figure.
- the model is divided into three parts, namely the model coding layer, the head entity extraction layer, and the relationship-tail entity extraction layer.
- the model architecture diagram refers to Figure 3. The specific content of each part is as follows:
- the model coding layer of the present invention adopts the Chinese pre-training model BERT-wwm-ext based on the full-word Mask on a larger-scale corpus by the Harbin Institute of Technology Xunfei Joint Laboratory. This model has obtained further performance improvements in multiple benchmark tests. . Using this model, the input text can be converted into the form of feature vectors.
- This layer is mainly composed of two BiLSTMs with the same structure, with the feature vector x i output by the bert encoding layer as input, and the output is extracted to the start and end marks of the entity.
- W s , W e are the weight matrices that two classifiers can train
- b s be are their respective bias vectors. It is the sign of the starting position of the entity. When its value is close to 1, it means that the position is the starting position of the entity; It is the sign of the end position of the entity. When its value is close to 1, it means that the position is the end position of the entity; in the example in the figure, for the text "the parties should have the corresponding capacity for civil rights and capacity for civil conduct", the head entity is extracted
- the entities that can be extracted by the layer are "parties" and "with corresponding capacity for civil rights and capacity for civil conduct”. The start and end positions of the entity.
- This layer is similar to the head entity extraction layer, and is also composed of two BiLSTMs with the same structure.
- the input of this layer model is no longer just the sentence feature vector h s , it also incorporates the head entity code extracted by the previous layer.
- k represents the kth head entity. Will as the input vector for this layer.
- the specific formula is:
- the middle part is the training process of the model.
- the input of the model is the text of the contract law.
- the Bert encoding layer After passing through the Bert encoding layer, the head entity extraction layer, and the relationship-tail entity extraction layer to obtain the triplet output by the model, and then use the given
- the fixed loss function is continuously iteratively optimized, and when the value of the loss function tends to be stable, the iteration is stopped, the training of the model is completed, and the trained model is saved.
- Step 3 Judging the relationship between text sentences.
- Inter-sentence relations include four kinds of relations: "condition”, “turning”, “parallel” and “cause and effect”. Among them, 85 cases of causal relations, 194 cases of conditional relations, 34 cases of turning relations and 8 cases of parallel relations were obtained.
- the parties shall have the corresponding capacity for civil rights and capacity for civil conduct" when entering into a contract, the "conditional relationship" of the inter-sentence relationship extracted from the two short sentences.
- Step 4 Triple composition and map visualization, including:
- the compound triples corresponding to the legal text of the contract can be obtained.
- the final triplet form that can be obtained is ((parties, conclude, contract), condition, (parties, should, have the corresponding capacity for civil rights and capacity for civil conduct)).
- all the extracted triples are integrated and spliced to obtain a complete knowledge map of contract law.
- FIG. 5 is a partial schematic diagram of the visualized contract law knowledge map.
- the completed contract law knowledge graph can be combined with deep learning technology to realize the functions of question-and-answer reasoning and related recommendations in the field of contract law.
- a legal knowledge graph construction system based on entity-relationship joint extraction including:
- the triplet data set building module is used to split legal text sentences into short sentences, and can complete the default subject in short sentences, and finally extract triplets from short sentences to construct triplet data set;
- the model building and training module is used to separately construct the model encoding layer, the head entity extraction layer and the relationship-tail entity extraction layer in the model architecture, and obtain a model capable of extracting triples through training;
- the inter-sentence relationship judgment module is used to judge the relationship between each short sentence for the legal text sentence that has not been split into short sentences;
- the knowledge map visualization module is used to combine the triples extracted from the model with the relationship between sentences in the text to obtain the compound triples corresponding to the legal text, and to realize the visualization of the legal knowledge map.
- a terminal device comprising a memory, a processor, and a computer program stored in the memory and operable on the processor, when the processor executes the computer program, the joint extraction based on entity relationship is realized.
- a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a legal knowledge map based on entity-relationship joint extraction are realized.
- the computer program can be divided into one or more modules/units, and the one or more modules/units are stored in the memory and executed by the processor to complete the construction of the knowledge map of the present invention method.
- the processor can be a central processing unit (Central Processing Unit, CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), off-the-shelf programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- the memory can be used to store computer programs and/or modules, and the processor implements the knowledge graph construction system of the present invention by running or executing the computer programs and/or modules stored in the memory, and calling the data stored in the memory. various functions.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Artificial Intelligence (AREA)
- Computational Linguistics (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Mathematical Physics (AREA)
- Computing Systems (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Multimedia (AREA)
- Animal Behavior & Ethology (AREA)
- Databases & Information Systems (AREA)
- Machine Translation (AREA)
Abstract
Description
Claims (8)
- 一种基于实体关系联合抽取的法律知识图谱构建方法,其特征在于,包括以下步骤:-三元组数据集的构建;将法律文本语句拆分成短句的形式;将短句中缺省的主语补全;从短句中抽取三元组,构建三元组数据集;-模型架构的设计和模型的训练;模型架构包括模型编码层、头实体抽取层以及关系-尾实体抽取层;具体的,模型编码层使用bert预训练模型;头实体抽取层使用两个BiLSTM作为二分类器,将文本的编码作为分类器的输入,输出信息中,第一个BiLSTM二分类器对应的实体起始位置输出为1,其余位置输出都为0,第二个BiLSTM二分类器对应的实体结束位置输出为1,其余位置输出都为0;关系-尾实体抽取层,将头实体的编码信息与句子的编码相结合作为输入,对于每个头实体,找到每个关系下可能存在的尾实体,最终得到完整的三元组;-文本句间关系判断;对于未进行短句拆分的法律文本语句,判断各个短句之间的关系;-三元组复合与图谱可视化;根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组;法律知识图谱的可视化。
- 根据权利要求1所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:构建三元组数据集时将主语和宾语分别作为头实体和尾实体,将谓语作为关系。
- 根据权利要求2所述基于实体关系联合抽取的法律知识图谱构建方法,其特征在于:根据标注的关系确定关系集合,合并语义相同或相似的关系。
- 一种基于实体关系联合抽取的法律知识图谱构建系统,其特征在于,包括:三元组数据集构建模块,用于将法律文本语句拆分成短句的形式,并能够将短句中缺省的主语补全,最终从短句中抽取三元组,构建三元组数据集;模型建立和训练模块,用于对模型架构中的模型编码层、头实体抽取层以及关系-尾实体抽取层分别构建,经过训练得到能够抽取出三元组的模型;句间关系判断模块,用于对未进行短句拆分的法律文本语句,判断各个短句之间的关系;知识图谱可视化模块,用于根据模型抽取出的三元组结合文本句间关系,得到法律文本对应的复合三元组,并实现法律知识图谱的可视化。
- 一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机程序,其特征在于:所述的处理器执行所述的计算机程序时实现如权利要求1至5中任意一项所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,其特征在于:所述的计算机程序被处理器执行时实现如权利要求1至5中任意一项所述基于实体关系联合抽取的法律知识图谱构建方法的步骤。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/956,864 US12530597B2 (en) | 2021-05-11 | 2022-09-30 | Method and device for constructing legal knowledge graph based on joint entity and relation extraction |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202110513432.2A CN113204649A (zh) | 2021-05-11 | 2021-05-11 | 基于实体关系联合抽取的法律知识图谱构建方法及设备 |
| CN202110513432.2 | 2021-05-11 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/956,864 Continuation US12530597B2 (en) | 2021-05-11 | 2022-09-30 | Method and device for constructing legal knowledge graph based on joint entity and relation extraction |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022237013A1 true WO2022237013A1 (zh) | 2022-11-17 |
Family
ID=77030936
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/116053 Ceased WO2022237013A1 (zh) | 2021-05-11 | 2021-09-01 | 基于实体关系联合抽取的法律知识图谱构建方法及设备 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12530597B2 (zh) |
| CN (1) | CN113204649A (zh) |
| WO (1) | WO2022237013A1 (zh) |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116361477A (zh) * | 2022-11-28 | 2023-06-30 | 中国电力科学研究院有限公司 | 信息驱动的电网运行态势知识图谱智能构建系统及方法 |
| CN116680392A (zh) * | 2023-06-05 | 2023-09-01 | 北京沃东天骏信息技术有限公司 | 一种关系三元组的抽取方法和装置 |
| CN116955560A (zh) * | 2023-07-21 | 2023-10-27 | 广州拓尔思大数据有限公司 | 基于思考链和知识图谱的数据处理方法及系统 |
| CN116992042A (zh) * | 2023-07-14 | 2023-11-03 | 珠海中科先进技术研究院有限公司 | 基于新型研发机构科技创新服务知识图谱系统的构建方法 |
| CN117033666A (zh) * | 2023-10-07 | 2023-11-10 | 之江实验室 | 一种多模态知识图谱的构建方法、装置、存储介质及设备 |
| CN117236432A (zh) * | 2023-09-26 | 2023-12-15 | 中国科学院沈阳自动化研究所 | 一种面向多模态数据的制造工艺知识图谱构建方法及系统 |
| CN118297065A (zh) * | 2024-03-01 | 2024-07-05 | 华中科技大学 | 一种低资源场景下基于提示学习的关系抽取方法 |
Families Citing this family (35)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111291185B (zh) * | 2020-01-21 | 2023-09-22 | 京东方科技集团股份有限公司 | 信息抽取方法、装置、电子设备及存储介质 |
| CN113204649A (zh) * | 2021-05-11 | 2021-08-03 | 西安交通大学 | 基于实体关系联合抽取的法律知识图谱构建方法及设备 |
| CN113779260B (zh) * | 2021-08-12 | 2023-07-18 | 华东师范大学 | 一种基于预训练模型的领域图谱实体和关系联合抽取方法及系统 |
| CN114004223B (zh) * | 2021-10-12 | 2022-05-24 | 北京理工大学 | 一种基于行为基的事件知识表示方法 |
| CN114118056A (zh) * | 2021-10-13 | 2022-03-01 | 中国人民解放军军事科学院国防工程研究院工程防护研究所 | 一种战争类研究报告的信息抽取方法 |
| CN113962214B (zh) * | 2021-10-25 | 2024-07-16 | 东南大学 | 基于eletric-bert的实体抽取方法 |
| CN114138979B (zh) * | 2021-10-29 | 2022-09-16 | 中南民族大学 | 基于词拓展无监督文本分类的文物安全知识图谱创建方法 |
| CN114357163B (zh) * | 2021-11-30 | 2025-09-02 | 腾讯科技(深圳)有限公司 | 文本类型识别方法、装置、计算机可读介质及电子设备 |
| CN115358227A (zh) * | 2022-04-13 | 2022-11-18 | 苏州空天信息研究院 | 一种基于短语增强的开放域关系联合抽取方法及系统 |
| CN114898384A (zh) * | 2022-06-01 | 2022-08-12 | 北京金山数字娱乐科技有限公司 | 合同识别方法及装置 |
| CN115168599B (zh) * | 2022-06-20 | 2023-06-20 | 北京百度网讯科技有限公司 | 多三元组抽取方法、装置、设备、介质及产品 |
| CN115114382B (zh) * | 2022-07-01 | 2024-09-06 | 沈阳航空航天大学 | 基于预训练模型与规则结合的武器装备实体关系抽取方法 |
| CN115495585B (zh) * | 2022-08-31 | 2026-01-09 | 上海海洋大学 | 一种基于知识图谱的花卉病虫害的本体建模方法和建模系统 |
| CN115757815A (zh) * | 2022-11-04 | 2023-03-07 | 北京中科凡语科技有限公司 | 知识图谱的构建方法、装置及存储介质 |
| CN115730078A (zh) * | 2022-11-04 | 2023-03-03 | 南京擎盾信息科技有限公司 | 用于类案检索的事件知识图谱构建方法、装置及电子设备 |
| CN116049422B (zh) * | 2022-12-07 | 2026-03-06 | 安徽大学 | 基于联合抽取模型的包虫病知识图谱构建方法及其应用 |
| CN116150404A (zh) * | 2023-03-03 | 2023-05-23 | 成都康赛信息技术有限公司 | 一种基于联合学习的教育资源多模态知识图谱构建方法 |
| CN116226408B (zh) * | 2023-03-27 | 2023-12-19 | 中国科学院空天信息创新研究院 | 农产品生长环境知识图谱构建方法及装置、存储介质 |
| CN116049148B (zh) * | 2023-04-03 | 2023-07-18 | 中国科学院成都文献情报中心 | 一种元出版环境下领域元知识引擎的构建方法 |
| CN117009478A (zh) * | 2023-06-25 | 2023-11-07 | 武汉光庭信息技术股份有限公司 | 一种基于软件知识图谱问答问句解析过程的算法融合方法 |
| CN116501830B (zh) * | 2023-06-29 | 2023-09-05 | 中南大学 | 一种生物医学文本的重叠关系联合抽取方法及相关设备 |
| CN116562303B (zh) * | 2023-07-04 | 2023-11-21 | 之江实验室 | 一种参考外部知识的指代消解方法及装置 |
| CN117252258B (zh) * | 2023-07-05 | 2025-11-11 | 浙江点创信息科技有限公司 | 基于IE-Triple的知识图谱构建方法 |
| CN116932660B (zh) * | 2023-07-11 | 2026-04-07 | 中国电信股份有限公司技术创新中心 | 元数据关系提取的建模方法、提取方法及相关设备 |
| CN117033420B (zh) * | 2023-10-09 | 2024-01-09 | 之江实验室 | 一种知识图谱同概念下实体数据可视化展示方法及装置 |
| CN117093728B (zh) * | 2023-10-19 | 2024-02-02 | 杭州同花顺数据开发有限公司 | 一种金融领域事理图谱构建方法、装置、设备及存储介质 |
| CN117371534B (zh) * | 2023-12-07 | 2024-02-27 | 同方赛威讯信息技术有限公司 | 一种基于bert的知识图谱构建方法及系统 |
| CN117874212B (zh) * | 2024-01-02 | 2024-07-09 | 中国司法大数据研究院有限公司 | 一种基于案情标签提取的法条推荐方法和装置 |
| CN117540035B (zh) * | 2024-01-09 | 2024-05-14 | 安徽思高智能科技有限公司 | 一种基于实体类型信息融合的rpa知识图谱构建方法 |
| CN118070886B (zh) * | 2024-04-19 | 2024-07-30 | 南昌工程学院 | 一种水库防洪应急预案知识图谱构建方法及系统 |
| CN118504551B (zh) * | 2024-05-10 | 2024-11-22 | 中国传媒大学 | 一种新闻人物的言论抽取方法、设备及介质 |
| CN118396118B (zh) * | 2024-05-11 | 2025-06-03 | 上海云阙智能科技有限公司 | 大语言模型输出幻觉矫正方法、系统、介质、电子设备 |
| CN118428369B (zh) * | 2024-05-17 | 2025-02-07 | 南京邮电大学 | 一种实体识别和关系抽取方法 |
| CN119149829B (zh) * | 2024-11-15 | 2025-03-07 | 江西跃山科技有限责任公司 | 一种基于知识图谱的启蒙图书推荐系统 |
| CN119719381B (zh) * | 2024-11-23 | 2025-10-31 | 北京计算机技术及应用研究所 | 一种借助知识图谱生成关系型数据的方法 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108874878A (zh) * | 2018-05-03 | 2018-11-23 | 众安信息技术服务有限公司 | 一种知识图谱的构建系统及方法 |
| CN109977237A (zh) * | 2019-05-27 | 2019-07-05 | 南京擎盾信息科技有限公司 | 一种面向法律领域的动态法律事件图谱构建方法 |
| CN110781254A (zh) * | 2020-01-02 | 2020-02-11 | 四川大学 | 一种案情知识图谱自动构建方法及系统及设备及介质 |
| US20200073933A1 (en) * | 2018-08-29 | 2020-03-05 | National University Of Defense Technology | Multi-triplet extraction method based on entity-relation joint extraction model |
| CN111858940A (zh) * | 2020-07-27 | 2020-10-30 | 湘潭大学 | 一种基于多头注意力的法律案例相似度计算方法及系统 |
| CN113204649A (zh) * | 2021-05-11 | 2021-08-03 | 西安交通大学 | 基于实体关系联合抽取的法律知识图谱构建方法及设备 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8805861B2 (en) * | 2008-12-09 | 2014-08-12 | Google Inc. | Methods and systems to train models to extract and integrate information from data sources |
| CN109902171B (zh) * | 2019-01-30 | 2020-12-25 | 中国地质大学(武汉) | 基于分层知识图谱注意力模型的文本关系抽取方法及系统 |
-
2021
- 2021-05-11 CN CN202110513432.2A patent/CN113204649A/zh active Pending
- 2021-09-01 WO PCT/CN2021/116053 patent/WO2022237013A1/zh not_active Ceased
-
2022
- 2022-09-30 US US17/956,864 patent/US12530597B2/en active Active
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108874878A (zh) * | 2018-05-03 | 2018-11-23 | 众安信息技术服务有限公司 | 一种知识图谱的构建系统及方法 |
| US20200073933A1 (en) * | 2018-08-29 | 2020-03-05 | National University Of Defense Technology | Multi-triplet extraction method based on entity-relation joint extraction model |
| CN109977237A (zh) * | 2019-05-27 | 2019-07-05 | 南京擎盾信息科技有限公司 | 一种面向法律领域的动态法律事件图谱构建方法 |
| CN110781254A (zh) * | 2020-01-02 | 2020-02-11 | 四川大学 | 一种案情知识图谱自动构建方法及系统及设备及介质 |
| CN111858940A (zh) * | 2020-07-27 | 2020-10-30 | 湘潭大学 | 一种基于多头注意力的法律案例相似度计算方法及系统 |
| CN113204649A (zh) * | 2021-05-11 | 2021-08-03 | 西安交通大学 | 基于实体关系联合抽取的法律知识图谱构建方法及设备 |
Non-Patent Citations (1)
| Title |
|---|
| ZHEPEI WEI; JIANLIN SU; YUE WANG; YUAN TIAN; YI CHANG: "A Novel Cascade Binary Tagging Framework for Relational Triple Extraction", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 22 June 2020 (2020-06-22), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081679963 * |
Cited By (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116361477A (zh) * | 2022-11-28 | 2023-06-30 | 中国电力科学研究院有限公司 | 信息驱动的电网运行态势知识图谱智能构建系统及方法 |
| CN116680392A (zh) * | 2023-06-05 | 2023-09-01 | 北京沃东天骏信息技术有限公司 | 一种关系三元组的抽取方法和装置 |
| CN116992042A (zh) * | 2023-07-14 | 2023-11-03 | 珠海中科先进技术研究院有限公司 | 基于新型研发机构科技创新服务知识图谱系统的构建方法 |
| CN116955560A (zh) * | 2023-07-21 | 2023-10-27 | 广州拓尔思大数据有限公司 | 基于思考链和知识图谱的数据处理方法及系统 |
| CN116955560B (zh) * | 2023-07-21 | 2024-01-05 | 广州拓尔思大数据有限公司 | 基于思考链和知识图谱的数据处理方法及系统 |
| CN117236432A (zh) * | 2023-09-26 | 2023-12-15 | 中国科学院沈阳自动化研究所 | 一种面向多模态数据的制造工艺知识图谱构建方法及系统 |
| CN117033666A (zh) * | 2023-10-07 | 2023-11-10 | 之江实验室 | 一种多模态知识图谱的构建方法、装置、存储介质及设备 |
| CN117033666B (zh) * | 2023-10-07 | 2024-01-26 | 之江实验室 | 一种多模态知识图谱的构建方法、装置、存储介质及设备 |
| CN118297065A (zh) * | 2024-03-01 | 2024-07-05 | 华中科技大学 | 一种低资源场景下基于提示学习的关系抽取方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113204649A (zh) | 2021-08-03 |
| US20230196127A1 (en) | 2023-06-22 |
| US12530597B2 (en) | 2026-01-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022237013A1 (zh) | 基于实体关系联合抽取的法律知识图谱构建方法及设备 | |
| CN111125331A (zh) | 语义识别方法、装置、电子设备及计算机可读存储介质 | |
| CN110309511B (zh) | 基于共享表示的多任务语言分析系统及方法 | |
| CN112100348A (zh) | 一种多粒度注意力机制的知识库问答关系检测方法及系统 | |
| WO2023184633A1 (zh) | 一种中文拼写纠错方法及系统、存储介质及终端 | |
| CN110321563A (zh) | 基于混合监督模型的文本情感分析方法 | |
| TW201403354A (zh) | 以資料降維法及非線性算則建構中文文本可讀性數學模型之系統及其方法 | |
| CN117577254A (zh) | 医疗领域语言模型构建及电子病历文本结构化方法、系统 | |
| US20220129768A1 (en) | Method and apparatus for training model, and method and apparatus for predicting text | |
| CN116304748B (zh) | 一种文本相似度计算方法、系统、设备及介质 | |
| CN113886601A (zh) | 电子文本事件抽取方法、装置、设备及存储介质 | |
| CN115359799A (zh) | 语音识别方法、训练方法、装置、电子设备及存储介质 | |
| CN117273012A (zh) | 电力知识语义分析系统及方法 | |
| CN113221564B (zh) | 训练实体识别模型的方法、装置、电子设备和存储介质 | |
| WO2024021334A1 (zh) | 关系抽取方法、计算机设备及程序产品 | |
| TW201905734A (zh) | 語意分析裝置、方法及其電腦程式產品 | |
| CN114117189B (zh) | 一种问题解析方法、装置、电子设备及存储介质 | |
| WO2024087297A1 (zh) | 文本情感分析方法、装置、电子设备及存储介质 | |
| CN116756267A (zh) | 事件时序关系识别方法、装置、设备及介质 | |
| CN115062146A (zh) | 基于BiLSTM结合多头注意力的中文重叠事件抽取系统 | |
| CN103530280A (zh) | 以数据降维法及非线性算则建构中文文本可读性模型的系统及其方法 | |
| CN119226513A (zh) | 基于大模型的文本处理方法、装置、设备、介质、程序产品及智能体 | |
| CN119295790A (zh) | 一种零样本图像分类系统及方法 | |
| CN119090010B (zh) | 一种人工智能交互的法律文本生成方法、系统及设备 | |
| Ruan et al. | DISCERN: Chain-of-Thought-Augmented Syntactic-based In-Context Learning for Chinese Semantic Error Detection |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21941580 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21941580 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21941580 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 24/06/2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21941580 Country of ref document: EP Kind code of ref document: A1 |