WO2021073259A1 - 基于知识驱动和数据驱动的分群方法及相关设备 - Google Patents
基于知识驱动和数据驱动的分群方法及相关设备 Download PDFInfo
- Publication number
- WO2021073259A1 WO2021073259A1 PCT/CN2020/111413 CN2020111413W WO2021073259A1 WO 2021073259 A1 WO2021073259 A1 WO 2021073259A1 CN 2020111413 W CN2020111413 W CN 2020111413W WO 2021073259 A1 WO2021073259 A1 WO 2021073259A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- knowledge
- target
- node
- processing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/02—Knowledge representation; Symbolic representation
- G06N5/022—Knowledge engineering; Knowledge acquisition
- G06N5/025—Extracting rules from data
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/70—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02A—TECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
- Y02A90/00—Technologies having an indirect contribution to adaptation to climate change
- Y02A90/10—Information and communication technologies [ICT] supporting adaptation to climate change, e.g. for weather forecasting or climate simulation
Definitions
- This application relates to the field of artificial intelligence, in particular to a knowledge-driven and data-driven clustering method and related equipment.
- Personalized precision treatment is currently an important research topic in the field of medical treatment, and it is also a development direction that is worth looking forward to in the field of medical research.
- Personalized precision treatment mainly starts from the group and provides personalized treatment according to the individual differences of patients. In this process, the most difficult thing is how to identify different groups and give sufficient medical knowledge for the personalized treatment of patients in the group. support.
- Existing grouping systems often use grouping methods that are completely data-driven or knowledge-driven, but both methods have many drawbacks.
- the core idea is to learn the rules of prescription for doctors in the data set based on machine learning or deep learning algorithms, and apply the learned rules to similar patients.
- the clustering method based entirely on knowledge-driven, the core of which is to construct rules based on relevant authoritative knowledge such as existing guidelines and consensus, and obtain corresponding auxiliary decision-making schemes according to the rules.
- relevant authoritative knowledge such as existing guidelines and consensus
- auxiliary decision-making schemes according to the rules.
- the use of a data-driven clustering method does not guarantee that the medical rules learned by the algorithm conform to medical knowledge, and it is impossible to make personalized recommendations for different patients in the same group to assist in decision-making.
- This application provides a knowledge-driven and data-driven clustering method and related equipment, which are used to create a target knowledge clustering decision tree, determine that the population to be classified belongs to the first group with similar characteristics, create a data clustering decision tree, and group
- the steps of dividing the population to be classified into the first group with similar characteristics or the second group with similar characteristics improve the accuracy of grouping the group to be classified into groups with similar characteristics, thereby providing reasonable, standardized and comprehensive and refined evidence-based medicine for doctors' decision-making evidence.
- the first aspect of the embodiments of the present application provides a knowledge-driven and data-driven clustering method, including: creating an initial knowledge clustering decision tree based on medical rule data, and filling sample data into the branches of the initial knowledge clustering decision tree
- the end-knowledge child node obtains the target knowledge grouping decision tree.
- the medical rule data is used to indicate the first medical discipline knowledge data obtained from the rule extraction process of the archived medical literature, medical consensus, medical guidelines, and the clinical experience of experts.
- the first grouping rule is used to group clinically similar feature groups, and the sample data includes clinical outcome data corresponding to the clinically similar feature groups following the medical knowledge rules; acquiring feature data of the group to be classified , And group the feature data through the target knowledge grouping decision tree, obtain the processing plan corresponding to the target end knowledge sub-node, and identify the number of categories of the processing plan, and the target end knowledge sub-node is the target end knowledge sub-node.
- the node position of the feature data on the target knowledge grouping decision tree, and the processing scheme is used to indicate the solution channel for groups of similar characteristics; if the number of categories of the processing scheme is identified as one, the to-be-classified A group of people belongs to a group with similar first characteristics, and the group with similar first characteristics is a group with similar characteristics to which the processing plan corresponding to the target end knowledge sub-node in the target knowledge grouping decision tree belongs; if the number of categories of the processing plan is identified as At least two, the data grouping decision tree is generated based on the second grouping rule, and the target layer number of the data sub-node of the data grouping decision tree corresponding to each processing scheme is determined, and the second grouping rule indicates according to The classification is based on the effectiveness of the recursive partitioning algorithm and the processing scheme; if the target level is the first level, the group to be classified is divided into the first feature similar group; if the target level is not the first level Layer, the group to be classified is divided into a group with similar
- a second aspect of the embodiments of the present application provides a clustering device based on knowledge-driven and data-driven.
- the clustering device based on knowledge-driven and data-driven includes: a first creation unit for creating initial knowledge based on medical rule data Group decision tree, and fill the sample data into the end knowledge sub-nodes of each branch of the initial knowledge group decision tree to obtain the target knowledge group decision tree.
- the medical rule data is used to indicate the archived medical literature, medical consensus, The first grouping rule obtained by extracting the medical discipline knowledge data of the medical guide and the clinical experience of the expert.
- the first grouping rule is used for grouping people with similar characteristics in the clinic, and the sample data includes following the medical knowledge rule The clinical outcome data corresponding to the clinically similar feature population; the identification unit is used to obtain the feature data of the population to be classified, and group the feature data through the target knowledge clustering decision tree to obtain the target end knowledge sub-node correspondence The target end knowledge sub-node is the node position of the feature data on the target knowledge grouping decision tree, and the processing plan is used to indicate the similarity The solution channel for characteristic groups; the second creation unit is used to generate a data grouping decision tree based on the second grouping rule if the number of categories of the processing scheme is at least two, and to determine that each processing scheme corresponds to the The target layer number of the data sub-nodes of the data grouping decision tree, the second grouping rule indicates classification according to the effectiveness of the recursive partitioning algorithm and processing scheme; the first dividing unit is used to if the target layer number is At the first level, the group to be classified is divided into the first group with
- the third aspect of the embodiments of the present application provides a knowledge-driven and data-driven clustering device, including a memory, a processor, and a computer program stored in the memory and running on the processor.
- the processing When the computer program is executed by the computer program, the following steps of the knowledge-driven and data-driven clustering method are performed: an initial knowledge clustering decision tree is created according to medical rule data, and sample data is filled into each branch of the initial knowledge clustering decision tree
- the end knowledge sub-nodes of, obtain the target knowledge group decision tree, and the medical rule data is used to indicate the first medical discipline knowledge data obtained from the rule extraction process of the archived medical literature, medical consensus, medical guidelines, and the clinical experience of experts
- a grouping rule, the first grouping rule is used to group clinically similar feature groups, and the sample data includes clinical outcome data corresponding to the clinically similar feature groups that follow the medical knowledge rules; to obtain the characteristics of the group to be classified Data, and group the feature data through the target knowledge grouping decision tree to obtain the processing plan corresponding to the target
- the fourth aspect of the embodiments of the present application provides a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute the steps of the knowledge-driven and data-driven clustering method as described below: Create an initial knowledge grouping decision tree based on the medical rule data, and populate the sample data to the end knowledge sub-nodes of each branch of the initial knowledge grouping decision tree to obtain the target knowledge grouping decision tree.
- the medical rule data is used to indicate the archived decision tree.
- the first grouping rule obtained by extracting the medical literature, medical consensus, medical guidelines, and expert clinical experience of medical discipline knowledge data the first grouping rule is used to group clinically similar groups of people
- the The sample data includes the clinical outcome data corresponding to the clinically similar feature population that follows the medical knowledge rules; the feature data of the population to be classified is obtained, and the feature data is grouped through the target knowledge clustering decision tree to obtain the target end knowledge sub Node corresponding processing plan, and identify the number of categories of the processing plan, the target end knowledge sub-node is the node position of the feature data on the target knowledge grouping decision tree, and the processing plan is used to indicate A solution channel for groups of people with similar characteristics; if the number of categories of the processing solution is identified as one, it is determined that the group to be classified belongs to the first group with similar characteristics, and the first group with similar characteristics is the target knowledge grouping decision tree Groups with similar characteristics to which the processing schemes corresponding to the target end knowledge sub-nodes belong; if the number of categories of the processing schemes is identified as at
- a target knowledge grouping decision tree is created; the processing solution corresponding to the target end knowledge sub-node on the target knowledge grouping decision tree is identified by identifying the characteristic data of the group to be classified When the number of categories is identified as one, it is determined that the population to be classified belongs to the first feature similar population corresponding to the target end knowledge sub-node in the target knowledge grouping decision tree; when the number of categories is identified as at least When two, the data grouping decision tree is generated based on the preset second grouping rule, and the target level of the data grouping decision tree is determined; the group to be classified is divided into the first feature according to the target level A similar group of people or a group of people with similar second characteristics corresponding to the target data node in the data grouping decision tree.
- the target knowledge grouping decision tree is created, the group to be classified is determined to belong to the group with similar first characteristics, the data grouping decision tree is created, and the group to be classified is divided into the group with similar first characteristics or the group with similar second characteristics according to the number of target layers. Steps such as population, improve the accuracy of grouping the population to be classified into groups of similar characteristics, and then provide reasonable, standardized and comprehensive and refined evidence-based medical evidence for doctors' decision-making.
- FIG. 1 is a schematic diagram of an embodiment of a knowledge-driven and data-driven grouping method in an embodiment of this application;
- FIG. 2 is a schematic diagram of outputting people with similar first characteristics in an embodiment of this application
- FIG. 3 is a schematic diagram of outputting people with similar second characteristics in an embodiment of this application.
- FIG. 4 is a schematic diagram of clustering using knowledge clustering decision tree and data clustering decision tree in an embodiment of the application
- FIG. 5 is a schematic diagram of another embodiment of a knowledge-driven and data-driven grouping method in an embodiment of the application.
- Fig. 6 is a schematic diagram of an embodiment of a knowledge-driven and data-driven clustering device in an embodiment of the application;
- FIG. 7 is a schematic diagram of another embodiment of a knowledge-driven and data-driven grouping device in an embodiment of this application.
- FIG. 8 is a schematic diagram of another embodiment of a knowledge-driven and data-driven clustering device in an embodiment of this application.
- the embodiments of the application provide a knowledge-driven and data-driven clustering method and related equipment, which are used to create a target knowledge clustering decision tree, determine that the population to be classified belongs to the first group with similar characteristics, create a data clustering decision tree, and based on the target layer In order to improve the accuracy of grouping the population to be classified into groups with similar characteristics, it will provide a reasonable, standardized and comprehensive and refined cycle for doctors’ decision-making. Proof of medical evidence.
- steps or modules may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or equipment.
- the division of modules in this application is only a logical division In actual applications, there may be other divisions when implemented. For example, multiple modules may be combined or integrated in another system, or some features may be ignored or not implemented.
- FIG. 1 a flowchart of a knowledge-driven and data-driven clustering method provided by an embodiment of the present application.
- the following provides an example of a knowledge-driven and data-driven clustering method provided in this application.
- the types are restricted, including:
- the medical rule data is used to indicate the archived medical
- the first grouping rule is the first grouping rule obtained by extracting the medical discipline knowledge data of the literature, medical consensus, medical guide and the clinical experience of experts.
- the first grouping rule is used to group similar clinical patients.
- the sample data includes following the medical knowledge rules Clinical outcome data.
- the medical discipline knowledge data of medical guidelines, medical consensus, medical literature, and experts' clinical experience are processed by data cleaning and data protocol processing to ensure the quality of medical discipline knowledge data.
- Rule extraction is performed on the medical discipline knowledge data that has been cleaned and processed by data specification.
- the extracted rules are used to construct coarse-grained grouping rules (that is, medical rule data), and the coarse-grained grouping rules are used to construct the initial knowledge grouping decision tree. Get the target knowledge grouping decision tree.
- the clinical outcome data ie sample data
- the clinical outcome data that follows the rules of medical knowledge is filled in the end of each branch of the corresponding knowledge grouping decision tree with the disease information of the clinically similar characteristic groups corresponding to the group.
- the end knowledge sub-nodes namely end knowledge sub node 1, end knowledge sub node 2, end knowledge sub node 3, end knowledge sub node 4, end knowledge sub node 5, end knowledge sub node 6 and end knowledge sub node 7, knowledge grouping
- the clinical outcome data outside the decision tree is clinical outcome data that does not follow the rules of medical knowledge
- the end knowledge sub-nodes of each branch of the knowledge grouping decision tree include at least one processing solution corresponding to disease information.
- the coarse-grained clustering rules obtained from medical guidelines, medical consensus, medical literature, and the clinical experience of experts in the medical discipline knowledge data and filling the first sample data into the end knowledge sub-nodes of each branch of the knowledge clustering decision tree are constructed based on Knowledge group decision tree for medical subject knowledge data and clinical outcome data.
- the content of the above figure 2 and the example content are for reference only, and the accuracy of the content and actual operation are not considered.
- the target end knowledge sub-node is The node position of the feature data in the decision tree of the target knowledge grouping, and the processing plan is used to indicate the solution channel for the group of similar features.
- the terminal obtains the characteristic data of the group to be classified, and uses the target knowledge grouping decision tree to group the group to be classified according to the characteristic data based on medical discipline knowledge, and obtains the group to be classified to the target end knowledge sub-node where the target knowledge grouping decision tree is located , And obtain the processing plan of the similar characteristic group corresponding to the target end knowledge sub-node, and identify the number of categories of the processing plan to determine whether to further group the similar characteristic group of the target end knowledge sub-node.
- the treatment plan is used to indicate a solution channel for people with similar characteristics, for example, a clinical treatment plan or prescription for a disease corresponding to a population with similar characteristics.
- the group to be classified belongs to the group with the first similar characteristics, and the group with the first similar characteristics is the target knowledge grouping decision tree, and the corresponding processing schemes belong to similar characteristics. crowd.
- the group to be classified is divided into groups with similar characteristics (ie, groups with similar first characteristics) corresponding to the target end knowledge sub-node where the target knowledge grouping decision tree is located, and only the group with similar characteristics is obtained.
- the grouping decision tree of the target knowledge is used to group the group of people to be classified into groups with similar first characteristics.
- the number of categories of the treatment plan of the end child node of each branch is one by analyzing whether the similarity of the medication prescriptions between the treatment plans is at a preset judgment value.
- a preset judgment value is 97%-100%.
- the similarity of medication prescriptions is 96%
- the similarity of medication prescriptions of treatment plan A and treatment plan C is 98%
- the similarity of medication prescriptions of treatment plan B and treatment plan C is 97.5%
- the end child node 1 is processed
- the number of categories of the scheme is 2.
- the end knowledge sub-node When it is analyzed that the number of categories of the processing plan of the end knowledge sub-node of each branch is one, the end knowledge sub-node does not perform subsequent data-driven grouping model creation operations, indicating that the population of the end knowledge sub-node is of the same type People with similar characteristics (ie, people with similar first characteristics).
- the content of the above examples is for reference only, and the accuracy of the content and actual operation are not considered.
- the data grouping decision tree is generated based on the second grouping rule, and the target layer number of the data sub-nodes of the data grouping decision tree corresponding to each processing scheme is determined.
- the dichotomy rule indicates classification based on the effectiveness of the recursive partitioning algorithm and processing scheme.
- the data grouping decision tree learns the data in the data set through machine learning or deep learning algorithms.
- the grouping method is to group people with similar clinical characteristics. Taking the target end knowledge child node as the parent node, starting with the parent node, the second grouping rule, which is classified according to the effectiveness of the recursive partitioning algorithm and processing scheme, branches the clinical outcome data of the population with similar characteristics corresponding to the data root node to Create a data clustering decision tree, and determine the target level of the data sub-nodes of the data clustering decision tree corresponding to each processing scheme.
- Data-driven clustering is carried out based on the medical subject knowledge data in the knowledge-driven clustering to further cluster similar groups of people and improve the accuracy of clustering.
- the similar feature groups of the first level parent node are branched, and the second level data node A is grouped into The effectiveness of the treatment plan is less than 90% for the second characteristic group of people with similar characteristics.
- the second data node is grouped into the second characteristic group whose effectiveness of the treatment plan is more than 90%.
- C Data node grouping is the second group with similar characteristics in which the effectiveness of the processing plan is 80% to 85%, and the effectiveness of grouping data nodes into the processing plan is 86% to 100%.
- the second feature of the ninety is similar to the crowd.
- the effectiveness includes but is not limited to the effectiveness of the drug efficacy.
- the target level is used to determine whether the group to be classified is divided into the target knowledge grouping decision tree to group the group to be classified into the group with similar characteristics or the data grouping decision tree.
- the target knowledge grouping decision tree is used to group the groups to be classified into groups with similar first characteristics, and then further grouping groups with similar characteristics to the second group.
- the data sub-nodes of each branch of each layer are from top to bottom, whether the treatment plan for the group with similar characteristics of the parent node and the treatment plan for the group with similar characteristics of the child node have the effectiveness of drug efficacy
- the difference is analyzed, and the layer number corresponding to the child node corresponding to the significant difference in the effectiveness of the treatment plan is used as the selected layer number of the data grouping decision tree, and it will be placed on the end child node of the knowledge grouping decision tree corresponding to the parent node Mark the target layer number.
- the target layer number is used as the basis for whether to call the data grouping decision tree and the data-driven grouping of similar groups of people through the data grouping decision tree.
- at least one of the cure rate or apparent rate or improvement rate or inefficiency or other indicators of drug effectiveness can be used to analyze the effectiveness of drug efficacy.
- the group to be classified is divided into groups with similar first characteristics.
- the group to be classified is divided into the target knowledge grouping decision tree to group the group to be classified into groups with similar first characteristics, and the data is no longer passed
- the grouping decision tree grouping groups the treated groups of people into groups.
- the group to be classified is divided into the second group with similar characteristics.
- the second group with similar characteristics is the group with similar characteristics to which the processing scheme corresponding to the target data node in the data grouping decision tree belongs.
- Each target data node corresponds to different groups of people with similar characteristics.
- the target level is the first level, that is, the parent node of the data grouping decision tree
- the group with similar characteristics is regarded as the second group with similar characteristics obtained by grouping the group to be classified by combining the target knowledge grouping decision tree and the data grouping decision tree. For example: as shown in Figure 3, if the target level is 2, the group to be classified is divided into the second feature similar group corresponding to the data node B, if the target level is 3, the group to be classified is divided into node C and The second characteristic group corresponding to the D node is similar.
- the group to be classified is divided into corresponding groups with similar characteristics, and the judgment based on the target level is to divide the group to be classified into the target knowledge grouping Whether the decision tree groups the group to be classified into groups with similar first characteristics, or the data grouping decision tree to group the target knowledge grouping decision tree to group the group to be classified into groups with similar first characteristics, and then further grouping groups with similar characteristics to obtain the second group with similar characteristics,
- the medical assistant decision-making system obtains the first evidence-based medical information corresponding to the population with similar first characteristics that provides medical knowledge and supplements the corresponding real-world evidence according to the first feature similar group obtained by the method of this application, according to the method of this application
- the obtained second characteristic similar group obtains the second evidence-based medicine information corresponding to the second characteristic similar group providing medical knowledge, supplementing corresponding real-world evidence, and combining data-driven grouping. It not only defines similar patients, but also provides personalized evidence-based medical evidence information
- the target knowledge grouping decision tree is created, the group to be classified is determined to belong to the group with similar first characteristics, the data grouping decision tree is created, and the group to be classified is divided into the group with similar first characteristics or the group with similar second characteristics according to the number of target layers. Steps such as population, improve the accuracy of grouping the population to be classified into groups of similar characteristics, and then provide reasonable, standardized and comprehensive and refined evidence-based medical evidence for doctors' decision-making.
- Another embodiment of the knowledge-driven and data-driven clustering method in the embodiment of the present application includes:
- the medical rule data is used to indicate the archived medical
- the first grouping rule is the first grouping rule obtained by extracting the medical discipline knowledge data of the literature, medical consensus, medical guide and the clinical experience of experts.
- the first grouping rule is used to group similar clinical patients.
- the sample data includes following the medical knowledge rules Clinical outcome data.
- the above-mentioned creation of the initial knowledge grouping decision tree based on the medical rule data, and filling the sample data to the end knowledge sub-nodes of each branch of the initial knowledge grouping decision tree, before obtaining the target knowledge grouping decision tree further includes:
- the system, multiple platforms, and multiple medical systems acquire medical subject knowledge data and clinical data, and perform data processing on medical subject knowledge data and clinical data to obtain medical rule data.
- Data processing includes data cleaning processing, data integration processing, and data protocol Processing, data transformation processing, feature extraction processing, data desensitization processing and classification processing.
- data preprocessing includes data cleaning and data processing.
- the quality of medical subject knowledge data is guaranteed to facilitate subsequent data analysis and data mining of medical subject knowledge data .
- the sensitive private data in the medical data can be reliably protected, and the quality of the medical data will not be affected on the basis of ensuring the reliable protection of sensitive private data.
- a structured query language SQL may be executed on a page to ensure that each medical data is desensitized and does not interfere with each other.
- the above identification of the number of categories of the processing schemes specifically includes: calculating the similarity between the processing schemes to obtain multiple similarity values, and comparing whether the multiple similarity values are all within a preset range; if If the comparison result is yes, then there is one type of statistical processing plan; if the comparison result is no, then there are multiple types of statistical processing plan.
- the type of the processing scheme can be determined.
- the similarity includes, but is not limited to, the similarity of medication prescriptions, the similarity of medication prescriptions combined with drug usage, and the similarity of disease information combined with medication prescriptions.
- the group to be classified belongs to the group with the first similar characteristics, and the group with the first similar characteristics is the target knowledge grouping decision tree and the corresponding processing schemes belong to similar characteristics. crowd.
- the dichotomy rule indicates classification based on the effectiveness of the recursive partitioning algorithm and processing scheme.
- the above-mentioned generating a data clustering decision tree based on the second clustering rule, and determining the target layer number of the data sub-nodes of the data clustering decision tree corresponding to each processing scheme specifically includes: Step 1: Combining the target end knowledge As the data root node of the data clustering decision tree, the child node calculates the effectiveness score of the processing plan corresponding to the data root node; Step 2: Branch the data root node according to the effectiveness score of each processing plan, and according to the recursive partition algorithm Divide each processing scheme into the data sub-nodes of each branch in turn, and continue to analyze the effectiveness of the processing scheme corresponding to each data sub-node; Step 3: If the effectiveness of the processing scheme corresponding to the current data sub-node If the score is greater than the first threshold, branch the current data sub-node; Step 4: Repeat step 3 until the validity score of the processing scheme corresponding to each data sub-node is less than or equal to the first threshold, and data grouping is obtained Decision
- the data grouping decision tree is constructed. For example, if the first threshold is 60 points, the processing scheme of the data root node is divided into 95 points or less and 95 points or more (that is, groups of people with similar characteristics are grouped), and the validity scores below 95 points are all greater than 60 points, then The treatment plan corresponding to the validity score of 95 points or less is branched, and the validity score of 95 points or more is greater than 60 points, then the treatment plan corresponding to the validity score of 95 points or more is branched, and so on, until the end The validity score of the processing scheme corresponding to the data sub-node corresponding to the branch is less than or equal to 60 points.
- the level of the end data leaf node in the data grouping decision tree is the target level.
- the effectiveness score of the processing plan can be used as a parameter, and the division range value of all the effectiveness scores obtained by measurement or measurement for the processing plan can be used as all observations.
- the parameters are estimated to obtain the parameter estimation value, and the stability of the parameter in different partitions is evaluated according to the parameter estimation value; if the evaluation result is unstable, the partition variable is analyzed and the stability is selected The lowest partition variable is evaluated for stability until the evaluation result is stable, and the partition variable is used to indicate the regression of the partition; if the evaluation result is stable, the stability of the parameter in different partitions will be stopped; based on the stability of the parameter in different partitions , According to the dichotomy or variable adaptive method to obtain the optimized region segmentation rules, the dichotomy is used to indicate that the number of selected sub-nodes is two, and the variable adaptive method is used to indicate that the number of selected sub-nodes is equal to 2.
- the method of the number of partitions obtain the sub-nodes corresponding to the regional segmentation rules according to the regional segmentation rules, repeat the execution on the sub-nodes corresponding to the regional segmentation rules, evaluate the stability of the parameters in different partitions according to the parameter estimates and adapt according to the dichotomy or variable adaptation
- the method obtains the optimized region segmentation rules until the parameters are stable in different partitions after evaluation, to group the people to be classified and the people with similar characteristics corresponding to the target end sub-nodes in the target knowledge grouping decision tree to obtain a data grouping decision tree.
- the model-based recursive partition and the clustering rule strongly related to clinical outcome data are used to create a data clustering decision tree.
- each node is associated with the model; the stability fluctuation analysis of the parameters is performed to evaluate whether the node needs to be split; if any partition variable has significant instability, the node is divided into local optimal ends , And repeat the process.
- the estimation of the parameters can be performed by the score function or the estimation function.
- the evaluation of the stability of the parameters in different zones can be captured through the empirical fluctuation process.
- the above-mentioned taking the target end knowledge sub-node as the data root node of the data clustering decision tree, and calculating the validity score of the processing scheme corresponding to the data root node specifically includes: using the target end knowledge sub-node as the data clustering decision tree The data root node of the data root node; obtain the target feature information of the processing plan corresponding to the data root node, split the target feature information into factor factors, calculate the correlation degree of the factor factors at different levels, the degree of influence on the preset influencing factors, and factor factors Uncontrolled edge relevance; analyze the relevance, influence and edge relevance of factor factors through preset compressible criteria and comparable criteria to obtain the confounding factor; calculate the regression coefficient of the confounding factor and the preset probability of influencing factors Exponential transformation value, and use the exponential transformation value as the judgment score to obtain the validity of the processing scheme corresponding to the data root node
- the target feature information can be the disease information in the processing plan of the parent node of the data grouping decision tree, the preset influencing factor is the disease, the influence degree is the risk of the disease, and the effectiveness is the drug effectiveness;
- the confounding factor refers to a Factors that are related to exposure and disease. It is not only an independent risk factor for the disease being studied, it must also be a risk factor in the non-exposure group.
- the drug effectiveness analysis is performed on the treatment plan corresponding to the population with similar characteristics to obtain the sub-nodes with statistical analysis significance, so that the knowledge grouping decision tree is grouped quickly and accurately through the data grouping decision tree, and then the knowledge grouping decision tree is grouped and then further grouped.
- the possible confounding factors are first selected, and then the confounding factors are added to the multivariate logistic regression analysis. Finally, the exponential transformation value of the regression coefficient of the drug variable is used as the final effectiveness judgment score. To obtain multiple valid values. Among them, compressible criteria, comparable criteria, virtual fact models, and graph models can be used as the identification criteria of the confounding factor to obtain the confounding factor.
- target level is the first level, classify the group of people to be classified into groups with similar first characteristics.
- the target layer number is not the first layer, determine whether the effectiveness score of the processing scheme corresponding to the target layer number is greater than a second threshold, and the second threshold is greater than the first threshold.
- the data grouping decision tree needs to be used to further group the people to be classified.
- Obtain the effectiveness score of the processing scheme corresponding to the target level and determine whether the effectiveness score is greater than the second threshold, so as to ensure the accuracy of grouping the population to be classified. For example, as shown in FIG. 3, the target layer number is the second layer, then the effectiveness score of the processing scheme corresponding to the second group of similar characteristics of the node B is obtained, and whether the effectiveness score is greater than the second threshold is analyzed.
- the effectiveness score of the processing scheme corresponding to the target layer number is greater than the second threshold, use the data leaf node where the processing scheme with the effectiveness score greater than the second threshold is located as the target data node, and divide the population to be classified into A group with similar characteristics to which the processing scheme corresponding to the target data node belongs is obtained, and a second group with similar characteristics is obtained.
- the target layer corresponding to the processing scheme's effectiveness score is greater than the second threshold.
- the target data node in the data decision tree is located in the target data node corresponding to the group with similar characteristics as the second characteristic to divide the group to be classified Similar people.
- the second threshold is 70 points
- the effectiveness score of the processing scheme corresponding to the target number of layers is 80 points
- the data leaf node with the effectiveness score of the corresponding processing scheme of 80 points is taken as the target data node, and the data node to be classified is
- the group is divided into groups with similar characteristics to which the processing scheme corresponding to the target data node belongs (ie, groups with similar second characteristics).
- the effectiveness score of the processing scheme corresponding to the target layer number is less than or equal to the second threshold
- the population with similar characteristics corresponding to the target layer number is not the final group of the group to be classified, and the last one of the target layer number needs to be analyzed.
- the number of layers corresponds to the effectiveness score of the treatment plan.
- the feature-similar population corresponding to the processing solution with a validity score less than or equal to the second threshold has a low matching degree with the group to be classified, and the nodes corresponding to the validity score less than or equal to the second threshold or later nodes will not be analyzed. To improve operation efficiency and accuracy of grouping.
- the effectiveness score of the processing scheme corresponding to the upper level of the target level is greater than the second threshold, the data leaf node where the processing scheme is located at the upper level of the target level is taken as the target data node, and The group to be classified is divided into groups with similar characteristics to which the processing scheme corresponding to the target data node belongs, and a second group with similar characteristics is obtained.
- the second feature similar population of the target node of the data grouping decision tree is a group with similar second features, and the second feature similar group corresponding to this node is a combination of knowledge-driven and data-driven populations to be classified and divided into groups that are strongly related to clinical outcome data. .
- the effectiveness score of the processing scheme corresponding to the upper level of the target level is less than or equal to the second threshold, whether the effectiveness score of the processing scheme corresponding to the upper level of the target level is greater than the first Two thresholds for judgment, and so on, until it is analyzed that the validity score of the corresponding processing plan is greater than the second threshold, and the characteristics of the target node of the data grouping decision tree corresponding to the processing plan greater than the second threshold are similar.
- the population is the second group with similar characteristics to be finally divided into the group to be classified.
- the target level is not the first level
- after the group to be classified is divided into groups with similar second characteristics, it specifically includes: calling a preset interpretation model for groups with similar first characteristics or second characteristics. Similar groups of people undergo interpretation processing and visual chart statistical analysis processing to obtain groups of similar interpretable characteristics.
- the interpretation processing includes effectiveness ranking processing, matching degree calculation processing and matching degree ranking processing.
- the effectiveness of the first group with similar characteristics or the second group with similar characteristics can be sorted in descending order of value to obtain the first ranking information; calculate the difference between the group to be classified and the first group
- the matching degree of the group with similar characteristics or the group with similar second characteristics, and the matching degree is sorted according to the value from the largest to the smallest to obtain the second ranking information; according to the first ranking information and/or the second ranking information, the first feature is sorted
- the similar group or the second group with similar characteristics are sorted, and the first group with similar characteristics or the second group with similar characteristics are marked according to the first ranking information and/or the second ranking information; the group with the first similar characteristics after the marking or The second group of people with similar characteristics performs the statistical analysis and processing of the visual chart to generate the visual chart, and the group of people with similar interpretable characteristics is obtained.
- the first group with similar characteristics or the second group with similar characteristics can be used to facilitate the user to easily obtain the grouping data of the group with similar characteristics
- the medical assistant decision-making system obtains the first evidence-based medical information corresponding to the population with similar first characteristics that provides medical knowledge and supplements the corresponding real-world evidence according to the first feature similar group obtained by the method of this application, according to the method of this application After obtaining the second evidence-based medical information corresponding to the people with similar second characteristics, providing medical knowledge, supplementing the corresponding real-world evidence, and combining data-driven grouping with the second evidence-based medicine information, including: calling the created explanation
- the model obtains the first feature information of the first sample data, the second feature information of the first evidence-based medicine evidence information, and/or the first feature information of the second evidence-based medicine evidence information, and the first sample data indicates not to follow Clinical outcome data of medical knowledge rules; analyze the first characteristic information and the second characteristic information to obtain the first analysis data, and analyze the first characteristic information and the third characteristic information to obtain the second analysis data; Perform statistics and analysis on the first feature information, the second feature information, and the third feature information to generate a visualization chart; call the created knowledge graph, obtain the first physical information
- the explanation data will be displayed in the form of a combination of graphics and text, so that users can better understand the content of the interpretable data.
- Explainable data is obtained by explaining the evidence information of the first evidence-based medicine or the evidence information of the second evidence-based medicine, so as to effectively provide more personalized, accurate and effective real-world evidence for assisting doctors in decision-making.
- the target knowledge grouping decision tree is created, the group to be classified is determined to belong to the group with similar first characteristics, the data grouping decision tree is created, and the group to be classified is divided into the group with similar first characteristics or the group with similar second characteristics according to the target level. Steps such as population, improve the accuracy of grouping the population to be classified into groups of similar characteristics, and then provide reasonable, standardized and comprehensive and refined evidence-based medical evidence for doctors' decision-making.
- An example of a data-driven clustering device includes:
- the first creating unit 601 is used to create an initial knowledge grouping decision tree according to the medical rule data, and fill the sample data to the end knowledge sub-nodes of each branch of the initial knowledge grouping decision tree to obtain the target knowledge grouping decision tree, and the medical rule data is used Indicate the first grouping rule obtained by extracting the medical discipline knowledge data of archived medical literature, medical consensus, medical guidelines, and the clinical experience of experts.
- the first grouping rule is used to group clinically similar groups of people,
- the sample data includes the clinical outcome data corresponding to the clinically similar populations that follow the rules of medical knowledge;
- the recognition unit 602 is used to obtain the characteristic data of the group to be classified, and group the characteristic data through the target knowledge grouping decision tree to obtain the processing plan corresponding to the target end knowledge sub-node, and to identify the number of categories of the processing plan.
- the knowledge sub-node is the node position of the feature data on the target knowledge grouping decision tree, and the processing plan is used to indicate the solution channel for the group of similar characteristics;
- the determining unit 603 is configured to determine that the group to be classified belongs to the first group with similar characteristics if the number of categories of the processing scheme is one, and the first group with similar characteristics is the processing scheme corresponding to the target end knowledge sub-node in the target knowledge grouping decision tree Groups with similar characteristics;
- the second creating unit 604 is configured to generate a data grouping decision tree based on the second grouping rule if the number of categories of the processing scheme is at least two, and determine the data sub-nodes of the data grouping decision tree corresponding to each processing scheme.
- the number of target layers at the location, the second grouping rule indicates classification according to the effectiveness of the recursive partitioning algorithm and processing scheme;
- the first dividing unit 605 is configured to divide the group to be classified into groups with similar first characteristics if the target level is the first level;
- the second division unit 606 is configured to, if the number of the target layer is not the first layer, divide the group to be classified into a group with similar second characteristics, and the second group with similar characteristics belongs to the processing scheme corresponding to the target data node in the data grouping decision tree People with similar characteristics, each target data node corresponds to different groups with similar characteristics.
- the target knowledge grouping decision tree is created, the group to be classified is determined to belong to the group with similar first characteristics, the data grouping decision tree is created, and the group to be classified is divided into the group with similar first characteristics or the group with similar second characteristics according to the number of target layers. Steps such as population, improve the accuracy of grouping the population to be classified into groups of similar characteristics, and then provide reasonable, standardized and comprehensive and refined evidence-based medical evidence for doctors' decision-making.
- another embodiment of the knowledge-driven and data-driven clustering device in the embodiment of the present application includes:
- the first creating unit 601 is used to create an initial knowledge grouping decision tree according to the medical rule data, and fill the sample data to the end knowledge sub-nodes of each branch of the initial knowledge grouping decision tree to obtain the target knowledge grouping decision tree, and the medical rule data is used Indicate the first grouping rule obtained by extracting the medical knowledge data of archived medical literature, medical consensus, medical guidelines, and the clinical experience of experts.
- the first grouping rule is used to group similar patients in the clinic, samples
- the data includes the clinical outcome data corresponding to the clinically similar people who follow the rules of medical knowledge;
- the recognition unit 602 is used to obtain the characteristic data of the group to be classified, and group the characteristic data through the target knowledge grouping decision tree to obtain the processing plan corresponding to the target end knowledge sub-node, and to identify the number of categories of the processing plan.
- the knowledge sub-node is the node position of the feature data on the target knowledge grouping decision tree, and the processing plan is used to indicate the solution channel for the group of similar characteristics;
- the determining unit 603 is configured to determine that the group to be classified belongs to the first group with similar characteristics if the number of categories of the processing scheme is one, and the first group with similar characteristics is the processing scheme corresponding to the target end knowledge sub-node in the target knowledge grouping decision tree Groups with similar characteristics;
- the second creating unit 604 is configured to generate a data grouping decision tree based on the second grouping rule if at least two types of processing schemes are identified, and determine where each processing scheme corresponds to the data sub-node of the data grouping decision tree
- the second grouping rule instructs to classify according to the effectiveness of the recursive partitioning algorithm and processing scheme
- the first dividing unit 605 is configured to divide the group to be classified into groups with similar first characteristics if the target level is the first level;
- the judging unit 606 is configured to determine whether the target layer number is not the first layer, and determine whether the validity score of the processing scheme corresponding to the target layer number is greater than a second threshold, and the second threshold is greater than the first threshold;
- the third dividing unit 607 is configured to, if the effectiveness score of the processing scheme corresponding to the target layer number is greater than the second threshold, use the data leaf node where the processing scheme with the effectiveness score greater than the second threshold is located as the target data node, and Divide the group to be classified into groups with similar characteristics to which the processing solution corresponding to the target data node belongs, and obtain a second group with similar characteristics;
- the analyzing unit 608 is configured to analyze the effectiveness score of the processing solution corresponding to the upper layer of the target layer if the effectiveness score of the processing solution corresponding to the target layer number is less than or equal to the second threshold;
- the third dividing unit 609 is configured to, if the validity score of the processing scheme corresponding to the upper level of the target level is greater than the second threshold, use the data leaf node where the upper level of the target level corresponds to the processing scheme as The target data node is divided into groups with similar characteristics to which the processing scheme corresponding to the target data node belongs, and a second group with similar characteristics is obtained.
- the second creating unit 604 is specifically configured to: Step 1: Use the target end knowledge sub-node as the data root node of the data grouping decision tree, and calculate the validity score of the processing scheme corresponding to the data root node Step 2: Branch the data root node according to the validity score of each processing plan, divide each processing plan into the data sub-nodes of each branch according to the recursive partition algorithm, and continue to each Perform validity analysis on the processing plan corresponding to the data sub-node; Step 3: If the validity score of the processing plan corresponding to the current data sub-node is greater than the first threshold, branch the current data sub-node; Step 4: Repeat Step 3 is performed until the validity score of the processing scheme corresponding to each data sub-node is less than or equal to the first threshold, and a data grouping decision tree is obtained.
- the data grouping decision tree includes a plurality of data sub-nodes. Each end data leaf node of the multiple data sub-nodes corresponds to a processing plan; Step 5: Determine the target level of the end data leaf node corresponding to each processing plan in the data grouping decision tree.
- the second creating unit 604 is specifically configured to: use the target end knowledge sub-node as the data root node of the data grouping decision tree; obtain the target feature information of the processing scheme corresponding to the data root node, and set the target
- the characteristic information is divided into factor factors; the correlation degree of factor factors at different levels, the influence degree on the preset influencing factors and the marginal correlation degree not controlled by the factor factor are calculated; the factors are analyzed through the preset compressible criteria and comparable criteria
- the correlation degree, influence degree, and edge correlation degree of the factors are used to obtain the confounding factor; the exponential transformation value of the regression coefficient of the confounding factor and the preset influencing factor occurrence probability is calculated, and the exponential transformation value is used as the judgment score to obtain the processing corresponding to the root node of the data
- a knowledge-driven and data-driven clustering device also includes: a processing unit for obtaining medical subject knowledge data and clinical data from multiple systems, multiple platforms, and multiple medical systems, and analyzing medical subject knowledge data Perform data processing with clinical data to obtain medical rule data.
- Data processing includes data cleaning processing, data integration processing, data protocol processing, data transformation processing, feature extraction processing, data desensitization processing, and classification processing.
- the processing unit is specifically used to: obtain medical subject knowledge data from multiple systems and multiple platforms, and obtain clinical data from multiple medical systems; perform data preprocessing on medical subject knowledge data and clinical data respectively, and data preprocessing includes data cleaning Processing, data integration processing, data protocol processing and data transformation processing; feature extraction of medical subject knowledge data after data preprocessing to obtain feature information, and data desensitization processing and classification of data preprocessing clinical data to get the first Sample data and third sample data; analyze the characteristic information, classify the characteristic information according to preset classification rules, and obtain medical rule data.
- the preset classification rules include the same and/or similar diseases.
- the identification unit 602 is specifically configured to: calculate the similarity between the processing schemes to obtain multiple similarity values, and compare whether the multiple similarity values are all within a preset range; if the comparison result is yes, perform statistical processing There is one type of scheme; if the comparison result is no, there are multiple types of statistical processing schemes.
- a knowledge-driven and data-driven clustering device further includes: a calling unit for calling a preset interpretation model to perform interpretation processing and visual chart statistical analysis processing for people with similar first characteristics or second groups with similar characteristics, Groups with similar interpretable characteristics are obtained, and the interpretation processing includes effectiveness ranking processing, matching degree calculation processing and matching degree ranking processing.
- the calling unit is specifically used to: call the preset interpretation model to sort the effectiveness of the first group with similar characteristics or the second group with similar characteristics in descending order of value to obtain the first ranking information;
- the matching degree of a group with similar characteristics or a group with similar second characteristics, and the degree of matching is sorted in the order of the value from the largest to the smallest to obtain the second ranking information; according to the first ranking information and/or the second ranking information, the first ranking information
- the group with similar characteristics or the second group with similar characteristics is sorted, and the first group with similar characteristics or the second group with similar characteristics are marked according to the first ranking information and/or the second ranking information; the first group with similar characteristics after the marking is marked Or the second group with similar characteristics performs statistical analysis and processing on the visual chart to generate a visual chart to obtain groups with similar interpretable characteristics.
- the target knowledge grouping decision tree is created, the group to be classified is determined to belong to the group with similar first characteristics, the data grouping decision tree is created, and the group to be classified is divided into the group with similar first characteristics or the group with similar second characteristics according to the number of target layers. Steps such as population, improve the accuracy of grouping the population to be classified into groups of similar characteristics, and then provide reasonable, standardized and comprehensive and refined evidence-based medical evidence for doctors' decision-making.
- FIG. 8 is a schematic structural diagram of a knowledge-driven and data-driven clustering device according to an embodiment of the present application.
- the knowledge-driven and data-driven clustering device 800 may have relatively large differences due to different configurations or performances, and may include One or more processors (central processing units, CPU) 801 (for example, one or more processors) and memory 809, and one or more storage media 808 for storing application programs 807 or data 806 (for example, one or more storage mediums 808) Storage device).
- the memory 809 and the storage medium 808 may be short-term storage or persistent storage.
- the program stored in the storage medium 808 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the sign-in management device.
- the processor 801 may be configured to communicate with the storage medium 808, and execute a series of instruction operations in the storage medium 808 on the knowledge-driven and data-driven clustering device 800.
- the knowledge-driven and data-driven clustering device 800 may also include one or more power supplies 802, one or more wired or wireless network interfaces 803, one or more input and output interfaces 804, and/or, one or more operating systems 805, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc.
- operating systems 805 such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc.
- FIG. 8 does not constitute a limitation on the knowledge-driven and data-driven clustering device, and may include more or less than that shown in the figure. Components, or a combination of certain components, or different component arrangements.
- the processor 801 can execute the first creation unit 601, the recognition unit 602, the second creation unit 603, the first analysis unit 604, the first marking unit 605, the second analysis unit 606, the second marking unit 607, and the The functions of an output unit 608, a second output unit 609, a classification unit 610, and a first call unit 611.
- the processor 801 is a control center of a knowledge-driven and data-driven clustering device, and can perform processing in accordance with a knowledge-driven and data-driven clustering method.
- the processor 801 uses various interfaces and lines to connect the various parts of the entire knowledge-driven and data-driven clustering device, by running or executing software programs and/or modules stored in the memory 809, and calling data stored in the memory 809 , Perform various functions and process data based on knowledge-driven and data-driven clustering equipment, so as to realize the function of improving the usability of the case retrieval system.
- the storage medium 808 and the memory 809 are both carriers for storing data.
- the storage medium 808 may refer to an internal memory with a small storage capacity but a fast speed, and the storage medium 809 may have a large storage capacity but a slow storage speed. External memory.
- the memory 809 may be used to store software programs and modules.
- the processor 801 executes various functional applications and data processing of the knowledge-driven and data-driven clustering device 800 by running the software programs and modules stored in the memory 809.
- the memory 809 may mainly include a storage program area and a storage data area.
- the storage program area may store an operating system, an application program required by at least one function (obtain characteristic data of the group to be classified, etc.), etc.;
- the data created by the use of the management device (the feature data is grouped through the target knowledge grouping decision tree, the processing plan corresponding to the target end child node is obtained, and the number of categories of the processing plan is identified, etc.), etc.
- the memory 809 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
- a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
- the computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
- Computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
- Computer instructions may be transmitted from a website, computer, server, or data center through a cable (such as Coaxial cable, optical fiber, twisted pair) or wireless (such as infrared, wireless, microwave, etc.) to transmit to another website site, computer, server or data center.
- the computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center integrated with one or more available media.
- the usable medium may be a magnetic medium (for example, a floppy disk, a hard disk, and a magnetic tape), an optical medium (for example, an optical disc), or a semiconductor medium (for example, a solid state disk (SSD)).
Landscapes
- Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Public Health (AREA)
- Artificial Intelligence (AREA)
- Medical Informatics (AREA)
- Health & Medical Sciences (AREA)
- Evolutionary Computation (AREA)
- Physics & Mathematics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Biomedical Technology (AREA)
- Computational Linguistics (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Databases & Information Systems (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Pathology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Epidemiology (AREA)
- General Health & Medical Sciences (AREA)
- Primary Health Care (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本申请涉及人工智能领域,公开了一种基于知识驱动和数据驱动的分群方法及相关设备,提高了将待分类人群分群到相似特征人群的精准度。本申请方法包括:创建目标知识分群决策树;识别待分类人群的特征数据在所述目标知识分群决策树上目标末端知识子节点对应的处理方案的类别数量;当识别到所述类别数量为一个时,确定所述待分类人群属于目标知识分群决策树中目标末端知识子节点对应的第一特征相似人群;当识别到所述类别数量为至少两个时,基于预置第二分群规则生成数据分群决策树,并确定所述数据分群决策树的目标层数;根据所述目标层数将所述待分类人群划分为所述第一特征相似人群或所述数据分群决策树中目标数据节点对应的第二特征相似人群。
Description
本申请要求于2019年10月18日提交中国专利局、申请号为201910991392.5、发明名称为“基于知识驱动和数据驱动的分群方法及相关设备”的中国专利申请的优先权,其全部内容通过引用结合在申请中。
本申请涉及人工智能领域,尤其涉及基于知识驱动和数据驱动的分群方法及相关设备。
个性化精准治疗是目前医学治疗领域一个重要的研究课题,也是医学研究领域一个值得期待的发展方向。而个性化精准治疗主要从群体出发,并根据患者的个体差异提供个性化治疗,在这个过程中,最困难的是如何确定不同群体,并针对群体中患者的个性化治疗给出充足的医学知识支撑。现有的分群系统中,常使用完全基于数据驱动或者完全基于知识驱动的分群方法,但两种方法都存在较多弊端。
完全基于数据驱动的分群方法中,其核心思想是基于机器学习或深度学习算法学习数据集中医生开药规则,并将学习到的规则应用于同类病人之中。而完全基于知识驱动的分群方法,其核心是参照已有指南、共识等相关权威知识构建规则,根据规则获取相应的辅助决策方案。然而,发明人发现,不同于其他行业,使用完全基于数据驱动的分群方法并不能保证算法学习到的医学规则符合医学知识,并且无法做到针对相同群体中不同病人进行个性化推荐对应的辅助决策方案;使用完全基于知识驱动的推荐方法则存在着无法对推荐的多种辅助决策方案进行合理的选择,并存在指南、共识等知识中完全无法获取的辅助信息,如:药品价格。因而,导致将待分类人群分群到相似特征人群的精准度低。
发明内容
本申请提供了一种基于知识驱动和数据驱动的分群方法及相关设备,用于创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
本申请实施例的第一方面提供一种基于知识驱动和数据驱动的分群方法,包括:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
本申请实施例的第二方面提供一种用于基于知识驱动和数据驱动的分群装置,所述基于知识驱动和数据驱动的分群装置包括:第一创建单元,用于根据医学规则数据创建初始知识分群 决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;识别单元,用于获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;第二创建单元,用于若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;第一划分单元,用于若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;第二划分单元,用于若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
本申请实施例的第三方面提供了一种基于知识驱动和数据驱动的分群装置,包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时执行如下所述的基于知识驱动和数据驱动的分群方法的步骤:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
本申请实施例的第四方面提供了一种计算机可读存储介质,包括指令,当所述指令在计算机上运行时,使得计算机执行如下所述的基于知识驱动和数据驱动的分群方法的步骤:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策 树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
相较于现有技术,本申请实施例提供的技术方案中,通过创建目标知识分群决策树;识别待分类人群的特征数据在所述目标知识分群决策树上目标末端知识子节点对应的处理方案的类别数量;当识别到所述类别数量为一个时,确定所述待分类人群属于目标知识分群决策树中目标末端知识子节点对应的第一特征相似人群;当识别到所述类别数量为至少两个时,基于预置第二分群规则生成数据分群决策树,并确定所述数据分群决策树的目标层数;根据所述目标层数将所述待分类人群划分为所述第一特征相似人群或所述数据分群决策树中目标数据节点对应的第二特征相似人群。本申请实施例,通过创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
图1为本申请实施例中基于知识驱动和数据驱动的分群方法的一个实施例示意图;
图2为本申请实施例中输出第一特征相似人群的一种示意图;
图3为本申请实施例中输出第二特征相似人群的一种示意图;
图4为本申请实施例中采用知识分群决策树和数据分群决策树分群的一种示意图;
图5为本申请实施例中基于知识驱动和数据驱动的分群方法的另一个实施例示意图;
图6为本申请实施例中基于知识驱动和数据驱动的分群装置的一个实施例示意图;
图7为本申请实施例中基于知识驱动和数据驱动的分群装置的另一个实施例示意图;
图8为本申请实施例中基于知识驱动和数据驱动的分群设备的另一个实施例示意图。
本申请实施例提供了一种基于知识驱动和数据驱动的分群方法及相关设备,用于创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
为了使本技术领域的人员更好地理解本申请方案,下面将结合本申请实施例中的附图,对本申请实施例进行描述。
应当理解,此处所描述的具体实施例仅用以解释本申请,并不用于限定本申请。本申请的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的实施例能够以除了在这里图示或描述的内容以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或模块的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或模块,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或模块,本申请中所出现的模块的划分,仅仅是一种逻辑上的划分,实际应用中实现时可以有另外的划分方式,例如多个模块可以结合成或集成在另一个系统中,或一些特征可以忽略,或不执行。
请参阅图1,本申请实施例提供的一种基于知识驱动和数据驱动的分群方法的流程图,以下对本申请提供一种基于知识驱动和数据驱动的分群方法进行举例说明,本申请不对执行主体的类型作限制,具体包括:
101、根据医学规则数据创建初始知识分群决策树,并将样本数据填充至初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,医学规则数据用于指示对已归档的医 学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,第一分群规则用于对临床上的相似患者进行分群,样本数据包括遵循医学知识规则的临床结局数据。
具体的,对医学指南、医学共识、医学文献和专家的临床经验的医学学科知识数据进行数据清理处理和数据规约处理,以保证医学学科知识数据的质量。对经过清理处理和数据规约处理的医学学科知识数据进行规则提取,将提取的规则构建较粗粒度的分群规则(即医学规则数据),用较粗粒度的分群规则构建初始知识分群决策树,以得到目标知识分群决策树。例如,如图2所示,将临床结局数据中遵循医学知识规则的临床结局数据(即样本数据)对应分群的临床上的相似特征人群的病症信息填充在对应的知识分群决策树各分支的末端知识子节点上,即末端知识子节点1、末端知识子节点2、末端知识子节点3、末端知识子节点4、末端知识子节点5、末端知识子节点6和末端知识子节点7,知识分群决策树之外的临床结局数据为不遵循医学知识规则的临床结局数据,则知识分群决策树各分支的末端知识子节点(末端知识子节点1、末端知识子节点2、末端知识子节点3、末端知识子节点4、末端知识子节点5、末端知识子节点6和末端知识子节点7)包含至少一种以上的病症信息对应的处理方案。通过从医学指南、医学共识、医学文献和专家的临床经验的医学学科知识数据获取的较粗粒度的分群规则和将第一样本数据填充到知识分群决策树各分支的末端知识子节点构建基于医学学科知识数据和临床结局数据的知识分群决策树。上述图2内容和举例内容仅作举例说明参考,其内容的准确性和实际操作与否不作考虑。
102、获取待分类人群的特征数据,并通过目标知识分群决策树对特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对处理方案的类别数量进行识别,目标末端知识子节点为特征数据在目标知识分群决策树上的节点位置,处理方案用于指示对相似特征人群的解决渠道。
终端获取待分类人群的特征数据,并通过目标知识分群决策树基于医学学科知识根据特征数据将待分类人群进行分群,获得待分类人群被分群到目标知识分群决策树所处的目标末端知识子节点,并获取目标末端知识子节点对应的相似特征人群的处理方案,对处理方案的类别数量进行识别,以判断是否对目标末端知识子节点的特征相似人群进行进一步分群。其中,处理方案用于指示对相似特征人群的解决渠道,例如,临床上的相似特征人群对应病症的治疗方案或者处方。
103、若识别到处理方案的类别数量为一个,则确定待分类人群属于第一特征相似人群,第一特征相似人群为目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群。
终端若识别到处理方案的类别数量为一个,则将待分类人群划分到所位于目标知识分群决策树所处的目标末端知识子节点对应的特征相似人群(即第一特征相似人群),得到仅通过目标知识分群决策树对待分类人群进行分群的第一特征相似人群。
例如,具体的,可通过处理方案之间的用药处方的相似度是否处于预设判断值来进行分析来分析各分支的末端子节点的处理方案的类别数量是否为一个。如图2的末端子节点1的处理方案有三个,分别为处理方案甲、处理方案乙和处理方案丙,预设判断值为相似度为97%~100%,处理方案甲与处理方案乙的用药处方的相似度为96%,处理方案甲与处理方案丙的用药处方的相似度为98%,处理方案乙与处理方案丙的用药处方的相似度为97.5%,则末端子节点1的处理方案的类别数量为2个。当分析到各分支的末端知识子节点的处理方案的类别数量是一个时,该末端知识子节点不进行后续的数据驱动分群模型的创建的操作,表示该末端知识子节点的人群是同一类的特征相似人群(即第一特征相似人群)。上述举例内容仅作举例说明参考,其内容的准确性和实际操作与否不作考虑。
104、若识别到处理方案的类别数量至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数,第二分群规则指示根据递归分区算法和处理方案的有效性进行分类。
具体的,当分析到各分支的末端知识子节点的处理方案的类别数量不是一个时,进入后续 的数据分群决策树的创建的操作,数据分群决策树通过机器学习或者深度学习算法学习数据集中的分群方式以对临床上的相似特征人群进行分群。以目标末端知识子节点作为父节点,以父节点开始,根据递归分区算法和处理方案的有效性进行分类的第二分群规则对数据根节点对应的相似特征人群的临床结局数据进行分支,以创建数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数。通过基于在知识驱动分群中的医学学科知识数据进行数据驱动分群,以对相似人群进行进一步分群,提高分群的精度。
例如,如图3所示,根据第一层的数据根节点的相似特征人群的处理方案的有效性对第一层的父节点的相似特征人群进行分支,在第2层将甲数据节点分群为处理方案的有效性为百分之九十以下的第二特征相似人群,将乙数据节点分群为处理方案的有效性为百分之九十以上的第二特征相似人群,在第3层将丙数据节点分群为处理方案的有效性为百分之八十至百分之八十五的第二特征相似人群,将丁数据节点分群为处理方案的有效性为百分之八十六至百分之九十的第二特征相似人群。
其中,有效性包括但不限于药物疗效有效性,目标层数用于判断是将待分类人群划分为目标知识分群决策树对待分类人群进行分群所得的第一特征相似人群还是通过数据分群决策树对目标知识分群决策树对待分类人群进行分群所得的第一特征相似人群再进行进一步分群所得的第二特征相似人群。从数据分群决策树的父节点开始,每层的各分支的数据子节点由上至下,对父节点的特征相似人群的处理方案和子节点的特征相似人群的处理方案的药物疗效有效性是否存在差异进行分析,以处理方案的药物疗效有效性存在显著差异对应的子节点对应的层数作为数据分群决策树的选择层数,并将在父节点对应的知识分群决策树的该末端子节点上标记该目标层数。以目标层数作为是否调用数据分群决策树且通过数据分群决策树对相似人群进行基于数据驱动分群的层数的依据。其中,可通过治愈率或显效率或好转率或无效率或其他药物有效性的指标中的至少一项进行药物疗效有效性分析。
105、若目标层数为第一层,则将待分类人群划分为第一特征相似人群。
具体的,当目标层数为第一层,即数据分群决策树的父节点时,将待分类人群划分为目标知识分群决策树对待分类人群进行分群所得的第一特征相似人群,不再通过数据分群决策树分群对待分类人群进行分群。
106、若目标层数不为第一层,则将待分类人群划分为第二特征相似人群,第二特征相似人群为数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
具体的,当目标层数为第一层,即数据分群决策树的父节点时,获取目标层数在数据分群决策树对应的层数的末端子节点,并获取该末端子节点对应的特征相似人群,将该特征相似人群作为结合目标知识分群决策树和数据分群决策树对待分类人群进行分群所得的第二特征相似人群。例如:如图3所示,若目标层数为2,则将待分类人群划分到乙数据节点对应的第二特征相似人群,若目标层数为3,则将待分类人群划分到丙节点和丁节点对应的第二特征相似人群。
例如,如图4所示,结合知识分群决策树和数据决策树对待分类人群进行分群,将待分类人群划分到对应的特征相似人群,根据目标层数判断是将待分类人群划分为目标知识分群决策树对待分类人群进行分群所得的第一特征相似人群还是通过数据分群决策树对目标知识分群决策树对待分类人群进行分群所得的第一特征相似人群再进行进一步分群所得的第二特征相似人群,医学辅助决策系统根据本申请的方法所获得的第一特征相似人群,得到第一特征相似人群对应的提供医学知识以及补充相应的真实世界证据的第一循证医学信息,根据本申请的方法所获得的第二特征相似人群,得到第二特征相似人群对应的提供医学知识、补充相应的真实世界证据和结合数据驱动进行分群的第二循证医学信息。不仅定义相似患者,还提供针对相同群体中不同病人的符合医学学科知识的个性化循证医学证据信息,以辅助医生决策。
本申请实施例,通过创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
请参阅图5,本申请实施例中,501至505的方法可参见101至105,此处不再赘述。本申请实施例中基于知识驱动和数据驱动的分群方法的另一个实施例包括:
501、根据医学规则数据创建初始知识分群决策树,并将样本数据填充至初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,第一分群规则用于对临床上的相似患者进行分群,样本数据包括遵循医学知识规则的临床结局数据。
可选的,上述的根据医学规则数据创建初始知识分群决策树,并将样本数据填充至初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树之前,还包括:从多个系统、多个平台和多个医疗系统获取医学学科知识数据和临床数据,并对医学学科知识数据和临床数据进行数据处理,获得医学规则数据,数据处理包括数据清理处理、数据集成处理、数据规约处理、数据变换处理、特征提取处理、据脱敏处理和分类处理。具体包括:从多个系统和多个平台获取医学学科知识数据,以及从多个医疗系统获取临床数据;分别对医学学科知识数据和临床数据进行数据预处理,数据预处理包括数据清理处理、数据集成处理、数据规约处理和数据变换处理;对经过数据预处理的医学学科知识数据进行特征提取得到特征信息,以及对经过数据预处理的临床数据进行数据脱敏处理和分类得到第一样本数据和第三样本数据;分析特征信息,按照预设分类规则对特征信息进行分类,获得医学规则数据,预设分类规则包括相同和/或相似病症。通过对医学学科知识数据进行数据清理处理、数据集成处理、数据规约处理和数据变换处理的数据预处理,以保证医学学科知识数据的质量,以便于后续对医学学科知识数据的数据分析和数据挖掘。通过对医疗数据中的某些敏感信息按照脱敏规则进行数据的变形,以对医疗数据中的敏感隐私数据进行可靠保护,在保证敏感隐私数据可靠保护的基础上对医疗数据的质量不产生影响。其中,可通过对一个页面可能执行一条结构化查询语言(structured query language,SQL),以保证各个医疗数据进行脱敏互不干扰。
502、获取待分类人群的特征数据,并通过目标知识分群决策树对特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对处理方案的类别数量进行识别,目标末端知识子节点为特征数据在目标知识分群决策树上的节点位置,处理方案用于指示对相似特征人群的解决渠道。
可选的,上述的对处理方案的类别数量进行识别,具体还包括:计算处理方案之间的相似度获得多个相似度值,并比较多个相似度值是否均在预设范围内;若比较结果为是,则统计处理方案的类别为一个;若比较结果为否,则统计处理方案的类别为多个。通过判断处理方案之间的多个相似度值是否均在预设范围内,以判断处理方案的类别。例如,相似度包括但不限于用药处方的相似度、用药处方结合药品用途的相似度和病症信息结合用药处方的相似度。通过判断处理方案的之间的类型,根据处理方案的类型判断是否构建数据驱动分群模型,以便于根据输入的就诊数据的不同进行不同的操作,进而减少不必要的操作和提高效率,以及保证精准分群。
503、若识别到处理方案的类别数量为一个,则确定待分类人群属于第一特征相似人群,第一特征相似人群为目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群。
504、若识别到处理方案的类别数量至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数,第二分群规则指示根据递归分区算法和处理方案的有效性进行分类。
可选的,上述的基于第二分群规则生成数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数,具体包括:步骤一:将目标末端知识子节点作为数据分群决策树的数据根节点,计算数据根节点对应的处理方案的有效性分值;步骤二:根据每个处理方案的有效性分值对数据根节点进行分支,根据递归分区算法将每个处理方案依次划分到每个分支的数据子节点中,并继续对每个数据子节点对应的处理方案进行有效性分析;步骤三:若当前数据子节点所对应的处理方案的有效性分值大于第一阈值,则对当前数据子节点进行分 支;步骤四:重复执行步骤三,直至每个数据子节点所对应的处理方案的有效性分值小于或等于第一阈值,得到数据分群决策树,数据分群决策树包括多个数据子节点,多个数据子节点中每个末端数据叶子节点对应有一个处理方案;步骤五:确定每个处理方案对应的末端数据叶子节点在数据分群决策树中所处的目标层数。根据目标末端知识子节点对应的处理方案的有效性分值和递归分区算法构建数据分群决策树。例如,第一阈值为60分,将数据根节点的处理方案分支为95分以下和95分以上(即对特征相似人群进行分群),95分以下的有效性分值均大于60分,则将95分以下的有效性分值对应的处理方案进行分支,95分以上的有效性分值均大于60分,则将95分以上的有效性分值对应的处理方案进行分支,依次类推,直到最后分支对应的数据子节点所对应的处理方案的有效性分值小于或等于60分,此时的末端数据叶子节点在数据分群决策树中所处的层数即为目标层数。其中,数据分群决策树在分支时,可通过将处理方案的有效性分值作为参数,将通过测量或测定所得到的对处理方案进行划分的所有有效性分值的划分范围值作为所有观测值,据预置的目标函数和所有观测值对参数进行估计处理得到参数估计值,根据参数估计值评估参数在不同分区的稳定性;若评估结果为不稳定,则分析分区变量,并选择稳定性最低的分区变量进行稳定性评估,直到评估结果为稳定,分区变量用于指示分区的回归量;若评估结果为稳定,则停止评估参数在不同分区的稳定性;基于参数在不同分区的稳定性,根据二分法或变量自适应法获取优化的区域分割规则,二分法用于指示所选子节点的数目为2的一分为二方法,变量自适应法用于指示所选子节点的数目等于分区数量的方法;根据区域分割规则获取与区域分割规则对应的子节点,在与区域分割规则对应的子节点重复执行根据参数估计值评估参数在不同分区的稳定性和根据二分法或变量自适应法获取优化的区域分割规则,直到在评估到参数在不同分区稳定,以将待分类人群和目标知识分群决策树中目标末端子节点对应的特征相似人群进行分群,得到数据分群决策树。采用基于模型的递归分区与临床结局数据强相关的分群规则创建数据分群决策树。通过利用模型的目标函数估计参数和分割点,对每个节点的参数不稳定性进行相应的模型评分,评估应该使用哪个变量进行分区,以利于创建数据分群决策树的非线性关系和自动检测解释变量之间的互动。在创建的过程中:每个节点都与模型相关联;对参数进行稳定性波动分析以评估节点是否需要分裂;若任何一个分区变量存在显著的不稳定性,则将节点分割为局部最优端,并重复该过程。参数的估计可通过得分函数或估计函数进行估计处理。参数在不同分区的稳定性的评估,可通过经验波动过程来捕捉。
可选的,上述的将目标末端知识子节点作为数据分群决策树的数据根节点,计算数据根节点对应的处理方案的有效性分值,具体包括:将目标末端知识子节点作为数据分群决策树的数据根节点;获取数据根节点对应的处理方案的目标特征信息,将目标特征信息拆分成因素因子,计算因素因子在不同水平上的关联度、对预置影响因素的影响度以及因素因子不控制的边缘关联度;通过预置的可压缩准则和可比较准则分析因素因子的关联度、影响度和边缘关联度,得到混杂因子;计算混杂因子与预置影响因素发生概率的回归系数的指数变换值,并以指数变换值作为判定得分,得到数据根节点对应的处理方案的有效性
分值,指数变换值的计算公式如下:
p(y=kx;w)为指数变换值,w为混杂因子与预置影响因素发生的概率,x为混杂因子,k为预置影响因素的类别数量,y为输出的预测值,y是k的条件概率,K为k的最大值,T为时间周期。例如,目标特征信息可为数据分群决策树的父节点的处理方案中的病症信息,预置影响因素为疾病,影响度为关于疾病的危险度,有效性为药物有效性;混杂因子是指一个与暴露和疾病都有联系的因子,它既是所研究疾病的独立危险因子,在非暴露组中它也必定是一个危险因子。对特征相似人群对应的处理方案进行药物有效性分析以获取存在统计分析意义的子节点,从而快速而准确地通过数据分群决策树对知识分群决策树分群之后再进一步分群。在药物有效性分析过程中,先选取出可能的混杂因子,然后将该混杂因子加入到多变量逻辑回归分析中,最后将药物变量的回归系数的 指数变换值作为最终的有效性的判定得分,以获取多个有效值。其中,可通过可压缩准则、可比较准则、虚拟事实模型、图模型作为混杂因子的识别准则,以获取混杂因子。
505、若目标层数为第一层,则将待分类人群划分为第一特征相似人群。
506、若目标层数不为第一层,判断目标层数对应处理方案的有效性分值是否大于第二阈值,第二阈值大于第一阈值。
当识别到目标层数不为第一层时,需要通过数据分群决策树对待分类人群进行进一步分群。获取目标层数对应处理方案的有效性分值,并判断有效性分值是否大于第二阈值,以保证对待分类人群分群的精准度。例如,如图3所示,目标层数为第2层,则获取乙节点的第二特征相似人群对应的处理方案的有效性分值,则分析该有效性分值是否大于第二阈值。
507、若目标层数对应处理方案的有效性分值大于第二阈值,则将有效性分值大于第二阈值的处理方案所位于的数据叶子节点作为目标数据节点,并将待分类人群划分为目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
当目标层数对应处理方案的有效性分值大于第二阈值时,不需要对待分类人群和特征相似人群进行进一步分群,该处已为结合知识驱动和数据驱动对待分类人群进行与临床结局数据强相关的分群所得的最终的特征相似人群,以目标层数对应处理方案的有效性分值大于第二阈值所位于数据决策树的目标数据节点对应特征相似人群作为将待分类人群划分的第二特征相似人群。例如,第二阈值为70分,目标层数对应处理方案的有效性分值为80分,则将对应处理方案的有效性分值为80分的数据叶子节点作为目标数据节点,并将待分类人群划分为目标数据节点对应的处理方案所属的特征相似人群(即第二特征相似人群)。
508、若目标层数对应处理方案的有效性分值小于或者等于第二阈值,则分析目标层数的上一层数对应处理方案的有效性分值。
当目标层数对应处理方案的有效性分值小于或者等于第二阈值时,该目标层数对应的特征相似人群还不是对待分类人群分群所划分所属的最终人群,需要分析目标层数的上一层数对应处理方案的有效性分值。有效性分值小于或者等于第二阈值的处理方案对应的特征相似人群与待分类人群的匹配度低,对有效性分值小于或者等于第二阈值对应的节点或之后的节点不再进行分析,以提高操作效率和分群的精准度。
509、若目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将目标层数的上一层数对应处理方案所位于的数据叶子节点作为目标数据节点,并将待分类人群划分为目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
若目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将待分类人群划分到目标层数的上一层数对应处理方案的有效性分值大于第二阈值对应的数据分群决策树的目标节点的第二特征相似人群,该节点对应的第二特征相似人群为结合知识驱动和数据驱动对待分类人群进行与临床结局数据强相关的分群划分的最终的特征相似人群。若目标层数的上一层数对应处理方案的有效性分值小于或等于第二阈值,则对目标层数的上一层数的上一层数对应处理方案的有效性分值是否大于第二阈值进行判断,以此类推判断,直到分析到对应处理方案的有效性分值大于第二阈值,以处理方案的有效性分值大于第二阈值对应的数据分群决策树的目标节点的特征相似人群作为将待分类人群最终划分的第二特征相似人群。
可选的,上述的若目标层数不为第一层,则将待分类人群划分为第二特征相似人群之后,具体还包括:调用预置的解释模型对第一特征相似人群或第二特征相似人群进行解释处理和可视化图表统计分析处理,得到可解释性特征相似人群,解释处理包括有效性排序处理、匹配度计算处理和匹配度排序处理。具体的,可通过调用预置的解释模型对第一特征相似人群或第二特征相似人群的有效性按照值从大到小的顺序进行排序,得到第一排序信息;计算待分类人群与第一特征相似人群或第二特征相似人群的匹配度,并对匹配度按照值从大到小的顺序进行排序,得到第二排序信息;根据第一排序信息和/或第二排序信息对第一特征相似人群或第二特征相似人群进行排序,并将根据第一排序信息和/或第二排序信息标记在第一特征相似人群或第二特征相似人群上;对标记之后的第一特征相似人群或第二特征相似人群进行可视化图表统计分析处理,生成可视化图表,得到可解释性特征相似人群。通过调用预置的解释模型对第一特征 相似人群或第二特征相似人群,以便于用户便捷地获取将待分类人群划分的特征相似人群各项分群数据。
医学辅助决策系统根据本申请的方法所获得的第一特征相似人群,得到第一特征相似人群对应的提供医学知识以及补充相应的真实世界证据的第一循证医学信息,根据本申请的方法所获得的第二特征相似人群,得到第二特征相似人群对应的提供医学知识、补充相应的真实世界证据和结合数据驱动进行分群的第二循证医学信息之后具体的,包括:调用已创建的解释模型分别获取第一样本数据的第一特征信息、第一循证医学证据信息的第二特征信息和/或第二循证医学证据信息的第一特征信息,第一样本数据指示不遵循医学知识规则的临床结局数据;对第一特征信息和第二特征信息进行分析,得到第一分析数据,以及对第一特征信息和和第三特征信息进行分析,得到第二分析数据;对第一特征信息、第二特征信息和第三特征信息进行统计和分析,生成可视化图表;调用已创建的知识图谱,获取知识图谱中与第二特征信息匹配的第一实物信息,以及获取知识图谱中与第三特征信息匹配的第二实物信息;将第一特征信息、第二特征信息、第三特征信息、第一分析数据、第二分析数据、第一实物信息、第二实物信息和可视化图表生成图文结合的分析报告,获得可解释性数据。将解释说明数据生成图文结合的形式显示,以便于用户更好地获悉可解释性数据的内容。通过对第一循证医学证据信息或第二循证医学证据信息进行解释说明得到可解释性数据,以有效地为辅助医生决策提供更加个性化精确化有效性的真实世界证据。
本申请实施例,通过创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
上面对本申请实施例中基于知识驱动和数据驱动的分群方法进行了描述,下面对本申请实施例中基于知识驱动和数据驱动的分群装置进行描述,请参阅图6,本申请实施例中基于知识驱动和数据驱动的分群装置的一个实施例包括:
第一创建单元601,用于根据医学规则数据创建初始知识分群决策树,并将样本数据填充至初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,第一分群规则用于对临床上的相似特征人群进行分群,样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;
识别单元602,用于获取待分类人群的特征数据,并通过目标知识分群决策树对特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对处理方案的类别数量进行识别,目标末端知识子节点为特征数据在目标知识分群决策树上的节点位置,处理方案用于指示对相似特征人群的解决渠道;
确定单元603,用于若识别到处理方案的类别数量为一个,则确定待分类人群属于第一特征相似人群,第一特征相似人群为目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;
第二创建单元604,用于若识别到处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数,第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;
第一划分单元605,用于若目标层数为第一层,则将待分类人群划分为第一特征相似人群;
第二划分单元606,用于若目标层数不为第一层,则将待分类人群划分为第二特征相似人群,第二特征相似人群为数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
本申请实施例,通过创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
请参阅图7,本申请实施例中基于知识驱动和数据驱动的分群装置的另一个实施例包括:
第一创建单元601,用于根据医学规则数据创建初始知识分群决策树,并将样本数据填充至初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,第一分群规则用于对临床上的相似患者进行分群,样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;
识别单元602,用于获取待分类人群的特征数据,并通过目标知识分群决策树对特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对处理方案的类别数量进行识别,目标末端知识子节点为特征数据在目标知识分群决策树上的节点位置,处理方案用于指示对相似特征人群的解决渠道;
确定单元603,用于若识别到处理方案的类别数量为一个,则确定待分类人群属于第一特征相似人群,第一特征相似人群为目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;
第二创建单元604,用于若识别到处理方案的类别数量至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应数据分群决策树的数据子节点所处的目标层数,第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;
第一划分单元605,用于若目标层数为第一层,则将待分类人群划分为第一特征相似人群;
判断单元606,用于目标层数不为第一层,判断目标层数对应处理方案的有效性分值是否大于第二阈值,第二阈值大于第一阈值;
第三划分单元607,用于若目标层数对应处理方案的有效性分值大于第二阈值,则将有效性分值大于第二阈值的处理方案所位于的数据叶子节点作为目标数据节点,并将待分类人群划分为目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群;
分析单元608,用于若目标层数对应处理方案的有效性分值小于或者等于第二阈值,则分析目标层数的上一层数对应处理方案的有效性分值;
第三划分单元609,用于若目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将目标层数的上一层数对应处理方案所位于的数据叶子节点作为目标数据节点,并将待分类人群划分为目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
可选的,第二创建单元604具体用于:步骤一:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的处理方案的有效性分值;步骤二:根据每个处理方案的有效性分值对所述数据根节点进行分支,根据递归分区算法将每个处理方案依次划分到每个分支的数据子节点中,并继续对每个数据子节点对应的处理方案进行有效性分析;步骤三:若当前数据子节点所对应的处理方案的有效性分值大于第一阈值,则对所述当前数据子节点进行分支;步骤四:重复执行步骤三,直至每个数据子节点所对应的处理方案的有效性分值小于或等于所述第一阈值,得到数据分群决策树,所述数据分群决策树包括多个数据子节点,所述多个数据子节点中每个末端数据叶子节点对应有一个处理方案;步骤五:确定每个处理方案对应的末端数据叶子节点在所述数据分群决策树中所处的目标层数。
可选的,第二创建单元604具体用于:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点;获取所述数据根节点对应的处理方案的目标特征信息,将目标特征信息拆分成因素因子;计算因素因子在不同水平上的关联度、对预置影响因素的影响度以及因素因子不控制的边缘关联度;通过预置的可压缩准则和可比较准则分析因素因子的关联度、影响度和边缘关联度,得到混杂因子;计算混杂因子与预置影响因素发生概率的回归系数的指数变换值,并以指数变换值作为判定得分,得到数据根节点对应的处理方案的有效性分值,指数变换值的计算公式如下:
p(y=kx;w)为指数变换值,w为混杂因子与预置影响因素发生的概率,x为混杂因子,k为预置影响因素的类别数量,y为输出的预测值,y是k 的条件概率,K为k的最大值,T为时间周期。
可选的,基于知识驱动和数据驱动的分群装置,还包括:处理单元,用于从多个系统、多个平台和多个医疗系统获取医学学科知识数据和临床数据,并对医学学科知识数据和临床数据进行数据处理,获得医学规则数据,数据处理包括数据清理处理、数据集成处理、数据规约处理、数据变换处理、特征提取处理、据脱敏处理和分类处理。处理单元具体用于:从多个系统和多个平台获取医学学科知识数据,以及从多个医疗系统获取临床数据;分别对医学学科知识数据和临床数据进行数据预处理,数据预处理包括数据清理处理、数据集成处理、数据规约处理和数据变换处理;对经过数据预处理的医学学科知识数据进行特征提取得到特征信息,以及对经过数据预处理的临床数据进行数据脱敏处理和分类得到第一样本数据和第三样本数据;分析特征信息,按照预设分类规则对特征信息进行分类,获得医学规则数据,预设分类规则包括相同和/或相似病症。
可选的,识别单元602具体用于:计算处理方案之间的相似度获得多个相似度值,并比较多个相似度值是否均在预设范围内;若比较结果为是,则统计处理方案的类别为一个;若比较结果为否,则统计处理方案的类别为多个。
可选的,基于知识驱动和数据驱动的分群装置,还包括:调用单元,用于调用预置的解释模型对第一特征相似人群或第二特征相似人群进行解释处理和可视化图表统计分析处理,得到可解释性特征相似人群,解释处理包括有效性排序处理、匹配度计算处理和匹配度排序处理。调用单元具体用于:调用预置的解释模型对第一特征相似人群或第二特征相似人群的有效性按照值从大到小的顺序进行排序,得到第一排序信息;计算待分类人群与第一特征相似人群或第二特征相似人群的匹配度,并对匹配度按照值从大到小的顺序进行排序,得到第二排序信息;根据第一排序信息和/或第二排序信息对第一特征相似人群或第二特征相似人群进行排序,并将根据第一排序信息和/或第二排序信息标记在第一特征相似人群或第二特征相似人群上;对标记之后的第一特征相似人群或第二特征相似人群进行可视化图表统计分析处理,生成可视化图表,得到可解释性特征相似人群。
本申请实施例,通过创建目标知识分群决策树、确定待分类人群属于第一特征相似人群、创建数据分群决策树和根据目标层数将待分类人群划分为第一特征相似人群或第二特征相似人群等步骤,提高将待分类人群分群到相似特征人群的精准度,进而为医生决策提供合理规范性且全面性精细化的循证医学证据。
上面图6至图7从模块化功能实体的角度对本申请实施例中的基于知识驱动和数据驱动的分群设备进行详细描述,下面从硬件处理的角度对本申请实施例中基于知识驱动和数据驱动的分群设备进行详细描述。
图8是本申请实施例提供的一种基于知识驱动和数据驱动的分群设备的结构示意图,该基于知识驱动和数据驱动的分群设备800可因配置或性能不同而产生比较大的差异,可以包括一个或一个以上处理器(central processing units,CPU)801(例如,一个或一个以上处理器)和存储器809,一个或一个以上存储应用程序807或数据806的存储介质808(例如一个或一个以上海量存储装置)。其中,存储器809和存储介质808可以是短暂存储或持久存储。存储在存储介质808的程序可以包括一个或一个以上模块(图示没标出),每个模块可以包括对签到管理设备中的一系列指令操作。更进一步地,处理器801可以设置为与存储介质808通信,在基于知识驱动和数据驱动的分群设备800上执行存储介质808中的一系列指令操作。
基于知识驱动和数据驱动的分群设备800还可以包括一个或一个以上电源802,一个或一个以上有线或无线网络接口803,一个或一个以上输入输出接口804,和/或,一个或一个以上操作系统805,例如Windows Serve,Mac OS X,Unix,Linux,FreeBSD等等。本领域技术人员可以理解,图8中示出的基于知识驱动和数据驱动的分群设备结构并不构成对基于知识驱动和数据驱动的分群设备的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。处理器801可以执行上述实施例中第一创建单元601、识别单元602、第二创建单元603、第一分析单元604、第一标记单元605、第二分析单元606、第二标记单元607、第一输出单元608、第二输出单元609、分类单元610和第一调用单元611的功能。
下面结合图8对基于知识驱动和数据驱动的分群设备的各个构成部件进行具体的介绍:
处理器801是基于知识驱动和数据驱动的分群设备的控制中心,可以按照基于知识驱动和数据驱动的分群方法进行处理。处理器801利用各种接口和线路连接整个基于知识驱动和数据驱动的分群设备的各个部分,通过运行或执行存储在存储器809内的软件程序和/或模块,以及调用存储在存储器809内的数据,执行基于知识驱动和数据驱动的分群设备的各种功能和处理数据,从而实现提高案件检索系统的可用性的功能。存储介质808和存储器809都是存储数据的载体,本申请实施例中,存储介质808可以是指储存容量较小,但速度快的内存储器,而存储器809可以是储存容量大,但储存速度慢的外存储器。
存储器809可用于存储软件程序以及模块,处理器801通过运行存储在存储器809的软件程序以及模块,从而执行基于知识驱动和数据驱动的分群设备800的各种功能应用以及数据处理。存储器809可主要包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需的应用程序(获取待分类人群的特征数据等)等;存储数据区可存储根据签到管理设备的使用所创建的数据(通过目标知识分群决策树对特征数据进行分群,得到目标末端子节点对应的处理方案,并对处理方案的类别数量进行识别等)等。此外,存储器809可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他易失性固态存储器件。在本申请实施例中提供的基于知识驱动和数据驱动的分群方法程序和接收到的数据流存储在存储器中,当需要使用时,处理器801从存储器809中调用。
在计算机上加载和执行计算机程序指令时,全部或部分地产生按照本申请实施例的流程或功能。计算机可以是通用计算机、专用计算机、计算机网络、或者其他可编程装置。计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一计算机可读存储介质传输,例如,计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、双绞线)或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。计算机可读存储介质可以是计算机能够存储的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。可用介质可以是磁性介质,(例如,软盘、硬盘、磁带)、光介质(例如,光盘)、或者半导体介质(例如固态硬盘(solid state disk,SSD))等。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统,装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
Claims (20)
- 一种基于知识驱动和数据驱动的分群方法,其中,所述方法包括:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
- 根据权利要求1所述的基于知识驱动和数据驱动的分群方法,其中,所述基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,包括:步骤一:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的样本数据中处理方案的有效性分值;步骤二:根据每个处理方案的有效性分值对所述数据根节点进行分支,根据递归分区算法将每个处理方案依次划分到每个分支的数据子节点中,并继续对每个数据子节点对应的处理方案进行有效性分析;步骤三:若当前数据子节点所对应的处理方案的有效性分值大于第一阈值,则对所述当前数据子节点进行分支;步骤四:重复执行步骤三,直至每个数据子节点所对应的处理方案的有效性分值小于或等于所述第一阈值,得到数据分群决策树,所述数据分群决策树包括多个数据子节点,所述多个数据子节点中每个末端数据叶子节点对应有一个处理方案;步骤五:确定每个处理方案对应的末端数据叶子节点在所述数据分群决策树中所处的目标层数。
- 根据权利要求2所述的基于知识驱动和数据驱动的分群方法,其中,所述将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的处理方案的有效性分值,包括:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点;获取所述数据根节点对应的样本数据中处理方案的目标特征信息,将所述目标特征信息拆分成因素因子;计算所述因素因子在不同水平上的关联度、对预置影响因素的影响度以及所述因素因子不控制的边缘关联度;通过预置的可压缩准则和可比较准则分析所述关联度、所述影响度和所述边缘关联度,得到混杂因子;计算所述混杂因子与所述预置影响因素发生概率的回归系数的指数变换值,并以所述指数 变换值作为判定得分,得到所述数据根节点对应的处理方案的有效性分值,所述指数变换值的计算公式如下:所述p(y=k|x;w)为所述指数变换值,所述w为所述混杂因子与所述预置影响因素发生的概率,所述x为所述混杂因子,所述k为所述预置影响因素的类别数量,所述y为输出的预测值,所述y是k的条件概率,所述K为k的最大值,所述T为时间周期。
- 根据权利要求2所述的基于知识驱动和数据驱动的分群方法,其中,所述若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,包括:若所述目标层数不为第一层,判断所述目标层数对应处理方案的有效性分值是否大于第二阈值,所述第二阈值大于所述第一阈值;若所述目标层数对应处理方案的有效性分值大于第二阈值,则将有效性分值大于第二阈值的处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群;若所述目标层数对应处理方案的有效性分值小于或者等于第二阈值,则分析所述目标层数的上一层数对应处理方案的有效性分值;若所述目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将所述目标层数的上一层数对应处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
- 根据权利要求1所述的基于知识驱动和数据驱动的分群方法,其中,所述根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树之前,所述方法还包括:从多个系统、多个平台和多个医疗系统获取医学学科知识数据和临床数据,并对所述医学学科知识数据和所述临床数据进行数据处理,获得医学规则数据,所述数据处理包括数据清理处理、数据集成处理、数据规约处理、数据变换处理、特征提取处理、据脱敏处理和分类处理。
- 根据权利要求1所述的基于知识驱动和数据驱动的分群方法,其中,所述对所述处理方案的类别数量进行识别,包括:计算所述处理方案之间的相似度获得多个相似度值,并比较多个所述相似度值是否均在预设范围内;若比较结果为是,则统计所述处理方案的类别为一个;若比较结果为否,则统计所述处理方案的类别为多个。
- 根据权利要求1-6中任一项所述的基于知识驱动和数据驱动的分群方法,其中,所述若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群之后,所述方法还包括:调用预置的解释模型对所述第一特征相似人群或所述第二特征相似人群进行解释处理和可视化图表统计分析处理,得到可解释性特征相似人群,所述解释处理包括有效性排序处理、匹配度计算处理和匹配度排序处理。
- 一种基于知识驱动和数据驱动的分群装置,其中,所述装置包括:第一创建单元,用于根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;识别单元,用于获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所 述处理方案用于指示对相似特征人群的解决渠道;确定单元,用于若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;第二创建单元,用于若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类各分支的末端子节点的;第一划分单元,用于若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;第二划分单元,用于若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
- 一种基于知识驱动和数据驱动的分群设备,其中,所述基于知识驱动和数据驱动的分群设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如下所述的基于知识驱动和数据驱动的分群方法的步骤:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
- 根据权利要求9所述的基于知识驱动和数据驱动的分群设备,其中,所述基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,包括以下步骤:步骤一:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的样本数据中处理方案的有效性分值;步骤二:根据每个处理方案的有效性分值对所述数据根节点进行分支,根据递归分区算法将每个处理方案依次划分到每个分支的数据子节点中,并继续对每个数据子节点对应的处理方案进行有效性分析;步骤三:若当前数据子节点所对应的处理方案的有效性分值大于第一阈值,则对所述当前数据子节点进行分支;步骤四:重复执行步骤三,直至每个数据子节点所对应的处理方案的有效性分值小于或等 于所述第一阈值,得到数据分群决策树,所述数据分群决策树包括多个数据子节点,所述多个数据子节点中每个末端数据叶子节点对应有一个处理方案;步骤五:确定每个处理方案对应的末端数据叶子节点在所述数据分群决策树中所处的目标层数。
- 根据权利要求10所述的基于知识驱动和数据驱动的分群设备,其中,所述将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的处理方案的有效性分值,包括以下步骤:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点;获取所述数据根节点对应的样本数据中处理方案的目标特征信息,将所述目标特征信息拆分成因素因子;计算所述因素因子在不同水平上的关联度、对预置影响因素的影响度以及所述因素因子不控制的边缘关联度;通过预置的可压缩准则和可比较准则分析所述关联度、所述影响度和所述边缘关联度,得到混杂因子;计算所述混杂因子与所述预置影响因素发生概率的回归系数的指数变换值,并以所述指数变换值作为判定得分,得到所述数据根节点对应的处理方案的有效性分值,所述指数变换值的计算公式如下:所述p(y=k|x;w)为所述指数变换值,所述w为所述混杂因子与所述预置影响因素发生的概率,所述x为所述混杂因子,所述k为所述预置影响因素的类别数量,所述y为输出的预测值,所述y是k的条件概率,所述K为k的最大值,所述T为时间周期。
- 根据权利要求10所述的基于知识驱动和数据驱动的分群设备,其中,所述若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,包括以下步骤:若所述目标层数不为第一层,判断所述目标层数对应处理方案的有效性分值是否大于第二阈值,所述第二阈值大于所述第一阈值;若所述目标层数对应处理方案的有效性分值大于第二阈值,则将有效性分值大于第二阈值的处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群;若所述目标层数对应处理方案的有效性分值小于或者等于第二阈值,则分析所述目标层数的上一层数对应处理方案的有效性分值;若所述目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将所述目标层数的上一层数对应处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
- 根据权利要求9所述的基于知识驱动和数据驱动的分群设备,其中,所述根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树之前,还包括以下步骤:从多个系统、多个平台和多个医疗系统获取医学学科知识数据和临床数据,并对所述医学学科知识数据和所述临床数据进行数据处理,获得医学规则数据,所述数据处理包括数据清理处理、数据集成处理、数据规约处理、数据变换处理、特征提取处理、据脱敏处理和分类处理。
- 根据权利要求9所述的基于知识驱动和数据驱动的分群设备,其中,所述对所述处理方案的类别数量进行识别,包括以下步骤:计算所述处理方案之间的相似度获得多个相似度值,并比较多个所述相似度值是否均在预设范围内;若比较结果为是,则统计所述处理方案的类别为一个;若比较结果为否,则统计所述处理方案的类别为多个。
- 根据权利要求9-14中任一项所述的基于知识驱动和数据驱动的分群设备,其中,所述若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群之后,还包括以下步骤:调用预置的解释模型对所述第一特征相似人群或所述第二特征相似人群进行解释处理和可视化图表统计分析处理,得到可解释性特征相似人群,所述解释处理包括有效性排序处理、匹配度计算处理和匹配度排序处理。
- 一种计算机可读存储介质,其中,包括指令,当所述指令在计算机上运行时,使得计算机执行如下所述的基于知识驱动和数据驱动的分群方法的步骤:根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树,所述医学规则数据用于指示对已归档的医学文献、医学共识、医学指南和专家的临床经验的医学学科知识数据进行规则提取处理所得的第一分群规则,所述第一分群规则用于对临床上的相似特征人群进行分群,所述样本数据包括遵循医学知识规则的临床上的相似特征人群对应的临床结局数据;获取待分类人群的特征数据,并通过所述目标知识分群决策树对所述特征数据进行分群,得到目标末端知识子节点对应的处理方案,并对所述处理方案的类别数量进行识别,所述目标末端知识子节点为所述特征数据在所述目标知识分群决策树上的节点位置,所述处理方案用于指示对相似特征人群的解决渠道;若识别到所述处理方案的类别数量为一个,则确定所述待分类人群属于第一特征相似人群,所述第一特征相似人群为所述目标知识分群决策树中目标末端知识子节点对应的处理方案所属的特征相似人群;若识别到所述处理方案的类别数量为至少两个,则基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数,所述第二分群规则指示根据递归分区算法和处理方案的有效性进行分类;若所述目标层数为第一层,则将所述待分类人群划分为所述第一特征相似人群;若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群,所述第二特征相似人群为所述数据分群决策树中目标数据节点对应的处理方案所属的特征相似人群,每个目标数据节点对应不同的特征相似人群。
- 根据权利要求16所述的计算机可读存储介质,其中,所述基于知识驱动和数据驱动的分群的指令被计算机执行所述基于第二分群规则生成数据分群决策树,并确定每个处理方案对应所述数据分群决策树的数据子节点所处的目标层数的步骤时,包括以下步骤:步骤一:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的样本数据中处理方案的有效性分值;步骤二:根据每个处理方案的有效性分值对所述数据根节点进行分支,根据递归分区算法将每个处理方案依次划分到每个分支的数据子节点中,并继续对每个数据子节点对应的处理方案进行有效性分析;步骤三:若当前数据子节点所对应的处理方案的有效性分值大于第一阈值,则对所述当前数据子节点进行分支;步骤四:重复执行步骤三,直至每个数据子节点所对应的处理方案的有效性分值小于或等于所述第一阈值,得到数据分群决策树,所述数据分群决策树包括多个数据子节点,所述多个数据子节点中每个末端数据叶子节点对应有一个处理方案;步骤五:确定每个处理方案对应的末端数据叶子节点在所述数据分群决策树中所处的目标层数。
- 根据权利要求17所述的计算机可读存储介质,其中,所述基于知识驱动和数据驱动的分群的指令被计算机执行所述将所述目标末端知识子节点作为所述数据分群决策树的数据根节点,计算所述数据根节点对应的处理方案的有效性分值的步骤时,包括以下步骤:将所述目标末端知识子节点作为所述数据分群决策树的数据根节点;获取所述数据根节点对应的样本数据中处理方案的目标特征信息,将所述目标特征信息拆 分成因素因子;计算所述因素因子在不同水平上的关联度、对预置影响因素的影响度以及所述因素因子不控制的边缘关联度;通过预置的可压缩准则和可比较准则分析所述关联度、所述影响度和所述边缘关联度,得到混杂因子;计算所述混杂因子与所述预置影响因素发生概率的回归系数的指数变换值,并以所述指数变换值作为判定得分,得到所述数据根节点对应的处理方案的有效性分值,所述指数变换值的计算公式如下:所述p(y=k|x;w)为所述指数变换值,所述w为所述混杂因子与所述预置影响因素发生的概率,所述x为所述混杂因子,所述k为所述预置影响因素的类别数量,所述y为输出的预测值,所述y是k的条件概率,所述K为k的最大值,所述T为时间周期。
- 根据权利要求17所述的计算机可读存储介质,其中,所述基于知识驱动和数据驱动的分群的指令被计算机执行所述若所述目标层数不为第一层,则将所述待分类人群划分为第二特征相似人群的步骤时,包括以下步骤:若所述目标层数不为第一层,判断所述目标层数对应处理方案的有效性分值是否大于第二阈值,所述第二阈值大于所述第一阈值;若所述目标层数对应处理方案的有效性分值大于第二阈值,则将有效性分值大于第二阈值的处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群;若所述目标层数对应处理方案的有效性分值小于或者等于第二阈值,则分析所述目标层数的上一层数对应处理方案的有效性分值;若所述目标层数的上一层数对应处理方案的有效性分值大于第二阈值,则将所述目标层数的上一层数对应处理方案所位于的数据叶子节点作为目标数据节点,并将所述待分类人群划分为所述目标数据节点对应的处理方案所属的特征相似人群,得到第二特征相似人群。
- 根据权利要求16所述的计算机可读存储介质,其中,所述基于知识驱动和数据驱动的分群的指令被计算机执行所述根据医学规则数据创建初始知识分群决策树,并将样本数据填充至所述初始知识分群决策树各分支的末端知识子节点,得到目标知识分群决策树之前,还包括以下步骤:从多个系统、多个平台和多个医疗系统获取医学学科知识数据和临床数据,并对所述医学学科知识数据和所述临床数据进行数据处理,获得医学规则数据,所述数据处理包括数据清理处理、数据集成处理、数据规约处理、数据变换处理、特征提取处理、据脱敏处理和分类处理。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910991392.5A CN110929752B (zh) | 2019-10-18 | 2019-10-18 | 基于知识驱动和数据驱动的分群方法及相关设备 |
| CN201910991392.5 | 2019-10-18 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021073259A1 true WO2021073259A1 (zh) | 2021-04-22 |
Family
ID=69849217
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/111413 Ceased WO2021073259A1 (zh) | 2019-10-18 | 2020-08-26 | 基于知识驱动和数据驱动的分群方法及相关设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110929752B (zh) |
| WO (1) | WO2021073259A1 (zh) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113590923A (zh) * | 2021-07-28 | 2021-11-02 | 深圳市酷开网络科技股份有限公司 | 人群圈定任务拆分方法、装置、设备及存储介质 |
| CN114238551A (zh) * | 2021-12-15 | 2022-03-25 | 泰康保险集团股份有限公司 | 知识的化简与展示方法、装置、电子设备及可读介质 |
| CN114925833A (zh) * | 2022-04-20 | 2022-08-19 | 中国人民解放军91977部队 | 一种基于能力数据底图的目标状态规律知识挖掘方法 |
| CN115186113A (zh) * | 2022-09-07 | 2022-10-14 | 粤港澳大湾区数字经济研究院(福田) | 一种筛选指南文本的方法、装置、设备及存储介质 |
| CN119557399A (zh) * | 2024-11-11 | 2025-03-04 | 北京百度网讯科技有限公司 | 应答数据的生成方法、装置及计算机程序产品 |
Families Citing this family (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110929752B (zh) * | 2019-10-18 | 2023-06-20 | 平安科技(深圳)有限公司 | 基于知识驱动和数据驱动的分群方法及相关设备 |
| CN111666494B (zh) * | 2020-05-13 | 2022-08-12 | 平安科技(深圳)有限公司 | 分群决策模型生成、分群处理方法、装置、设备及介质 |
| CN111696661B (zh) * | 2020-05-13 | 2024-09-24 | 平安科技(深圳)有限公司 | 患者分群模型构建方法、患者分群方法及相关设备 |
| CN112016979B (zh) * | 2020-09-08 | 2023-07-18 | 平安科技(深圳)有限公司 | 用户分群方法、装置、设备和计算机可读存储介质 |
| CN112116028B (zh) * | 2020-09-29 | 2024-04-26 | 联想(北京)有限公司 | 模型决策解释实现方法、装置及计算机设备 |
| CN114611598A (zh) * | 2022-03-07 | 2022-06-10 | 兴业消费金融股份公司 | 一种分群处理方法及装置 |
| CN114415975B (zh) * | 2022-03-28 | 2022-06-21 | 支付宝(杭州)信息技术有限公司 | 针对知识图谱进行划分分区的方法及装置 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130226616A1 (en) * | 2011-10-13 | 2013-08-29 | The Board of Trustees for the Leland Stanford, Junior, University | Method and System for Examining Practice-based Evidence |
| CN104699939A (zh) * | 2013-12-05 | 2015-06-10 | 国际商业机器公司 | 组合知识和数据驱动的洞察的患者风险分层的方法和系统 |
| CN108492887A (zh) * | 2018-04-13 | 2018-09-04 | 合肥工业大学 | 医疗知识图谱构建方法及装置 |
| CN108984761A (zh) * | 2018-07-19 | 2018-12-11 | 南昌工程学院 | 一种基于模型和领域知识驱动的信息处理系统 |
| US20190272654A1 (en) * | 2018-03-02 | 2019-09-05 | International Business Machines Corporation | Systems and methods for constructing clinical pathways within a gui |
| CN110929752A (zh) * | 2019-10-18 | 2020-03-27 | 平安科技(深圳)有限公司 | 基于知识驱动和数据驱动的分群方法及相关设备 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106384282A (zh) * | 2016-06-14 | 2017-02-08 | 平安科技(深圳)有限公司 | 构建决策模型的方法和装置 |
| CN109308564A (zh) * | 2018-08-01 | 2019-02-05 | 平安科技(深圳)有限公司 | 人群绩效等级识别方法、装置、存储介质及计算机设备 |
| CN109992699B (zh) * | 2019-02-28 | 2023-08-11 | 平安科技(深圳)有限公司 | 用户群的优化方法及装置、存储介质、计算机设备 |
-
2019
- 2019-10-18 CN CN201910991392.5A patent/CN110929752B/zh active Active
-
2020
- 2020-08-26 WO PCT/CN2020/111413 patent/WO2021073259A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130226616A1 (en) * | 2011-10-13 | 2013-08-29 | The Board of Trustees for the Leland Stanford, Junior, University | Method and System for Examining Practice-based Evidence |
| CN104699939A (zh) * | 2013-12-05 | 2015-06-10 | 国际商业机器公司 | 组合知识和数据驱动的洞察的患者风险分层的方法和系统 |
| US20190272654A1 (en) * | 2018-03-02 | 2019-09-05 | International Business Machines Corporation | Systems and methods for constructing clinical pathways within a gui |
| CN108492887A (zh) * | 2018-04-13 | 2018-09-04 | 合肥工业大学 | 医疗知识图谱构建方法及装置 |
| CN108984761A (zh) * | 2018-07-19 | 2018-12-11 | 南昌工程学院 | 一种基于模型和领域知识驱动的信息处理系统 |
| CN110929752A (zh) * | 2019-10-18 | 2020-03-27 | 平安科技(深圳)有限公司 | 基于知识驱动和数据驱动的分群方法及相关设备 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113590923A (zh) * | 2021-07-28 | 2021-11-02 | 深圳市酷开网络科技股份有限公司 | 人群圈定任务拆分方法、装置、设备及存储介质 |
| CN114238551A (zh) * | 2021-12-15 | 2022-03-25 | 泰康保险集团股份有限公司 | 知识的化简与展示方法、装置、电子设备及可读介质 |
| CN114925833A (zh) * | 2022-04-20 | 2022-08-19 | 中国人民解放军91977部队 | 一种基于能力数据底图的目标状态规律知识挖掘方法 |
| CN115186113A (zh) * | 2022-09-07 | 2022-10-14 | 粤港澳大湾区数字经济研究院(福田) | 一种筛选指南文本的方法、装置、设备及存储介质 |
| CN119557399A (zh) * | 2024-11-11 | 2025-03-04 | 北京百度网讯科技有限公司 | 应答数据的生成方法、装置及计算机程序产品 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110929752B (zh) | 2023-06-20 |
| CN110929752A (zh) | 2020-03-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2021073259A1 (zh) | 基于知识驱动和数据驱动的分群方法及相关设备 | |
| CN111949787B (zh) | 基于知识图谱的自动问答方法、装置、设备及存储介质 | |
| US11101029B2 (en) | Systems and methods for predicting medications to prescribe to a patient based on machine learning | |
| Karrar | The effect of using data pre-processing by imputations in handling missing values | |
| JP7076483B2 (ja) | データモデルを構築する方法、装置、デバイス及び媒体 | |
| EP3985559A1 (en) | Entity semantics relationship classification | |
| US10235633B2 (en) | Method and system for linking heterogeneous data sources | |
| CN113707297A (zh) | 医疗数据的处理方法、装置、设备及存储介质 | |
| CN113707303A (zh) | 基于知识图谱的医疗问题解答方法、装置、设备及介质 | |
| EP3582118A1 (en) | Method and apparatus for training classification model | |
| CN109918498B (zh) | 一种问题入库方法和装置 | |
| CN114647713A (zh) | 基于虚拟对抗的知识图谱问答方法、设备及存储介质 | |
| WO2019136993A1 (zh) | 文本相似度计算方法、装置、计算机设备和存储介质 | |
| WO2022222942A1 (zh) | 问答记录生成方法、装置、电子设备及存储介质 | |
| US11132622B2 (en) | Autonomous system for hypotheses generation | |
| CN108519971B (zh) | 一种基于平行语料库的跨语种新闻主题相似性对比方法 | |
| CN114330335B (zh) | 关键词抽取方法、装置、设备及存储介质 | |
| US20170193197A1 (en) | System and method for automatic unstructured data analysis from medical records | |
| US20210166822A1 (en) | Method and apparatus for selecting radiology reports for image labeling by modality and anatomical region of interest | |
| CA3164921A1 (en) | Unsupervised taxonomy extraction from medical clinical trials | |
| EP3576024A1 (en) | Accessible machine learning | |
| CN114139530A (zh) | 同义词提取方法、装置、电子设备及存储介质 | |
| CN112149409A (zh) | 医疗词云生成方法、装置、计算机设备及存储介质 | |
| CN113888265A (zh) | 产品推荐方法、装置、设备及计算机可读存储介质 | |
| CN112580348A (zh) | 政策文本关联性分析方法及系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20876666 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20876666 Country of ref document: EP Kind code of ref document: A1 |




