WO2011122897A2 - 사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법 - Google Patents

사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법 Download PDF

Info

Publication number
WO2011122897A2
WO2011122897A2 PCT/KR2011/002264 KR2011002264W WO2011122897A2 WO 2011122897 A2 WO2011122897 A2 WO 2011122897A2 KR 2011002264 W KR2011002264 W KR 2011002264W WO 2011122897 A2 WO2011122897 A2 WO 2011122897A2
Authority
WO
WIPO (PCT)
Prior art keywords
term
node
concept
query
graph
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2011/002264
Other languages
English (en)
French (fr)
Other versions
WO2011122897A3 (ko
Inventor
김홍기
김응희
송승재
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SNU R&DB Foundation
Original Assignee
SNU R&DB Foundation
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by SNU R&DB Foundation filed Critical SNU R&DB Foundation
Publication of WO2011122897A2 publication Critical patent/WO2011122897A2/ko
Publication of WO2011122897A3 publication Critical patent/WO2011122897A3/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor

Definitions

  • the present invention forms a concept grid for terms used in relation to each other in subjects such as a plurality of medical institutions, and supports a concept grid based query term mapping support system for mapping query terms to terms used in relation to each other. And to a method.
  • the present invention is a conceptual grid based on the terms related to the query term in a hierarchical graph represented by the term node, and using the query term as the root node to indicate the relevance of the upper term node to each term node with support or confidence.
  • the present invention relates to a query term mapping support system and method.
  • the present invention relates to a concept grid based query term mapping support system and method for displaying the relationship between terms used as support or confidence, but calculates the degree of support or confidence reflecting the grade of the subject using the term will be.
  • ⁇ colitis '' is called ⁇ appendicitis '', ⁇ appendicitis '', ⁇ intestinal otitis '', and ⁇ collectitis '', depending on the situation even though there is a medical term of ⁇ appendicitis.
  • '' Infarction is also referred to as 'cardiac infarction', 'infarction of heart', 'MI' or 'Heart Attack' depending on the situation.
  • 'address' it is used in the meaning of 'Address' in the department of clinic, but in the meaning of 'Chief Complain'.
  • the complexity of medical terms is due to the complexity of the medical field itself, and the interface with reference terminology (Reference Terminology) that distinguishes these medical terms and serves as a standard for communication between medical personnel or in different fields. Interface Terminology is required.
  • the terminology system used in the medical field is a collection of medical concepts used in a specific field, and includes a definition of the concept, a term referring to the concept, and a relationship between the concepts.
  • the medical term system can be classified into reference terminology, interface terminology, classification, etc. according to its purpose and characteristics, and the medical term used in the clinical field is local term. It is called.
  • Reference terminology is a set of formal representations of medical concepts and their inter-relationships. Reference term systems include SNOMED CT, RxNorm, NCI Thesaurus, ICNP, and the like.
  • Interface terminology is a terminology proposed to make it easier for medical professionals and clinical support professionals to acquire, collect, transmit, and process information in clinical documentation. This includes all terms on the terminology and treatment sheet.
  • the interface term system is a set of terminology and knowledge used by medical staff in a specific hospital to refer to a specific medical concept. It serves as an interface between the medical staff and the reference term system.
  • Classification is an interdisciplinary category that encompasses everything in order to integrate data at a predetermined level of differentiation for a specific purpose.
  • the purpose of classification is to conduct statistics or research in medical services.
  • Diseases KCD (Korean Standard Classification of Disease).
  • mapping is the variety of local terms used in the medical field. That is, within a hospital, local terms used by doctors in different departments are different, and local terms managed in each hospital are also different in each hospital. Local terms are terms that are directly used by medical staff such as doctors, and interface term system is a term system system that collects local terms, assigns meanings, defines them by concept, and mounts contextual information on medical concepts. Because it is a term used in, its influence is very large.
  • An object of the present invention is to solve the problems as described above, by analyzing the relationship between terms from terms used such as medical terms used in relation to each other to intuitively grasp the inherent relationship between the terms used in the field. It is to provide a conceptual grid-based query term mapping support system and method.
  • an object of the present invention is to form a concept grid for the terms used in a plurality of subjects in relation to each other, to support a mapping of the terms used in relation to each other for a given query term. It is to provide a term mapping support system and method.
  • an object of the present invention is to express the relationship between terms used as support or confidence, concept grid based query term mapping support system and method for calculating the degree of support or confidence reflecting the grade of the subject using the term To provide.
  • the present invention forms a concept grid for terms used in a plurality of subjects in relation to each other, and supports a concept grid based query that supports mapping of query terms to terms used in relation to each other.
  • a term mapping support system comprising: a terminology collection unit for receiving terminology used by each user as a set (hereinafter, referred to as terminology set); Form a conceptual grid graph that consists of nodes (hereinafter, conceptual nodes) that use the term set and the set of subjects that use the term set in relation to each other, and the term set of the lower concept node inherits all the term sets of the parent concept node.
  • Forming part And receiving a query term and generating a related term graph configured as a term node for the query term, wherein the term node corresponds to one of the concept nodes of the conceptual lattice graph, and the concept node corresponding to the term node corresponds to the term.
  • the root node (hereinafter referred to as query node) of the related term graph corresponds to the highest concept node including the query term, and uses the query term as its own term.
  • the term node (hereinafter referred to as association node) that does not correspond to the highest level conceptual node includes a term excluding terms included in all parent nodes of the concept node corresponding to the highest level conceptual node.
  • the parent node relationship between the query nodes is characterized by the same as the parent node relationship between the corresponding conceptual nodes.
  • the present invention is a concept grid-based query term mapping support system, the term node is characterized by having a degree of support or confidence calculated using the number of medical institutions of the conceptual node corresponding to the node.
  • the terminology collection unit receives a grade weight of each user, and the support or confidence level reflects the grade weight.
  • the present invention is a conceptual grid-based query term mapping support system, if the term set of the concept node corresponding to the query node includes a term other than the query term, the related term graph includes the other term.
  • the term node (hereinafter, referred to as an implicit node) is configured as a lower node of the query node, wherein the implication node corresponds to a concept node corresponding to the query node.
  • the present invention is a concept lattice-based query term mapping support system, wherein the rating weight value is an integer multiple, the use term set for each user is further overlapped by the rating weight, the support or confidence degree is obtained. It is done.
  • the support or confidence of the term node is calculated as the ratio of the sum of the grade weights of the subjects of the corresponding concept node.
  • the entire subject of the support when the concept grid graph is divided into at least two graphs, the entire subject of the support includes a concept node corresponding to the term node. It is characterized by being limited to the subject used in the graph.
  • the system receives the query term, and displays the relationship between the related terminology for the query term in the related term graph of the query term, and display the degree of relationship with the upper node of each term node It characterized in that it further comprises a query term display.
  • the present invention is a concept lattice-based query term mapping support system
  • the query term display unit is characterized in that the use terms and support or confidence level corresponding to each term node step by step according to the depth of the related term graph, characterized in that do.
  • the present invention is a concept grid-based query term mapping support system
  • the query term display unit is characterized in that it displays only the term node that the degree of support or confidence in the related term graph is greater than the minimum support or the minimum confidence.
  • the present invention forms a concept grid for the terms used in a plurality of subjects used in relation to each other, the concept grid-based query term mapping support method that supports mapping the query terms to the terms used in relation to each other Regarding, (a) receiving the terminology used in relation to each other in each user as a set (hereinafter term set); (b) A conceptual grid graph is formed of nodes (hereinafter, conceptual nodes) consisting of a term set and a set of subjects that use the term set in common. A term set of a child concept node inherits all of a term set of a parent concept node.
  • the node may comprise the step of including the query term.
  • the present invention is a concept grid-based query term mapping support method, the term node is characterized by having a degree of support or confidence calculated using the number of medical institutions of the conceptual node corresponding to the node.
  • step (a) the class weight of each user is input, and the support or the degree of confidence reflects the class weight.
  • the present invention is a concept lattice-based query term mapping support method, wherein the rating weight value is an integer multiple, the set of terms for each user is further duplicated by the rating weight, the support or confidence is obtained. It is done.
  • the support or confidence of the term node is calculated as the ratio of the sum of the grade weight of the subjects of the corresponding concept node.
  • the present invention is a concept lattice-based query term mapping support method, when the concept lattice graph is divided into at least two graphs, the entire subject of the support is divided into a concept node corresponding to the term node It is characterized by being limited to the subject used in the graph.
  • the present invention is a concept grid-based query term mapping support method, the method, (d) receives a query term and displays the relationship between the related terminology for the query term in the related term graph of the query term, It further comprises a query term step for indicating the degree of relationship with the parent node of each term node with support or confidence.
  • the present invention also relates to a computer-readable recording medium having recorded thereon a program for performing the concept lattice based query term mapping support method.
  • the relationship between the terms used to express the degree of support or confidence but calculates the degree of support or confidence by reflecting the grade of the subject using the term By doing so, the effect of deriving a more accurate relationship between terms can be obtained by reflecting the influence of each user in the field.
  • FIG. 1 is a diagram showing the configuration of an entire system for implementing the present invention.
  • FIG. 2 is a block diagram of the configuration of the concept grid based query term mapping support system according to an embodiment of the present invention.
  • FIG. 3 is a diagram illustrating a format context for a term used in relation to each other in a subject according to a first embodiment of the present invention.
  • FIG. 4 is a diagram illustrating a concept extracted from the format context of FIG. 3 according to the first embodiment of the present invention.
  • FIG. 5 is a diagram illustrating a conceptual lattice graph according to a first embodiment of the present invention.
  • FIG. 6 is a diagram illustrating a related term graph according to a first exemplary embodiment of the present invention.
  • FIG. 7 is a table illustrating support of related terms for a query term using a related term graph generated according to the first embodiment of the present invention.
  • FIG. 8 is a table illustrating a format context for a term used in relation to each other in a subject according to a second embodiment of the present invention.
  • FIG. 9 is a table illustrating an example of a graph of related terms and a support of related terms according to a second embodiment of the present invention.
  • FIG. 10 illustrates a format context having at least two different mapping information according to the third embodiment of the present invention.
  • FIG. 11 is a flowchart illustrating a concept grid-based query term mapping support method according to an embodiment of the present invention.
  • FIG. 12 is a diagram illustrating an algorithm for generating a related term graph according to an embodiment of the present invention.
  • the query term mapping support system may be implemented as a server system on a network or a program system on a computer terminal.
  • an example of the whole system for the implementation of the present invention is composed of the user terminal 10 and the query term mapping support system 30 and connected to each other network 20.
  • a database 40 for storing necessary data may be further provided.
  • the user subject terminal 10 is a general computing terminal such as a PC, a notebook, a netbook, a PDA, a mobile, etc. used by a subject such as a medical person or a medical institution.
  • the medical personnel may use the user terminal 10 to transmit a plurality of usage terms used in relation to each other in the medical institution to the mapping support system 30, or refer to related terms for a specific term (or query term). Request to the mapping support system 30.
  • the query term mapping support system 30 is a server that is connected to the network 20 as a conventional server and provides a service that searches for related terms with respect to the queried term and shows the relatedness.
  • the mapping support system 30 provides a registration service for registering a plurality of terms used by each user in relation to each other.
  • the mapping support system 30 may be implemented as a web server or a web application server that provides each service as a web page on the Internet.
  • mapping support system 30 suggests terms related to the query term, users (or users) map the terms for the query term with reference to the suggested results.
  • the relevance of the terms provided in the mapping support system 30 is not limited to the term mapping.
  • mapping can also be used for analyzing terminology. Therefore, the present invention is not limited to supporting mapping, but hereinafter, it is described as supporting mapping for convenience of description.
  • the database 40 is a general storage medium for storing data required by the mapping support system 30.
  • the database 40 includes conceptual grid graphs formed from the terms received from each user, or other related terms for each term. It stores related term graph of hierarchy.
  • FIG. 1b another example of the whole system for the implementation of the present invention is composed of a query term mapping support system 30 in the form of a program installed in the computer terminal 13. That is, each function of the mapping support system 30 is implemented as a computer program and is installed in the computer terminal 13 so that a plurality of terms used by the user 14 in association with each other are used by the computer terminal 13. If the input is received through the input device of the) or is requested to refer to the related terms for the query term, the result is output through the output device of the computer terminal 13. Meanwhile, data necessary for the mapping support system 30 is stored and used in a storage space such as a hard disk of the computer terminal 13.
  • the conceptual grid-based query term mapping support system 30 includes a terminology collection unit 31, a concept grid forming unit 32, and a query graph generator ( 33).
  • the query term display unit 32 may be further included.
  • it may be configured to further include a database (or storage space) 40 for storing data.
  • the terminology collection unit 31 receives the terminology used in association with each other from each user as a set. Terms used in the same sense are terms that are related to each other.
  • hospital A, B, C means appendicitis, 3 as appendicitis (t-1), appendenditis (t-2), and appendicitis (t-4) Eggplant is used, and hospital D uses only two types of appendicitis (t-1) and affendicity (t-2), and hospital E is used for appendicitis (t-1) and stooltail (t-3). Only two are used. That is, the terminology collection unit 31 receives a terminology set as shown in FIG. 3a from each subject.
  • the conceptual grid forming unit 32 displays the term set and a set of subjects that commonly use the term set as nodes (hereinafter referred to as concept nodes), and the term sets of the lower concept nodes all represent the term sets of the upper concept nodes. Construct an inherited conceptual grid.
  • the formal context K (G, M, I) consists of the subject set G and the term set M, and the binomial relationship I ⁇ G ⁇ M between G and M. That is, the elements of G and M represent the subjects of the format context and the terminology that each subject can use.
  • a user g uses the term m, it is represented by gIm or (g, m) ⁇ I, and g means that m is used.
  • the formal context may be represented in the form of a data table, and the heads of the rows and columns of the table are composed of subjects and terms used in the formal context.
  • the head of the rows and columns of the table are composed of subjects and terms used in the formal context.
  • Intent (O): ⁇ t ⁇ M
  • ⁇ o ⁇ O: (o, t) ⁇ I ⁇ O '
  • extent (T): ⁇ o ⁇ G
  • ⁇ t ⁇ T: (o, t ) ⁇ I ⁇ T '.
  • each concept is defined as a pair of the form (O, T), and in particular, the user set O is an extent of the term set T, and at the same time, the term set T becomes an intent of the user set O.
  • All concepts B (K) extracted from the format context of FIG. 3B are the same as those of FIG. 4.
  • c1 to c5 represent the concepts shown in FIG. 3B.
  • the concepts extracted from the given format context can define the parent-child concept relationship as definitions 3 and 4 based on extent or intent.
  • the superconcept-subconcept relation (O1, T1) ⁇ (O2, T2) is a kind of semi-sequential relationship. (partial order relation), defined as
  • a set of all concepts that exist in formal context K (G, M, I) is called B (K), and for concepts (X1, Y1) and (X2, Y2) ⁇ B (K), (X1, Y1) When (X3, Y3) satisfying ⁇ (X2, Y2) and satisfying (X1, Y1) ⁇ (X3, Y3) ⁇ (X2, Y2) does not exist in B (K), (X1, Y1 ) Is called the lower neighbor of (X2, Y2), (X2, Y2) is called the upper neighbor of (X1, Y1), and (X1, Y1) It is written as (X2, Y2).
  • Is called a concept lattice (Concept Lattice or Galois Lattice) and is expressed as L: (B (K), E ⁇ ).
  • a concept grid containing information extracted from the format context described in FIG. 3B and the hierarchical relationship between the concepts may be displayed as shown in FIG. 5 using a graph.
  • Each node of the graph represents a concept node, and a label above the concept node indicates a term (or intent) of the concept node, and a label below the node indicates a subject (or extent) of the concept node. That is, the conceptual grid graph displays a set of subjects that use the term set and the term set in common as nodes (hereinafter, conceptual nodes).
  • the lower conceptual node includes all of the terms used in its upper conceptual node. That is, the term set of the lower conceptual node inherits all the term sets of the upper conceptual node.
  • the parent concept node includes all the subjects of its child concept nodes.
  • the hospitals using the terms Q and R can be obtained by extent (Q) and extent (R), respectively.
  • extent (Q) ⁇ A, B, C ⁇
  • additional information may be extracted from the data of the format context given in FIG. 3B, "a hospital using appendicitis also uses appendicitis and offensiveness.”
  • Q is said to be associated with R and is written Q ⁇ Rminsup, minconf. Minsup, miconf ⁇ [0, 1].
  • minsup Minimum Support
  • minconf Minimum Confidence
  • the query graph generation unit 33 generates a related term graph composed of term nodes, wherein the term node corresponds to one of the concept nodes of the conceptual grid.
  • the concept node corresponding to the term node includes the query term.
  • the query term is a term that is queried by a user (or a medical person), and a related term graph is separately generated for each query term.
  • FIG. 6 is a related term graph generated by using the conceptual grid graph of FIG. 5.
  • FIG. 6A is a related term graph for the query term t-1
  • FIG. 6B is a related term graph for the query term t-2.
  • the related term graph is a graph expressed by extracting the implications and relations for the query terms using the concept grid graph and visualizing the weighted non-cyclic directional graph (or tree structure graph). That is, for a given query Q, it is a graph that extracts and displays all the information related to 'R' that satisfies the appropriate implication relationship Q ⁇ R and the association Q ⁇ Rminsup, minconf from the concept grid graph.
  • the root node (hereinafter referred to as query node) includes the query term on the top and corresponds to the top level conceptual node containing the query term.
  • the query node is obtained by definition 7.
  • node of the related term graph is expressed in the form of "(Q, (X, Y))” like the query node of definition 7.
  • node when the term node is (Q, (X, Y)), it is expressed as "the term node corresponds to the concept node (X, Y).”
  • terminal is Q. Or “the term node includes Q (as a term).”
  • term node (Q, (X, Y)) will also be referred to simply as "term node Q”.
  • concept c3 can be referred to as a concept suitable for query ⁇ t2 ⁇ , and ( ⁇ t2 ⁇ , ( ⁇ A, B, C, D ⁇ , (t-1, t-2)) is a query for query ⁇ t2 ⁇ . It can be called a node.
  • node ⁇ t-1 ⁇ is a query node
  • node ⁇ t-2 ⁇ is a query node and corresponds to a concept node ( ⁇ A, B, C, D ⁇ , ⁇ t-1, t-2 ⁇ ).
  • the query node will be displayed in a box.
  • the related term graph is a lower term of the term node (hereinafter referred to as a node) including the other term.
  • the implication node corresponds to a concept node corresponding to the query node.
  • the implication node is obtained by the following definition 8.
  • K (G, M, I), for a query node (Q, (X, Y)) for a given query Q ⁇ M, Q ⁇ Y ⁇ If ⁇ minsup ⁇ [0, 1], (Y ⁇ Q, (X, Y)) is defined as the implication node for the query Q.
  • the reason for not reflecting minconf is as follows.
  • node ⁇ t-1 ⁇ is an implication node, and the term node corresponds to the concept node ( ⁇ A, B, C, D ⁇ , ⁇ t1, t2 ⁇ ).
  • a term node (hereinafter referred to as an associated node) including a query term but corresponding to a concept node other than the top level includes a term excluding terms included in all of its parent nodes in the term set of the concept node corresponding to the term.
  • the parent node relationship between the association node or the query node is the same as the parent node relationship between the corresponding concept node.
  • the related term graph is composed of term nodes corresponding only to conceptual nodes in which the set of subjects has at least one element.
  • the associated node is obtained by definition 9.
  • X1 the extent of the concept (X1, Y1) in the relation (X1, Y1) ⁇ (X, Y), uses Y ⁇ Q by definition 3.
  • Y1 ⁇ Y is a set of subjects using additional terminology.
  • definition 3 parent-child relationship
  • definition 4 close-neighbor and near-neighbor
  • Y * is a term (or label) of a term node
  • (X, Y) is a concept or concept node corresponding to a term node
  • a set of all nodes is denoted as N (KQ) minsup, minconf.
  • the graph constructed from these nodes is called a related term graph.
  • the term node is expressed as a label (or term of the term node) in one of the manners.
  • 6A and 6B are graphs of related terms for query terms ⁇ t1 ⁇ and ⁇ t2 ⁇ , respectively, and the query node is represented by a square node, the implication node is a rhombic node, and the association node is a circular node.
  • the query term display unit 34 receives the query term and displays the relation between the related terminology for the query term as a related term graph of the query term, and displays the degree of relationship with the upper node of each term node. Mark with confidence.
  • the terminology and weights included in each term node are displayed.
  • the weights indicate support or confidence.
  • the ratio of the number of the subjects (the number of elements of the subject set) to the term node corresponding to the total number of subjects is displayed as a support.
  • the ratio of the number of users of the conceptual node corresponding to the query node to the number of users of the conceptual node corresponding to the query node is displayed with confidence.
  • 6A or 6B show support as weights in the related term graph (confidence is omitted). That is, the support (or support at node t-2) of the relationship from term node ⁇ t-1 ⁇ to term node ⁇ t-2 ⁇ is
  • 80% And the support of the association from the term node ⁇ t-1 ⁇ to the term node ⁇ t-3 ⁇ (or support at node t-3) is
  • 20%.
  • the query term display unit 34 displays the terminology and support or confidence level corresponding to each term node step by step according to the depth of the related term graph. For example, as shown in FIG. 7, the related term graph of FIG. 6A is ranked in the depth order to sequentially display the term and support or confidence level.
  • the term n1 ⁇ P is Q, and for k satisfying 1 ⁇ k ⁇ m ⁇ 1, r for all nk and nk + 1 ⁇ P,
  • the ordinal set P is defined as the analysis path of the related term graph, and is interpreted as 'term of n1 ⁇ term of n2 ⁇ ... ⁇ nm'.
  • the weight of the analysis path P is defined as the weight of r (nm-1, nm), R.
  • n1 ( ⁇ t-1 ⁇ , ( ⁇ A, B, C, D, E ⁇ , ⁇ t-1 ⁇ ))
  • n2 ( ⁇ t-2 ⁇ , ( ⁇ A, B, C , D ⁇ , ⁇ t-1, t-2 ⁇ )
  • n3 ( ⁇ t-4 ⁇ , ( ⁇ A, B, C ⁇ , ⁇ t-1, t-2, t-4 ⁇ )) .
  • weight information which is numerical information
  • '25% or more (exactly 60%) of hospitals using appendicitis tend to use advantage and appendicitis.
  • FIG. 8 is a table illustrating a format context for a term used in relation to each other in a subject according to a second embodiment
  • FIG. 9 is a table illustrating an example of a term graph and a degree of support for the term.
  • the terminology collecting unit 31 receives a grade weight of each subject as well as a set of terms used in relation to each other in the subject.
  • the grade weight is given according to the influence of the user in the medical field.
  • Terminology used in three-dimensional hospitals has a greater impact on the medical field than terminology used in primary hospitals.
  • the same primary hospital has a greater impact than the original hospital.
  • the influence is different according to the user of the term in the same hospital. For example, it also depends on doctors, nurses, and general nurses. Therefore, weights are determined differently by classifying each user or hospital according to the influence as described above.
  • the grade weight is determined by various variables, such as the number of beds, the order of hospitals (1st, 2nd, 3rd), the job of the mapping agent, the specialty.
  • the rating weight is reflected in the support or confidence of the term node in the related term graph.
  • the weighting method is used as a rating weight instead of the number of users when obtaining the support or confidence. There may be ways to obtain it.
  • FIG. 8 a set of terminology of a hospital E having a high grade weight is inserted in a formal context. In FIG. 8, an additional five sets of terms are inserted.
  • FIG. 9A is a related term graph for the query ⁇ t-1 ⁇ generated based on the format context of FIG. 8.
  • the support of the term node ⁇ t-2 ⁇ was reduced from 80% to 40%, and the support of the term node ⁇ t-3 ⁇ was increased from 20% to 60%. That is, only hospital E marks appendicitis as scabtail, but since hospital E has a very large weight weight, it is indicated that the scabtail is more relevant than Appenticitis.
  • the priority of the term relevance shown to the queryer is also indicated differently.
  • the spear tail salt t-3 is displayed more preferentially than the advantaget t-2.
  • the class weight is an integer multiple. That is, the terminology set for the terminology is duplicated by an integer multiple.
  • definition 9-2 The method of obtaining support or confidence obtained by class weight instead of the number of users is obtained by the following definition 9-2.
  • definition 9-2 the definition of support and reliability in definition 9 described above is based on the class weight, not the number of subjects.
  • the support of the concept is calculated as the ratio of the sum of the weights of the users to the sum of the weights of the users.
  • Confidence of the concept is calculated as the ratio of the sum of the weights of the users to the users of the associated concepts.
  • minsup minconf ⁇ [0, 1]
  • W h are the class weights of principal h.
  • the support of the term node is calculated as the ratio of the sum of the class weights of the users of the conceptual node corresponding to the sum of the weights of the entire users.
  • the term node's confidence level is calculated as the ratio of the sum of the weights of the users of the conceptual node corresponding to the query node to the sum of the weights of the users of the conceptual node corresponding to the query node.
  • mapping information of a term used by a user may be input as at least two different mapping information. That is, the terminology of ⁇ t-1, t-2, t-3, t-4 ⁇ and the terminology of ⁇ t-31, t-32, t-33 ⁇ are not only different from each other. The figures are respectively different from ⁇ A, B, C, D, E ⁇ and ⁇ X, Y, Z ⁇ .
  • the concept grid graph is formed separately for each mapping information of G1 and mapping information of G2.
  • G The number of ⁇ A, B, C, D, E, X, Y, Z ⁇
  • the concept of G1 and the concept of G2 may be different, so the accuracy of analysis is poor.
  • the conceptual grid graph is generated separately, it is preferable to set the total number of subjects differently in the separated conceptual grid graph unit. That is, it is used as the number of
  • the entire subject of use of the support is limited to the subjects included in the separated graph including the concept node corresponding to the term node.
  • a set of terms used in association with each user is input and a grade weight of each user is input (S10), and a set of users using the term set and the term set in common.
  • a concept lattice graph consisting of nodes (hereinafter referred to as concept nodes) is formed (S20).
  • concept nodes the term set of the lower conceptual node inherits all the term sets of the upper conceptual node.
  • the query term is input (S30), and a related term graph including term nodes is generated for the query term (S40).
  • the term node is configured to correspond to one of the concept nodes of the concept grid graph, and the concept node corresponding to the term node composes a related term graph to include the query term.
  • a detailed method of generating the related term graph is as shown in FIG. 12.
  • the related term graph is formed as follows.
  • the root node of the related term graph (hereinafter referred to as query node) corresponds to the highest concept node including the query term and is configured to include the query term as its own term.
  • the term node (hereinafter, referred to as association node) that does not correspond to the highest level concept node includes terms except for the terms included in all of the parent nodes in the term set of the concept node corresponding to the term.
  • the parent node relationship between the association node or the query node is configured as the parent node relationship between the corresponding concept node.
  • the related term graph is a lower term of the term node (hereinafter referred to as a node) including the other term. It is configured as a node, the implication node is configured to correspond to the concept node corresponding to the query node.
  • the support of the term node is obtained by the ratio of the total number of use subjects of the concept node corresponding to the total number of use subjects. In this case, the support reflects the grade weight.
  • the confidence level of the term node is obtained as the ratio of the number of subjects using the conceptual node to the number of subjects corresponding to the query node. In this case, the degree of confidence reflects the grade weight.
  • the grade weights are integer multiples, and the set of terms for each user is further overlapped by the grade weights to obtain the support or confidence.
  • node's support is calculated as a ratio of the sum of the class weights of the corresponding subjects to the sum of the class weights of all the subjects.
  • node's confidence level is calculated as the ratio of the sum of the weights of the users of the conceptual node corresponding to the query node to the sum of the weights of the users of the conceptual node corresponding to the query node.
  • the entire subject of use of the support is limited to the subjects included in the separated graph including the concept node corresponding to the term node.
  • the query term is input, and the relationship between the related terminology for the query term is displayed in the terminology graph of the terminology, and the degree of relationship with the upper node of each term node is displayed with support or confidence (S50).
  • the terminology and support for each term node are displayed step by step according to the depth of the related term graph.
  • the present invention is applied to the development of a query term mapping support system that supports the mapping of query terms to terms used in relation to each other by forming a concept grid for terms used in a plurality of subjects. This is possible.
  • the present invention consists of a graph of a hierarchical structure of terms related to a query term, and a query term mapping that displays the relevance of a higher term node as support or confidence in each term node with the query term as a root node.
  • a query term mapping that displays the relevance of a higher term node as support or confidence in each term node with the query term as a root node.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

다수의 사용주체에 의해 서로 관련 있게 사용되는 용어(사용용어)들을 수집하여, 특정 용어(질의용어)와 관련된 용어들을 관계 그래프로 표시하는 개념격자 기반 질의용어 매핑지원 시스템에 관한 것으로서, 각 사용주체에 의해 사용되는 사용용어를 집합(용어집합)으로 입력받아 용어집합과 사용주체 집합을 노드로 구성하는 개념격자 그래프를 형성하고, 개념격자 그래프를 이용하여 질의용어에 대한 관련용어 그래프를 생성하는 구성을 마련한다. 상기와 같은 시스템 및 방법에 의하여, 사용주체에 의해 서로 관련 있게 사용되는 용어 간의 관계를 자동으로 분석하여 그래프로 표시함으로써 사용용어 사이에 내재되어 있는 관계를 직관적으로 표시할 수 있다.

Description

사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법
본 발명은 다수의 의료기관 등 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 관한 것이다.
또한, 본 발명은 질의용어와 관련된 용어들을 용어노드로 표시되는 계층구조의 그래프로 구성하고, 질의용어를 루트노드로 하여 각 용어노드에 상위 용어노드의 관련성을 지지도 또는 확신도로 표시하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 관한 것이다.
또한, 본 발명은 사용용어 사이의 관련성을 지지도 또는 확신도로 표시하되, 상기 사용용어를 사용하는 사용주체의 등급을 반영하여 지지도 또는 확신도를 계산하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 관한 것이다.
일반적으로 우리나라에서 사용되고 있는 상당수의 의학 용어가 외래어 또는 복잡한 한자어 상태로 사용되고 있다. 또한 의료인들 사이에서도 동일한 개념에 대해 여러 가지 용어를 사용함으로써 원활한 정보 소통에 문제가 발생하고 있다.
이점 때문에 의학이 일반인에게는 이해하기 어렵고 쉽게 접근하기 어려운 학문으로 인식되고 있으며 환자와 의료인 사이의 원활한 정보 전달에도 많은 어려움이 있다. 이러한 문제와 함께 의학 용어가 표준화되어 있지 않아 병원 간의 진료정보 교류가 쉽지 않은 문제가 대두되고 있다.
한편, 국제적으로도 의학 용어(또는 의료용어)에 대한 표준화가 많이 이루어지고 있으나, 의료분야가 가지는 복잡성으로 인해 표준화가 쉽지 않은 실정이다.
동일한 개념(질환)을 진료과 별로 다른 의학용어를 사용하여 표현하기도 하며, 같은 의학용어를 사용하더라도 사용처에 따라 다른 의미를 가지는 경우가 있다. 예를 들어, 우리가 ‘맹장염’이라 흔히 일컫는 질병은 ‘Appendicitis’라는 의학용어가 있음에도 상황에 따라 ‘충수돌기염’, ‘충수염’, ‘막창자꼬리염’, ‘맹장염’으로 일컬어지며, ‘Myocardial Infarction(심근경색)’ 역시 상황에 따라 ‘cardiac infarction’, ‘infarction of heart’, ‘MI’, ‘Heart Attack’ 등으로 일컬어진다. 반면에 ‘주소’의 경우, 원무과에서는 ‘Address’의 의미로 사용하지만, 진료실에서는 ‘Chief Complain(주호소)’의 의미로 사용한다.
상기와 같이 의료용어가 복잡한 것은 의료분야 자체의 복잡성에 기인한 것으로, 이러한 의료용어를 구분하고 의료진 사이의 의사소통이나 서로 다른 분야의 의사소통을 위하여 기준이 되는 참조용어체계(Reference Terminology)와 인터페이스 용어체계(Interface Terminology)가 필요하다.
더 나아가서, 각 임상 영역에서 필요로 하는 정보를 적시에 효율적으로 수집, 분석하기 위해서는 해당 임상영역에서 필요로 하는 지식을 정의하고, 도메인 특화된 규칙을 통하여 필요 지식들을 수집하는 용어 모형을 통하여 의미적 상호운용성을 확보하는 것이 필수적이다.
의료분야에서 사용되는 용어체계(Terminology System)란 특정 분야에서 사용되는 의학 개념의 집합이며, 개념에 대한 정의, 개념을 일컫는 용어, 개념 간의 관계 등을 포함하는 시스템이다. 의학용어 체계는 그 목적성과 특징에 따라, 참조용어(Reference terminology), 인터페이스용어(Interface terminology), 분류체계(Classification) 등으로 구분할 수 있으며, 임상현장에서 사용하는 의학용어를 로컬용어(Local Term)라 한다.
참조용어 체계(Reference terminology)는 의학개념과 그들 간의 관계(Inter-relationships)의 정형화된 표현의 집합을 말한다. 참조용어체계에는, SNOMED CT, RxNorm, NCI Thesaurus, ICNP 등이 있다.
인터페이스 용어 체계(Interface terminology)의 경우 의료전문가 및 임상지원 전문가가 임상문서 내 정보를 보다 수월하게 획득 및 수집하여 전달, 처리할 수 있도록 제안된 용어체계로, 의료진이 진료 기록 작성 시점에서 사용하는 모든 용어와 진료서식지 상에 있는 모든 용어들이 이에 해당한다. 인터페이스 용어체계는 특정 병원의 의료진이 특정 의학개념을 일컬을 때 사용하는 여러 형태의 용어들과 지식의 집합이라 할 수 있으며, 이는 의료진과 참조용어체계를 인터페이스 해주는 역할을 한다.
분류체계(Classification)는 특정 목적에 맞게 미리 정해진 분화 수준에서 자료를 통합하기 위해 모든 것을 망라하는 상호 배반적인 범주, 분류의 목적은 의료서비스의 통계나 연구를 수행하는데 사용하며, ICD(Internatioanl Classification of Disease), KCD(Korean Standard Classification of Disease) 등이 있다.
의학개념에 대한 정형화된 표현들의 집합인 참조용어체계가 있다고 하더라도, 각각 병원에서 다양한 분야의 의료진이 사용하는 로컬용어는 상이할 수밖에 없기 때문에 그 차이를 매워 주기 위해 매핑이라는 작업이 필요하다.
그러나 매핑 작업에서 가장 큰 걸림돌은 의료 현장에서 사용되는 로컬용어의 다양성이다. 즉, 한 병원 내에서도 진료과별로 의사별로 사용되는 로컬용어가 다르고 각 병원에서 관리되고 있는 로컬 용어들도 병원마다 다르다. 로컬용어는 직접 의사 등 의료진이 사용하고 있는 용어이고, 인터페이스 용어 체계는 로컬용어들을 모아서 의미를 부여하고 개념별로 정의하여 의학개념에 대한 상황정보를 탑재할 수 있도록 하는 용어체계 시스템으로서, 모두 의료현장에서 사용되는 용어이기 때문에 그 영향력이 매우 크다.
따라서 참조용어 체계가 의료개념을 아무리 잘 표현할 수 있도록 구성되었더라도 현장에서 사용되는 의학용어와의 차이는 현저하다. 한편, 참조용어 체계를 이용하고자 하는 궁극적 목표 중 하나는 의료현장에서 난립하는 용어들 간의 의미전달을 위한 것이다.
의료현장에서 사용되는 용어를 포괄할 수 있는 의료용어체계를 만드는 가장 좋은 방법 중 하나는 의료현장에서 사용되는 가장 지배적인 용어로 점차 통일화시켜 나가는 것이고, 이를 위해서 현재 의료현장에서 사용되는 용어들의 사용 현황을 알아야 하는 것이 매우 중요하다.
의료현장에서 가장 많이 사용되는 용어들이 무엇인지 알게 되면, 각 의료진 또는 사용주체들은 지배적인 용어들을 점차 이용하려고 할 것이고 쏠림현상에 의해 하나 또는 소수의 최적 용어로 수렴될 수 있을 것이다.
그러나 단순히 얼마나 많은 의료진이 사용한다는 것보다 더 중요한 것은 어떤 의미로 사용되고 있는지를 파악하는 것이다. 의료용어는 동일한 용어라도 다르게 사용될 수 있기 때문이다.
의료현장에서 사용되는 각 용어들의 의미를 일일이 파악하는 것은 거의 불가능하므로, 의료현장 용어들의 관계를 정량적으로 분석하여 의료용어 간에 내재되어 있는 관계를 직관적으로 추출할 수 있는 자동화된 분석기술이 절실하다.
본 발명의 목적은 상술한 바와 같은 문제점을 해결하기 위한 것으로, 서로 관련 있게 사용되는 의료용어 등 사용용어들로부터 용어들 간의 관계를 분석함으로써 현장에서 사용되는 사용용어 간에 내재되어 있는 관계를 직관적으로 파악할 수 있는 개념격자 기반 질의용어 매핑지원 시스템 및 방법을 제공하는 것이다.
또, 본 발명의 목적은 다수의 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 주어진 질의용어에 대하여 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법을 제공하는 것이다.
또, 본 발명의 목적은 질의용어와 관련된 용어들을 용어노드로 표시되는 계층구조의 그래프로 구성하고, 질의용어를 루트노드로 하여 각 용어노드에 상위 용어노드의 관련성을 지지도 또는 확신도로 표시하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법을 제공하는 것이다.
또, 본 발명의 목적은 사용용어 사이의 관련성을 지지도 또는 확신도로 표시하되, 상기 사용용어를 사용하는 사용주체의 등급을 반영하여 지지도 또는 확신도를 계산하는 개념격자 기반 질의용어 매핑지원 시스템 및 방법을 제공하는 것이다.
상기 목적을 달성하기 위해 본 발명은 다수의 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 시스템에 관한 것으로서, 각 사용주체에 의해 서로 관련 있게 사용되는 용어를 집합(이하 용어집합)으로 입력받는 사용용어 수집부; 용어집합과 상기 용어집합을 서로 관련 있게 사용하는 사용주체 집합을 노드(이하 개념노드)로 구성하는 개념격자 그래프를 형성하되, 하위 개념노드의 용어 집합은 상위 개념노드의 용어집합을 모두 상속하는 개념격자 형성부; 및 질의용어를 입력받아 상기 질의용어에 대하여 용어노드로 구성하는 관련용어 그래프를 생성하되, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응되고, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함하는 질의그래프 생성부를 포함하는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 관련용어 그래프의 루트노드(이하 질의노드)는 상기 질의용어를 포함하는 최상위 개념노드와 대응되고, 자신의 용어로서 상기 질의용어를 포함하고, 상기 최상위 개념노드와 대응되지 않는 용어노드(이하 연관노드)는 자신에 대응되는 개념노드의 용어집합에서 자신의 모든 상위노드에 포함되는 용어들을 제외한 용어를 포함하고, 상기 연관노드 또는 상기 질의노드 사이의 상하위 노드 관계는 대응되는 개념노드 간의 상하위 노드 관계와 같은 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 용어노드는 자신과 대응되는 개념노드의 의료기관 집합개수를 이용하여 계산된 지지도 또는 확신도를 가지는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 사용용어 수집부는 각 사용주체의 등급 가중치를 입력받고, 상기 지지도 또는 확신도는 상기 등급 가중치를 반영하는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 질의노드에 대응되는 개념노드의 용어집합이 상기 질의용어 이외 다른 사용용어를 포함하면, 상기 관련용어 그래프는 상기 다른 사용용어를 포함하는 용어노드(이하 함의노드)를 상기 질의노드의 하위노드로서 구성하되, 상기 함의노드는 상기 질의노드에 대응되는 개념노드와 대응되는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 등급가중치는 정수배이고, 각 사용주체에 대한 사용용어 집합을 상기 등급가중치 만큼 더 중복하여, 상기 지지도 또는 확신도가 구해지는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 용어노드의 지지도 또는 확신도는 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산되는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체으로 한정되는 것을 특징으로 한다.
또, 본 발명은 상기 시스템은, 질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도로 표시해주는 질의용어 표시부를 더 포함하는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 질의용어 표시부는 상기 관련용어 그래프의 깊이에 따라 단계적으로 각 용어노드에 해당하는 사용용어 및 지지도 또는 확신도를 표시해주는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 시스템에 있어서, 상기 질의용어 표시부는 상기 관련용어 그래프에서 지지도 또는 확신도가 최소 지지도 또는 최소 확신도 이상인 용어노드만을 표시해주는 것을 특징으로 한다.
또한, 본 발명은 다수의 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 방법에 관한 것으로서, (a) 각 사용주체에서 서로 관련 있게 사용되는 사용용어를 집합(이하 용어집합)으로 입력받는 단계; (b) 용어집합과 상기 용어집합을 공통적으로 사용하는 사용주체 집합을 노드(이하 개념노드)로 구성하는 개념격자 그래프를 형성하되, 하위 개념노드의 용어 집합은 상위 개념노드의 용어 집합을 모두 상속하는 단계; 및 (c) 질의용어를 입력받아 상기 질의용어에 대하여 용어노드로 구성하는 관련용어 그래프를 생성하되, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응되고, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함하는 단계를 포함하는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 용어노드는 자신과 대응되는 개념노드의 의료기관 집합개수를 이용하여 계산된 지지도 또는 확신도를 가지는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 (a)단계에서, 각 사용주체의 등급 가중치를 입력받고, 상기 지지도 또는 확신도는 상기 등급 가중치를 반영하는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 등급가중치는 정수배이고, 각 사용주체에 대한 사용용어 집합을 상기 등급가중치 만큼 더 중복하여, 상기 지지도 또는 확신도가 구해지는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 용어노드의 지지도 또는 확신도는 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산되는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체로 한정되는 것을 특징으로 한다.
또, 본 발명은 개념격자 기반 질의용어 매핑지원 방법에 있어서, 상기 방법은, (d) 질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도 또는 확신도로 표시해주는 질의용어 단계를 더 포함하는 것을 특징으로 한다.
또한, 본 발명은 상기 개념격자 기반 질의용어 매핑지원 방법을 수행하는 프로그램을 기록한 컴퓨터로 읽을 수 있는 기록매체에 관한 것이다.
상술한 바와 같이, 본 발명에 따른 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 의하면, 의료기관 등 사용주체에서 서로 관련 있게 사용되는 의료용어 등 사용용어를 취합하여 사용용어 간의 사용관계를 자동으로 분석함으로써 사용용어 간에 내재되어 있는 관계를 직관적으로 추출할 수 있는 효과가 얻어진다.
한편, 현재 의료현장에서 사용되는 용어들의 사용 현황, 즉, 사용분포를 내포하여 관련 용어들의 관계를 표시해주거나 추천해주면, 의료현장에서 각 의료진 또는 사용주체들은 지배적인 용어를 선택하여 사용할 뿐만 아니라, 자신들이 판단하기에 의료개념을 가장 잘 표현하는 용어를 선택하게 될 것이다. 또한, 이들의 선택은 현장에서 사용되는 용어들의 관계에 영향(피드백)을 줄 것이다.
따라서 시간이 지남에 따라 용어간 적자생존에 의해 하나 또는 소수의 최적 용어로 쏠림현상이 일어날 것이고, 결과적으로, 다수의 의료진들이 서로 협력 혹은 경쟁을 통하여 얻게 되는 지적 능력에 의한 결과로 얻어진 집단적 능력인 집단지성(集團知性, Collective Intelligence)에 의해 최적 용어가 결정될 것이다. 이러한 과정은 향후 사용용어 표준화에도 크게 기여할 수 있다.
또한, 본 발명에 따른 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 의하면, 사용용어 사이의 관련성을 지지도 또는 확신도로 표시하되 상기 사용용어를 사용하는 사용주체의 등급을 반영하여 지지도 또는 확신도를 계산함으로써, 현장에서 각 사용주체가 미치는 영향력을 반영하여 보다 정확한 사용용어 간의 관계를 도출할 수 있는 효과가 얻어진다.
또한, 본 발명에 따른 개념격자 기반 질의용어 매핑지원 시스템 및 방법에 의하면, 각 사용주체에서 사용하고 있는 사용용어로 개념격자 그래프를 형성하고, 이를 기반으로 하여, 각 질의용어에 대한 관련용어 관계를 보다 빠르고 정확하게 도출할 수 있는 효과가 얻어진다. 즉, 사용용어의 개수와 사용주체의 수가 매우 많으므로 관련용어를 도출하는 과정이 매우 복잡하고 많은 시간이 소요될 수 있으나, 개념격자의 그래프 특성을 이용함으로써 계산량을 상당히 줄일 수 있다.
도 1은 본 발명을 실시하기 위한 전체 시스템의 구성을 도시한 도면이다.
도 2는 본 발명의 일실시예에 따른 개념격자 기반 질의용어 매핑지원 시스템의 구성에 대한 블록도이다.
도 3은 본 발명의 제1 실시예에 따른 사용주체에서 서로 관련 있게 사용되는 사용용어에 대한 형식문맥을 예시한 도면이다.
도 4는 본 발명의 제1 실시예에 따라 도 3의 형식문맥으로부터 추출한 개념을 예시한 도면이다.
도 5는 본 발명의 제1 실시예에 따른 개념격자 그래프를 예시도이다.
도 6은 본 발명의 제1 실시예에 따른 관련용어 그래프를 예시도이다.
도 7은 본 발명의 제1 실시예에 따라 생성된 관련용어 그래프를 이용하여 질의용어에 대한 관련용어의 지지도를 예시한 표이다.
도 8은 본 발명의 제2 실시예에 따른 사용주체에서 서로 관련 있게 사용되는 사용용어에 대한 형식문맥을 예시한 표이다.
도 9는 본 발명의 제2 실시예에 따른 관련용어 그래프의 예시도 및 관련용어의 지지도에 대한 표이다.
도 10은 본 발명의 제3 실시예에 따라 적어도 2개의 다른 매핑정보를 가지는 형식문맥을 예시한 도면이다.
도 11은 본 발명의 일실시예에 따른 개념격자 기반 질의용어 매핑지원 방법을 설명하는 흐름도이다.
도 12는 본 발명의 일실시예에 따라 관련용어 그래프를 생성하는 알고리즘을 예시한 도면이다.
이하, 본 발명의 실시를 위한 구체적인 내용을 도면에 따라서 설명한다.
또한, 본 발명을 설명하는데 있어서 동일 부분은 동일 부호를 붙이고, 그 반복 설명은 생략한다.
먼저, 본 발명을 실시하기 위한 전체 시스템의 구성의 예들에 대하여 도 1을 참조하여 설명한다.
도 1a 또는 도 1b에서 보는 바와 같이, 본 발명에 따른 질의용어 매핑지원 시스템은 네트워크 상의 서버 시스템 또는 컴퓨터 단말 상의 프로그램 시스템으로 실시될 수 있다.
도 1a와 같이, 본 발명의 실시를 위한 전체 시스템의 일례는 사용주체 단말(10)과 질의용어 매핑지원 시스템(30)으로 구성되고 서로 네트워크(20)로 연결된다. 또, 필요한 데이터를 저장하기 위한 데이터베이스(40)를 더 구비할 수 있다.
사용주체 단말(10)은 의료인 또는 의료기관 등 사용주체가 이용하는 PC, 노트북, 넷북, PDA, 모바일 등의 통상의 컴퓨팅 단말기이다. 의료인 등은 사용주체 단말(10)을 이용하여 자신의 의료기관에서 서로 관련 있게 사용되는 다수의 사용용어들을 매핑지원 시스템(30)으로 전송하거나, 특정 용어(또는 질의용어)에 대한 관련용어들을 참조해줄 것을 매핑지원 시스템(30)에 요청한다.
질의용어 매핑지원 시스템(30)은 통상의 서버로서 네트워크(20)에 연결되어 질의된 용어에 대하여 관련 용어들을 검색하여 관련성을 보여주는 서비스를 제공하는 서버이다. 또, 매핑지원 시스템(30)은 각 사용주체들이 서로 관련 있게 사용되는 다수의 사용용어들을 등록할 수 있는 등록서비스를 제공한다. 매핑지원 시스템(30)은 상기 각 서비스들을 인터넷 상의 웹페이지로 제공하는 웹서버 또는 웹어플리케이션 서버 등으로 구현될 수 있다.
매핑지원 시스템(30)은 질의용어에 대하여 이와 관련된 사용용어들을 제시하면, 사용자(또는 사용주체)들은 이 제시된 결과를 참조하여 질의용어에 대한 사용용어를 매핑한다. 그러나 매핑지원 시스템(30)에서 결과로 제공되는 용어의 관련성은 용어매핑에 한정되어 이용되는 것은 아니다. 예를 들어, 사용용어를 분석하거나 하는 등의 작업에도 이용될 수 있다. 따라서 매핑을 지원하는 것에 한정되지는 않지만, 이하에서 설명의 편의를 위해 매핑을 지원하는 것으로 기재한다.
데이터베이스(40)는 매핑지원 시스템(30)에서 필요한 데이터를 저장하는 통상의 저장매체로서, 각 사용주체으로부터 수신한 사용용어로부터 형성한 개념격자 그래프, 또는 각 사용용어별로 이와 관련된 다른 관련 사용용어들을 계층구조의 관련용어 그래프 등을 저장한다.
도 1b와 같이, 본 발명의 실시를 위한 전체 시스템의 다른 예는 컴퓨터 단말(13)에 설치되는 프로그램 형태의 질의용어 매핑지원 시스템(30)으로 구성된다. 즉, 매핑지원 시스템(30)의 각 기능들은 컴퓨터 프로그램으로 구현되어 컴퓨터 단말(13)에 설치되어, 사용자(14)에 의해 각 사용주체들이 서로 관련 있게 사용되는 다수의 사용용어들을 컴퓨터 단말(13)의 입력장치를 통해 입력받아 등록하거나, 질의용어에 대한 관련용어들을 참조해줄 것을 요청받으면 그 결과를 컴퓨터 단말(13)의 출력장치를 통해 출력한다. 한편, 매핑지원 시스템(30)에서 필요한 데이터들은 컴퓨터 단말(13)의 하드디스크 등 저장공간에 저장되어 이용된다.
다음으로, 본 발명의 제1 실시예에 따른 개념격자 기반 질의용어 매핑지원 시스템의 구성을 도 2를 참조하여 설명한다.
도 2에서 보는 바와 같이, 본 발명의 제1 실시예에 따른 개념격자 기반 질의용어 매핑지원 시스템(30)은 사용용어 수집부(31), 개념격자 형성부(32), 및 질의그래프 생성부(33)를 포함하여 구성된다. 추가적으로, 질의용어 표시부(32)를 더 포함하여 구성될 수도 있다. 또한, 데이터를 저장하기 위한 데이터베이스(또는 저장공간)(40)를 더 포함하여 구성될 수 있다.
먼저, 사용용어 수집부(31)는 각 사용주체으로부터 서로 관련 있게 사용되는 사용용어를 집합으로 입력받는다. 동일한 의미로 사용되는 사용용어들도 하나의 서로 관련 있게 사용하는 사용용어들이다.
도 3a에서 보는 바와 같이, 병원A, B, C는 맹장염의 의미로, 맹장염(t-1), 어펜디시티스(Appendicitis)(t-2), 및 충수돌기염(t-4)으로 3가지를 사용하고 있고, 병원D는 맹장염(t-1) 및 어펜디시티스(t-2)의 2가지만 사용하고, 병원E는 맹장염(t-1) 및 막창꼬리염(t-3)의 2가지만 사용하고 있다. 즉, 사용용어 수집부(31)는 각 사용주체으로부터 도3a와 같은 사용용어 집합을 입력받는다.
다음으로, 개념격자 형성부(32)는 용어집합과 상기 용어집합을 공통적으로 사용하는 사용주체 집합을 노드(이하 개념노드)로 표시하고, 하위 개념노드의 용어 집합은 상위 개념노드의 용어 집합을 모두 상속하는 개념격자 그래프를 형성한다.
이하에서, 개념격자 그래프를 형성하는 방법과 그 특성들을 보다 구체적으로 설명한다.
각 사용주체에 의해 서로 관련 있게 사용되는 사용용어들은 다음 정의 1과 같이 형식문맥(Formal context)으로 정의할 수 있다.
[정의 1]
형식문맥 K=(G, M, I)는 사용주체 집합 G와 용어집합 M, 그리고 G와 M사이의 이항관계 I⊆G×M로 구성된다. 즉, G와 M의 원소들은 각각 해당 형식문맥의 사용주체들과 각 사용주체들이 사용할 수 있는 사용용어들을 나타낸다. 또한, 어떤 사용주체 g가 사용용어 m을 사용하는 경우, gIm 또는 (g, m) ∈ I 로 나타내며, g는 m을 사용한다는 것을 의미한다.
도 3b에서 보는 바와 같이, 형식문맥(Formal context)은 데이터 테이블 형태로 나타낼 수 있으며, 해당 표의 행과 열의 헤드부분은 각각 형식문맥를 구성하는 사용주체들과 사용용어들로 구성된다. 또한, 데이터 테이블의 각 셀에 대해서는, 해당 셀에 관련된 사용주체과 사용용어가 이항관계 I를 만족할 경우에는 X표시하고, 이외의 경우에는 빈 공란으로 남겨둔다.
이와 같은 형식문맥으로부터 서로 관련 있게 사용되는 사용용어를 갖는 사용주체들을 클러스터링(Clustering)하여, 정보의 기본단위로서 연관개념들(Concepts)을 추출할 수 있다. 각 개념들은 (O, T)와 같은 형태의 쌍(pair)으로 정의되며, 보다 정형적인 정의는 정의 2와 같다.
[정의 2]
임의의 형식문맥 K=(G, M, I)에 대하여, O⊆G, T⊆M일 때, intent(O)=T ∧ extent(T)=O를 만족하는 (O, T)를 개념(formal concept)이라고 한다.
단, intent(O):={t∈M|∀o∈O:(o, t)∈I}=O',extent(T):={o∈G|∀t∈T:(o, t)∈I}=T'.
임의의 O⊆G에 대하여, intent(O)에 의해 O의 모든 사용주체들이 공통적으로 갖는 용어집합을 구할 수 있다. 예를 들면, 위의 도 3b의 형식문맥에 있어서, O={D, E}에 대하여, intent(O)={t-1, t-2, t3}이다. 한편, 임의의 T⊆M에 대하여, extent(T)에 의해 T의 사용용어들을 사용하는 사용주체들의 집합을 구할 수 있다. 예를 들면, T={t-1, t-2, t-4}에 대하여, extent(T)={A, B, C}이다.
즉, 각 개념들은 (O, T)와 같은 형태의 쌍(pair)으로 정의되며 특히, 사용주체 집합 O는 용어집합 T의 extent이며, 동시에, 용어집합 T는 사용주체 집합 O의 intent가 된다. 이와 같은 방법으로 형식문맥 K=(G, M, I)로부터 추출한 모든 개념들의 집합을 B(K):={(O, T)∈2G×2M|intent(O)=T∧ extent(T)=O}로 나타낸다. 도 3b의 형식문맥로부터 추출한 모든 개념들 B(K)는 도4와 같다. 이하에서, c1 내지 c5는 도 3b에 표기된 개념을 나타낸다.
위와 같이 주어진 형식문맥로부터 추출된 개념들은 extent 또는 intent를 토대로 상위-하위개념관계를 정의 3과 4 같이 정의할 수 있다.
[정의 3]
주어진 형식문맥의 임의의 개념 (O1, T1), (O2, T2)∈B(K)에 대하여, 상하위관계(Superconcept-Subconcept relation) (O1, T1)≤(O2, T2)는 일종의 반순서관계(partial order relation)로서 다음과 같이 정의된다.
(O1, T1)≤(O2, T2) ⇔ O1⊆O2(⇔ T1⊇T2).
[정의 4]
형식문맥 K=(G, M, I)에 존재하는 모든 개념 집합을 B(K)라 하고, 개념 (X1,Y1),(X2,Y2)∈B(K)에 대하여, (X1,Y1)<(X2,Y2)를 만족하고, (X1,Y1)<(X3,Y3)<(X2,Y2)를 만족하는 (X3,Y3)가 B(K)에 존재하지 않을 때, (X1,Y1)을 (X2,Y2)의 근접 하위 이웃(lower neighbor)이라 하고, (X2,Y2)를 (X1,Y1)의 근접 상위 이웃(upper neighbor)이라 하며 (X1,Y1)
Figure PCTKR2011002264-appb-I000001
(X2,Y2)라 표기한다.
형식문맥 K=(G, M, I)로부터 만들어진 모든 개념들 간의 상위-하위개념관계≤는 일종의 반순서관계(partial order relation)에 해당하며, 개념들 사이의 상하위관계에 의해 만들어진 계층적 개념구조를 개념격자(Concept Lattice 또는 Galois Lattice)라고 부르고 L:=(B(K), E≤)과 같이 표현한다.
도 3b에 기술된 형식문맥로부터 추출된 개념과 개념간의 상하위관계 정보를 담고 있는 개념격자를 그래프를 사용하여 도 5와 같이 표시할 수 있다. 그래프의 각 노드는 개념노드를 표시하고, 개념노드의 상위에 위치한 레이블은 개념노드의 사용용어(또는 intent)를, 하위에 위치한 레이블은 개념노드의 사용주체(또는 extent)을 각각 나타낸다. 즉, 개념격자 그래프는 용어집합과 상기 용어집합을 공통적으로 사용하는 사용주체 집합을 노드(이하 개념노드)로 표시한다.
상기 정의 3과 4에서 보는 바와 같이, 하위 개념노드는 자신의 상위 개념노드의 사용용어들을 모두 포함한다. 즉, 하위 개념노드의 용어 집합은 상위 개념노드의 용어 집합을 모두 상속한다. 반대로 상위 개념노드는 자신의 하위 개념노드의 사용주체를 모두 포함한다.
다음으로, 형식문맥(Formal Context)에서의 용어집합(또는 속성집합) 간의 관계 및 관계에 대한 특성을 설명한다.
먼저, 용어집합 간의 함의관계는 정의 5와 같이 정의된다.
[정의 5]
주어진 형식문맥 K=(G, M, I)의 임의의 두 사용용어 Q, R⊆M 이, extent(Q)⊆extent(R)를 만족하는 경우, "Q는 R을 함의한다."라고 부르며, Q⇒R로 표기한다.
임의의 Q, R⊆M에 대하여, extent(Q)와 extent(R)에 의해서 사용용어 Q와 R을 사용하는 병원들을 각각 구할 수 있다. 예를 들어, 도 3b의 형식문맥에 있어서, Q={t-4}, R={t-1, t-2}이라 하면 extent(Q)={A, B, C}, extent(R)={A, B, C, D}를 각각 구할 수 있으며, 이는 extent(Q)⊆extent(R)를 만족하므로, Q⇒R, 즉 {t-4}⇒{t-1, t-2}가 성립함을 알 수 있다. 다시 말해, 도 3b에 주어진 형식문맥의 데이터로부터 "충수돌기염을 사용하는 병원은 맹장염과 어펜티시티스도 역시 사용한다."라는 추가적인 정보를 추출해 낼 수 있다.
형식문맥에서의 용어집합 간의 연관관계를 지지도 및 확신도에 의해 정의 6과 같이 정의된다.
[정의 6]
주어진 형식문맥 K=(G, M, I)의 임의의 두 사용용어 Q, R⊆M 이,
Figure PCTKR2011002264-appb-I000002
≥minsup 와
Figure PCTKR2011002264-appb-I000003
≥minconf 를 만족하는 경우
Q는 R과 연관된다 라고 하며, Q→Rminsup,minconf 로 표기한다. 단 minsup, miconf∈[0, 1].
정의6에서 언급된
Figure PCTKR2011002264-appb-I000004
Figure PCTKR2011002264-appb-I000005
는, 각각 연관관계 Q→R의 지지도(support)와 확신도(confidence)라고 부르며, 특히, minsup(Minimum Support: 최소 지지도)와 minconf(Minimum Confidence: 최소 확신도)는 집합 Q가 R과 갖는 연관관계를 바라보는 분석가의 주관적인 경계값(Threshold)이다. 먼저 minsup는 Q와 R 사이에 존재하는 관계가 전체 사용주체들 중 적어도 얼마나 되는 사용주체 사이에서 성립할 경우에, 이 두 집합 간의 연관관계를 인정할 것인가에 대한 경계 값이다. 그리고 minconf는 Q라는 사용용어를 사용하는 병원 중 적어도 얼마나 되는 병원이 R이라는 사용용어 역시 가질 때, Q는 R과 연관관계에 있다고 판단 할 것인지에 대한 경계 값이다.
예를 들어, 표 2의 context에 있어서, Q={t-2}, R={t-4}이라 하고, minsup=0.5, minconf=0.7이라 하면, |extent({t-2, t-4})|/|G|=0.6≥minsup, |extent({t-2, t-4})|/|extent({t-2})|=0.75≥minconf 이므로, Q→R 즉, 50%이상의 지지도와 70%이상의 확신도에서 {t-4}→{t-2}는 성립한다고 할 수 있다. 이를 문맥상 '어펜티시티스를 사용하는 병원 중 70% 이상이 충수돌기염을 사용하는 경향을 보이며, 이러한 현상은 전체 병원 중 50% 이상의 병원으로부터 관찰된다' 라고 해석할 수 있다. 특히, minsup=0, minconf=1인 경우의 연관관계는 곧 함의관계를 나타내므로, 함의관계는 연관관계의 특수한 경우라 할 수 있다.
다음으로, 질의그래프 생성부(33)는 용어노드로 구성하는 관련용어 그래프를 생성하되, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응된다. 이때, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함한다. 질의용어는 사용자(또는 의료인)에 의해 질의되는 용어로서, 관련용어 그래프는 질의용어 마다 별도로 생성된다.
도 6은 도 5의 개념격자 그래프를 이용하여 생성한 관련용어 그래프로서, 도 6a는 질의용어 t-1에 대한 관련용어 그래프이고, 도 6b는 질의용어 t-2에 대한 관련용어 그래프이다.
관련용어 그래프는 개념격자 그래프를 활용하여 질의용어에 대한 함의관계 및 연관관계를 추출하고, 가중치가 부여된 비순환 유방향성 그래프(또는 트리구조의 그래프)로 가시화하여 표현되는 그래프이다. 즉, 주어진 질의 Q에 대해서, 적합한 함의관계 Q⇒R과 연관관계 Q→Rminsup, minconf를 만족하는 'R'에 관련된 제반 정보를 개념격자 그래프로부터 추출하여 표시하는 그래프이다.
관련용어 그래프에서, 루트노드(이하 질의노드)는 상의 질의용어를 포함하고 상기 질의용어를 포함하는 최상위 개념노드와 대응된다. 상기 질의노드는 정의 7에 의하여 구해진다.
[정의 7]
형식문맥 K=(G, M, I)에 대한 질의Q⊆M이 주어졌을 때, Q⊆Y를 만족하는 임의의 개념(X, Y)∈B(K)에 대해서, {(X1, Y1)∈B(K)|(X, Y)
Figure PCTKR2011002264-appb-I000006
(X1,Y1)∧Q⊆Y1}=
Figure PCTKR2011002264-appb-I000007
인 경우, 개념(X, Y)을 질의 Q에 적합한 개념이라고 부르고, (Q, (X, Y))를 질의 Q에 대한 질의노드라고 정의한다.
이하에서 보듯이, 관련용어 그래프의 용어노드는 정의 7의 질의노드와 같이 "(Q, (X, Y))"의 형태로 표현된다. 이하에서, 용어노드가 (Q, (X, Y))인 경우, "용어노드는 개념노드 (X, Y)에 대응된다."라고 표현하기로 한다. 또, "용어노드의 용어는 Q이다." 또는 "용어노드는 (용어로서) Q를 포함한다."라고 표현하기로 한다. 또한, "용어노드 (Q, (X, Y))"를 간단하게 "용어노드 Q"로도 표시하기로 한다.
예를 들어 도 3b의 형식문맥에 대하여, 질의 Q={t2}라 할 때, (X, Y)=({A, B, C}, {t-1, t-2, t-4})=c2라 하면, Q⊆Y는 만족하지만, {(X1, Y1)∈B(K)|c2
Figure PCTKR2011002264-appb-I000008
(X1,Y1)∧Q⊆Y1}={c3}≠
Figure PCTKR2011002264-appb-I000009
이므로 c2는 질의 {t2}에 적합한 개념이라 할 수 없다. 반면 (X, Y)=({A, B, C, D}, {t-1, t-2})=c3라 하면, Q⊆Y를 만족하고 {(X1, Y1)∈B(K)|c4
Figure PCTKR2011002264-appb-I000010
(X1,Y1)∧Q⊆Y1}=
Figure PCTKR2011002264-appb-I000011
이므로, 개념 c3을 질의 {t2}에 적합한 개념이라 할 수 있으며 ({t2}, ({A, B, C, D}, {t-1, t-2))를 질의 {t2}에 대한 질의노드라 할 수 있다.
도 6a에서 용어노드 {t-1}이 질의노드이고, 용어노드 {t-1}은 개념노드 c4=({A, B, C, D, E}, {t-1})에 대응된다. 또, 6b에서 용어노드 {t-2}가 질의노드이고, 개념노드 ({A, B, C, D}, {t-1, t-2})에 대응된다.
한편, 관련용어 그래프에서 질의노드는 네모로 표시하기로 한다.
또한, 상기 질의노드에 대응되는 개념노드의 용어집합이 상기 질의용어 이외 다른 사용용어를 포함하면, 상기 관련용어 그래프는 상기 다른 사용용어를 포함하는 용어노드(이하 함의노드)를 상기 질의노드의 하위노드로서 구성하되, 상기 함의노드는 상기 질의노드에 대응되는 개념노드와 대응된다. 상기 함의노드를 다음 정의 8에 의해 구해진다.
[정의 8]
형식문맥 K=(G, M, I)에서, 임의로 주어진 질의 Q⊆M에 대한 질의노드(Q, (X, Y))에 대하여, Q≠Y ∧
Figure PCTKR2011002264-appb-I000012
≥minsup∈[0, 1]인 경우, (Y\Q, (X, Y))를 질의 Q에 대한 함의노드라 정의한다.
함의관계를 표현하는 함의노드를 정의함에 있어 경계 값인 minsup을 반영하는 이유는, 앞서 개념격자의 설명에서 언급했듯이, 함의관계는 연관관계의 특수한 경우이기 때문이다. 즉 고유한 의미의 함의관계를 반영하는 함의노드는 minsup=0인 경우라 할 수 있다. 반면 함의관계를 표현하는 함의노드를 정의함에 있어, minconf를 반영하지 않은 이유는 다음과 같다. 질의 Q에 적합한 개념을 (X, Y)∈B(K)라 할 때, Q≠Y가 성립하여 함의노드가 존재할 경우, 정의2에 의하여, 항상 extent(Q)⊆extent(Q\Y), 다시 말해 Q⇒Q\Y를 만족하므로, 즉 minconf=1이므로 고려하지 않는다.
예를 들어 도 3b의 형식문맥에 대하여, 질의 Q={t2}, minsup=0.4일 때, 질의 Q에 적합한 개념은 c3이며, 질의노드는 ({t2}, c3)이다. 또한 질의 {t2}가 c3의 intent인 {t1, t2}와 같지 않으며, |{A, B, C, D}|/|G|=0.8≥minsup이므로, 질의 {t2}에 대한 함의노드는 ({t1}, ({A, B, C, D}, {t1, t2}))라고 할 수 있다.
도 6b에서 용어노드 {t-1}이 함의노드이고, 이 용어노드는 개념노드 ({A, B, C, D}, {t1, t2})에 대응된다.
또한, 질의용어를 포함하되 최상위가 아닌 개념노드와 대응되는 용어노드(이하 연관노드)는 자신에 대응되는 개념노드의 용어집합에서 자신의 모든 상위노드에 포함되는 용어들을 제외한 용어를 포함한다. 그리고 상기 연관노드 또는 상기 질의노드 사이의 상하위 노드 관계는 대응되는 개념노드 간의 상하위 노드 관계와 같다. 이때, 관련용어 그래프는 사용주체 집합이 적어도 하나의 원소를 가지는 개념노드에 대해서만 대응되는 용어노드로 구성한다. 구체적으로, 상기 연관노드는 정의 9에 의해 구해진다.
[정의 9]
형식문맥 K=(G, M, I)에 대한 주어진 질의 Q⊆M에 적합한 개념이 (X, Y)∈B(K) 일 때, 임의의 (X1, Y1), (X2, Y2)∈B(K)가, (X1, Y1)<(X, Y) ∧ (X2, Y2)≤(X, Y) ∧ (X1, Y1)
Figure PCTKR2011002264-appb-I000013
(X2, Y2)를 만족하고,
Figure PCTKR2011002264-appb-I000014
≥minsup,
Figure PCTKR2011002264-appb-I000015
≥minconf 이면, (Y1\Y2, (X1, Y1))를 질의Q에 대한 연관노드이다 라고 정의한다. 단 minsup, minconf∈[0, 1].
주어진 질의 Q에 적합한 개념이 (X, Y)라면, (X1, Y1)<(X, Y) 관계에 있는 개념 (X1, Y1)의 extent인 X1은, 정의 3에 의해, Y⊃Q를 사용용어로 가지면서 Y1\Y를 추가 사용용어를 사용하는 사용주체의 집합이라 할 수 있다. 이때, 정의 3(상하위관계)과 정의 4(근접하위이웃 및 근접상위이웃)를 활용하여, 주어진 질의Q와 연관관계에 있는 사용용어를 도출해 낼 수 있다. 즉, 관련용어 그래프의 용어노드 간의 상하위 노드 관계는 대응되는 (개념노드 그래프의) 개념노드 간의 상하위 노드 관계와 같다.
예를 들어 도 3b의 형식문맥에 대하여, 질의 Q={t1}, 경계 값 minsup=0.1, minconf=0.25이라 하면, 질의 {t1}에 적합한 개념은 c4가 된다. 이때 (X1, Y1)=c2, (X2, Y2)=c3이라 하면, c3<c4, c2≤c4, c2
Figure PCTKR2011002264-appb-I000016
c3을 만족하며, |X1|/|G|=0.6≥minsup, |X1|/|X|=0.6≥minconf가 성립하므로, Y1\Y2={t-1, t-2, t-4}\{t-1, t-2}={t-4}가 된다. 즉, Q={t1}이고, minsup=0.1, minconf=0.25인 경우, ({t-4}, ({A, B, C}, {t-1, t-2, t-4}))를 연관노드라 할 수 있다.
지금까지 정의한, 형식문맥 K(G, M, I)에 주어진 질의 Q⊆M, 각 경계 값 minsup, minconf에 대한, 질의노드, 함의노드 및 연관노드는, (Y*,(X,Y))와 같은 형태로 일반화하여 나타낼 수 있다(단, (X, Y)∈B(K), Y*⊆Y).
이 때, Y*를 용어노드의 용어(또는 레이블), (X, Y)를 용어노드에 대응되는 개념(Concept) 또는 개념노드라 하며, 모든 노드들의 집합을 N(KQ)minsup,minconf라 표기하며, 이들 노드로부터 구성되는 그래프를 관련용어 그래프라 한다. 이하에서 용어노드를 표시하는 방식 중 하나로 레이블(또는 용어노드의 용어)로 표시한다.
도 6a 및 도 6b는 각각 질의용어 {t1}과 {t2}에 대한 관련용어 그래프이고, 질의노드는 네모난 노드로, 함의노드는 마름모꼴 노드로, 연관노드는 원형 노드로 표시된다.
다음으로, 질의용어 표시부(34)는 질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도 또는 확신도로 표시한다. 바람직하게는, 각 용어노드에 포함된 사용용어 및 가중치를 표시해준다. 바람직하게는 가중치는 지지도 또는 확신도를 표시해준다.
바람직하게는, 용어노드에 전체 사용주체의 개수 대비 대응되는 개념노드의 사용주체들의 개수(사용주체 집합의 원소 개수)의 비율을 지지도로 표시한다. 또, 질의노드에 대응되는 개념노드의 사용주체들의 개수 대비 자신과 대응되는 개념노드의 사용주체들의 개수의 비율을 확신도로 표시한다.
용어노드 간의 연관관계를 정의 10에 의해 구해진다.
[정의 10]
형식문맥 K=(G, M, I)에 주어진 질의 Q⊆M, minsup, minconf에 관한 두 노드 n1=(Y1*,(X1,Y1)), n2=(Y2*,(X2,Y2))∈N(KQ)minsup,minconf 에 대하여 수학식 1과 같은 조건을 만족하면 n1은 n2와 질의 Q에 대해 연관관계에 있다고 정의하고 r(n1,n2)로 표기한다.
[수학식 1]
Figure PCTKR2011002264-appb-I000017
또한, 용어노드 간의 연관관계에 대한 관계 가중치는 정의 11과 같이 정의된다.
[정의 11]
형식문맥 K=(G, M, I)에 주어진 질의 Q⊆M, minsup, minconf에 관한 두 노드 n1=(Y1*,(X1,Y1)),n2=(Y2*,(X2,Y2))∈N(KQ)minsup,minconf에 대하여, 관계 r=(n1,n2)∈R이 존재할 때, 관계 r의 가중치를 W(r) = (
Figure PCTKR2011002264-appb-I000018
,
Figure PCTKR2011002264-appb-I000019
)라 정의하며, 를 관계 r의 지지도(support),
Figure PCTKR2011002264-appb-I000020
를 관계 r의 확신도(confidence)라 각각 정의한다.
정의 11에 의한 관계 가중치인 지지도 또는 확신도를 용어관련 그래프에 각 용어노드에 표시한다. 관련용어 그래프는 트리형태의 비선형 그래프이므로 n2는 항상 하나의 관계 r=(n1,n2)만을 갖는다. 따라서 r=(n1,n2)의 가중치를 노드 n2에 기재한다.
도 6a 또는 도 6b는 관련용어 그래프에 가중치로서 지지도를 표시하고 있다(확신도는 생략함). 즉, 용어노드 {t-1}에서 용어노드 {t-2}로의 연관관계의 지지도(또는 노드 t-2에서 지지도)는 |{A, B, C, D}|/|G|=80%이고, 용어노드 {t-1}에서 용어노드 {t-3}로의 연관관계의 지지도(또는 노드 t-3에서 지지도)는 |{E}|/|G|=20%이다.
도 6a의 관련용어 그래프에서, 용어노드 {t-1} 표로 나타낸 것이 도 7과 같다.
한편, 바람직하게는, 질의용어 표시부(34)는 관련용어 그래프의 깊이에 따라 단계적으로 각 용어노드에 해당하는 사용용어 및 지지도 또는 확신도를 표시해준다. 예를 들어, 도 7과 같이, 도 6a의 관련용어 그래프를 깊이 순서에 따라 순위를 두어 순차적으로 사용용어 및 지지도 또는 확신도를 표시해준다.
[정의 12]
형식문맥 K=(G, M, I), Q⊆M, minsup, minconf∈[0, 1]에 대한 용어관련그래프 (N(KQ)minsup,minconf, R)에 대하여, 임의의 순서집합(Ordered set)을 P={n1,n2,n3,…,nm}⊆N(KQ)minsup,minconf라 할 때, n1∈P의 용어는 Q이고, 1≤k≤m-1를 만족하는 k에 대하여, 모든 nk, nk+1∈P에 있어, r(nk,nk+1)∈R을 만족할 때, 순서집합 P를 관련용어 그래프의 해석 경로라 정의하며, 'n1의 용어→ n2의 용어∧...∧nm의 용어' 이라 해석한다. 또한 해석 경로 P의 가중치는 r(nm-1,nm)∈R 의 가중치라 정의한다.
도 6a에 표현된 관련용어 그래프의 해석 경로의 한 예를 살펴보면 다음과 같다. 용어노드는 각각 n1=({t-1},({A,B,C,D,E}, {t-1})), n2=({t-2},({A,B,C,D}, {t-1, t-2})), n3=({t-4},({A,B,C}, {t-1, t-2, t-4}))이다. 이를 보다 문맥에 맞게 해석하면, '맹장염을 사용하는 사용주체는 어펜티시티스와 충수돌기염을 사용하는 경향이 있다' 라고 해석할 수 있다.
또한 수치정보인 가중치 정보를 부과하여 해석하면 '맹장염를 사용하는 병원 중 25% 이상(정확히 60%)의 병원이 어펜티시티스와 충수돌기염을 사용하는 경향이 있다.' 라고 해석할 수 있다.
다음으로, 본 발명의 제2 실시예에 따라 사용주체의 등급 가중치를 이용하여 지지도 또는 확신도를 구하는 방식을 도 8과 도 9를 참조하여 설명한다. 도 8은 제2 실시예에 따른 사용주체에서 서로 관련 있게 사용되는 사용용어에 대한 형식문맥을 예시한 표이고, 도 9는 관련용어 그래프의 예시도 및 관련용어의 지지도에 대한 표이다.
사용용어 수집부(31)은 각 사용주체에서 서로 관련 있게 사용되는 사용용어를 집합과 더불어 각 사용주체의 등급 가중치를 입력받는다. 등급 가중치란 의료현장에서 사용주체의 영향력에 따라 부여된 것이다.
3차원 병원에서 사용하는 사용용어는 1차 병원에서 사용되는 사용용어 보다는 의료현장에서 미치는 영향력은 더 크다. 또한, 동일한 1차 병원이더라도 본원이 분원 보다 더 큰 영향력을 가지고 있다. 또한, 동일한 병원 내에서도 사용용어의 사용자에 따라 그 영향력이 다르다. 예를 들어, 의사, 수간호사, 일반 간호사 등에 따라 역시 다르다. 따라서 상기와 같은 영향력에 따라 각 사용자 또는 병원의 등급을 구분하여 가중치를 달리 정한다.
일실시예로서, 등급 가중치는 병상(bed) 수, 병원의 차수(1,2,3차), 매핑 행위자의 직책, 전문성(specialty) 등 다양한 변수에 의해 정해진다.
등급가중치는 관련용어 그래프의 용어노드의 지지도 또는 확신도에 반영하는데, 형식문맥을 작성할 때 해당 사용주체의 사용용어 집합을 중복시키는 방식과 지지도 또는 확신도를 구할 때 사용주체의 개수 대신 등급 가중치로 구하는 방식이 있을 수 있다.
전자의 방식의 예로서, 도 8에서 보는 바와 같이, 등급 가중치가 높은 병원 E의 사용용어 집합을 형식문맥(formal context)에 중복하여 삽입한다. 도 8에서는 추가적으로 5개의 사용용어 집합을 더 삽입하였다.
형식문맥에 중복 삽입되면, 개념격자 그래프의 개념노드의 사용용어 집합이 변경되고, 이에 따라 관련용어 그래프도 달리 생성된다. 예를 들면, 도 9a는 도 8의 형식문맥을 기반으로 생성된 질의 {t-1}에 대한 관련용어 그래프이다. 도 6a와 비교하면, 용어노드 {t-2}의 지지도는 80%에서 40%으로 감소되었고, 용어노드 {t-3}의 지지도는 20%에서 60%으로 증가되었다. 즉, 맹장염을 막창꼬리염으로 표시하는 병원은 병원 E 밖에 없으나, 병원 E의 등급 가중치가 매우 크므로, 막창꼬리염이 어펜티시티스(Appenticitis) 보다 더 관련성이 있는 사용용어로 표시된다.
따라서 질의자에게 보여주는 용어 관련성의 우선순위도 달리 표시된다. 도 9b를 도 7과 대비하여 보면, 막창꼬리염(t-3)이 어펜티시티스(t-2) 보다 더 우선적으로 표시되게 된다.
이 경우, 등급가중치는 정수배이다. 즉, 정수배 만큼 사용용어에 대한 사용용어 집합을 중복한다.
사용주체의 개수 대신 등급 가중치로 구하는 지지도 또는 확신도를 구하는 방식은 다음 정의 9-2에 의해 구해진다. 정의 9-2는 앞서 설명한 정의 9에서 지지도 및 신뢰도를 사용주체의 개수가 아닌 등급 가중치에 의해 정의된 것이다.
즉, 개념의 지지도는 전체 사용주체의 등급 가중치의 합 대비 자신의 사용주체들의 등급 가중치의 합의 비율로 계산된다. 개념의 확신도는 연관개념의 사용주체들의 등급 가중치의 합 대비 자신의 사용주체들의 등급 가중치의 합의 비율로 계산된다.
[정의 9-2]
형식문맥 K=(G, M, I)에 대한 주어진 질의 Q⊆M에 적합한 개념이 (X, Y)∈B(K) 일 때, 임의의 (X1, Y1), (X2, Y2)∈B(K)가, (X1, Y1)<(X, Y) ∧ (X2, Y2)≤(X, Y) ∧ (X1, Y1)
Figure PCTKR2011002264-appb-I000021
(X2, Y2)를 만족하고,
Figure PCTKR2011002264-appb-I000022
≥minsup,
Figure PCTKR2011002264-appb-I000023
≥minconf 이면, (Y1Y2, (X1, Y1))를 질의Q에 대한 연관노드이다 라고 정의한다.
단 minsup, minconf∈[0, 1], Wh는 사용주체 h의 등급가중치이다.
또한, 용어노드 간의 등급가중치를 반영한 연관관계에 대한 관계 가중치는 정의 11-2과 같이 정의된다.
[정의 11-2]
형식문맥 K=(G, M, I)에 주어진 질의 Q⊆M, minsup, minconf에 관한 두 노드 n1=(Y1*,(X1,Y1)),n2=(Y2*,(X2,Y2))∈N(KQ)minsup,minconf에 대하여, 관계 r=(n1,n2)∈R이 존재할 때, 관계 r의 가중치를 W(r) = (
Figure PCTKR2011002264-appb-I000024
,
Figure PCTKR2011002264-appb-I000025
)라 정의하며,
Figure PCTKR2011002264-appb-I000026
를 관계 r의 지지도(support),
Figure PCTKR2011002264-appb-I000027
를 관계 r의 확신도(confidence)라 각각 정의한다. 단, Wh는 사용주체 h의 등급가중치이고, Q' = extent(Q).
즉, 용어노드의 지지도는 전체 사용주체의 등급 가중치의 합 대비 자신과 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산된다. 용어노드의 확신도는 질의노드에 대응되는 개념노드의 사용주체들의 등급 가중치의 합 대비 자신과 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산된다.
앞서 예에서, 병원 A, B, C, D의 등급 가중치는 1로 정하고, 병원 E의 등급 가중치는 6으로 정하면, 중복삽입 방식과 동일한 결과가 도출된다.
다음으로, 본 발명의 제3 실시예에 따라 적어도 2개의 다른 매핑정보를 가지는 형식문맥을 이용하여 관련용어 그래프를 구하는 방식을 도 10을 참조하여 설명한다.
도 10에서 보는 바와 같이, 사용주체에서 사용하는 사용용어의 매핑정보가 적어도 2개의 다른 매핑정보로 입력되는 경우가 있다. 즉, {t-1, t-2, t-3, t-4}의 사용용어와 {t-31, t-32, t-33}의 사용용어는 서로 다를 뿐만 아니라, 이를 이용하는 사용주체 집합도 각각 {A, B, C, D, E}와 {X, Y, Z}로 각기 다르다.
이와 같은 경우, 하나의 형식개념(또는 매핑정보)으로 입력된다하더라도 개념격자 그래프는 G1의 매핑정보와 G2의 매핑정보 별로 따로 형성된다. 이때 지지도를 구할 때 이용되는 전체 사용주체의 수, 즉, |G| ({A, B, C, D, E, X, Y, Z}의 개수)를 이용한다면, G1의 개념과 G2의 개념이 다를 수 있으므로 분석의 정확도가 떨어진다.
따라서 개념격자 그래프가 분리되어 생성되면, 분리된 개념격자 그래프 단위로 전체 사용주체의 수를 달리 정하는 것이 바람직하다. 즉, 분리된 개념격자 그래프 G1에 대해서는 |G1|, 즉, {A, B, C, D, E}의 개수로 이용하고, 분리된 개념격자 그래프 G2에 대해서는 |G2|, 즉, {X, Y, Z}의 개수로 이용한다.
즉, 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체으로 한정한다.
등급 가중치를 이용하는 경우에는 동일하게 적용될 수 있다.
다음으로, 본 발명의 제4 실시예에 따른 개념격자 기반 질의용어 매핑지원 방법을 도 11과 도 12를 참조하여 설명한다.
도 11에서 보는 바와 같이, 먼저, 각 사용주체에서 서로 관련 있게 사용되는 사용용어를 집합 및 각 사용주체의 등급 가중치를 입력받아(S10), 용어집합과 상기 용어집합을 공통적으로 사용하는 사용주체 집합을 노드(이하 개념노드)로 구성하는 개념격자 그래프를 형성한다(S20). 이때, 하위 개념노드의 용어집합은 상위 개념노드의 용어집합을 모두 상속한다.
다음으로, 질의용어를 입력받아(S30), 질의용어에 대하여 용어노드로 구성하는 관련용어 그래프를 생성한다(S40). 이때, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응되도록 구성하고, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함하도록 관련용어 그래프를 구성한다. 관련용어 그래프를 생성하는 구체적인 방법은 도 12에서 보는 바와 같다.
상기 관련용어 그래프는 다음과 같이 형성한다. 상기 관련용어 그래프의 루트노드(이하 질의노드)는 상기 질의용어를 포함하는 최상위 개념노드와 대응되고, 자신의 용어로서 상기 질의용어를 포함하도록 구성한다. 상기 최상위 개념노드와 대응되지 않는 용어노드(이하 연관노드)는 자신에 대응되는 개념노드의 용어집합에서 자신의 모든 상위노드에 포함되는 용어들을 제외한 용어를 포함한다. 그리고 상기 연관노드 또는 상기 질의노드 사이의 상하위 노드 관계는 대응되는 개념노드 간의 상하위 노드 관계와 같이 구성한다.
또한, 상기 질의노드에 대응되는 개념노드의 용어집합이 상기 질의용어 이외 다른 사용용어를 포함하면, 상기 관련용어 그래프는 상기 다른 사용용어를 포함하는 용어노드(이하 함의노드)를 상기 질의노드의 하위노드로서 구성하되, 상기 함의노드는 상기 질의노드에 대응되는 개념노드와 대응되게 구성한다.
또한, 상기 용어노드의 지지도는 전체 사용주체의 개수 대비 자신과 대응되는 개념노드의 사용주체 집합개수의 비율로 구해진다. 이때, 상기 지지도는 상기 등급 가중치를 반영한다. 또, 상기 용어노드의 확신도는 질의노드에 대응되는 개념노드의 사용주체들의 개수 대비 자신과 대응되는 개념노드의 사용주체들의 개수의 비율로 구해진다. 이때, 상기 확신도는 상기 등급 가중치를 반영한다.
등급 가중치를 반영하는 방식의 일례로서, 등급가중치는 정수배이고, 각 사용주체에 대한 사용용어 집합을 상기 등급가중치 만큼 더 중복하여, 상기 지지도 또는 확신도를 구한다.
또, 다른 일례로서, 용어노드의 지지도는 전체 사용주체의 등급 가중치의 합 대비 자신과 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산된다. 용어노드의 확신도는 질의노드에 대응되는 개념노드의 사용주체들의 등급 가중치의 합 대비 자신과 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산된다.
한편, 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체으로 한정된다.
그리고 질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도 또는 확신도로 표시한다(S50). 바람직하게는, 상기 관련용어 그래프의 깊이에 따라 단계적으로 각 용어노드에 해당하는 사용용어 및 지지도를 표시한다.
상기 개념격자 기반 질의용어 매핑지원 방법 중 생략된 부분은 앞서 설명한 개념격자 기반 질의용어 매핑지원 시스템의 설명을 참조한다.
이상, 본 발명자에 의해서 이루어진 발명을 상기 실시 예에 따라 구체적으로 설명하였지만, 본 발명은 상기 실시 예에 한정되는 것은 아니고, 그 요지를 이탈하지 않는 범위에서 여러 가지로 변경 가능한 것은 물론이다.
본 발명은 다수의 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 질의용어 매핑지원 시스템을 개발하는 데 적용이 가능하다.
특히, 본 발명은 질의용어와 관련된 용어들을 용어노드로 표시되는 계층구조의 그래프로 구성하고, 질의용어를 루트노드로 하여 각 용어노드에 상위 용어노드의 관련성을 지지도 또는 확신도로 표시하는 질의용어 매핑지원 시스템을 개발하는 데 유용하다.

Claims (19)

  1. 다수의 사용주체에 의해 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 시스템에 있어서,
    각 사용주체에 의해 서로 관련 있게 사용되는 용어를 집합(이하 용어집합)으로 입력받는 사용용어 수집부;
    용어집합과 상기 용어집합을 서로 관련 있게 사용하는 사용주체 집합을 노드(이하 개념노드)로 구성하는 개념격자 그래프를 형성하되, 하위 개념노드의 용어 집합은 상위 개념노드의 용어집합을 모두 상속하는 개념격자 형성부; 및
    질의용어를 입력받아 상기 질의용어에 대하여 용어노드로 구성하는 관련용어 그래프를 생성하되, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응되고, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함하는 질의그래프 생성부를 포함하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  2. 제1항에 있어서,
    상기 관련용어 그래프의 루트노드(이하 질의노드)는 상기 질의용어를 포함하는 최상위 개념노드와 대응되고, 자신의 용어로서 상기 질의용어를 포함하고,
    상기 최상위 개념노드와 대응되지 않는 용어노드(이하 연관노드)는 자신에 대응되는 개념노드의 용어집합에서 자신의 모든 상위노드에 포함되는 용어들을 제외한 용어를 포함하고,
    상기 연관노드 또는 상기 질의노드 사이의 상하위 노드 관계는 대응되는 개념노드 간의 상하위 노드 관계와 같은 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  3. 제2항에 있어서,
    상기 용어노드는 자신과 대응되는 개념노드의 의료기관 집합개수를 이용하여 계산된 지지도 또는 확신도를 가지는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  4. 제3항에 있어서,
    상기 사용용어 수집부는 각 사용주체의 등급 가중치를 입력받고,
    상기 지지도 또는 확신도는 상기 등급 가중치를 반영하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  5. 제2항 또는 제4항에 있어서,
    상기 질의노드에 대응되는 개념노드의 용어집합이 상기 질의용어 이외 다른 사용용어를 포함하면, 상기 관련용어 그래프는 상기 다른 사용용어를 포함하는 용어노드(이하 함의노드)를 상기 질의노드의 하위노드로서 구성하되, 상기 함의노드는 상기 질의노드에 대응되는 개념노드와 대응되는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  6. 제4항에 있어서,
    상기 등급가중치는 정수배이고, 각 사용주체에 대한 사용용어 집합을 상기 등급가중치 만큼 더 중복하여, 상기 지지도 또는 확신도가 구해지는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  7. 제4항에 있어서,
    상기 용어노드의 지지도 또는 확신도는 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산되는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  8. 제6항 또는 제7항에 있어서,
    상기 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체으로 한정되는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  9. 제6항 또는 제7항에 있어서, 상기 시스템은,
    질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도 또는 확신도로 표시해주는 질의용어 표시부를 더 포함하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  10. 제9항에 있어서,
    상기 질의용어 표시부는 상기 관련용어 그래프의 깊이에 따라 단계적으로 각 용어노드에 해당하는 사용용어 및 지지도 또는 확신도를 표시해주는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  11. 제10항에 있어서,
    상기 질의용어 표시부는 상기 관련용어 그래프에서 지지도 또는 확신도가 최소 지지도 또는 최소 확신도 이상인 용어노드만을 표시해주는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 시스템.
  12. 다수의 사용주체에서 서로 관련 있게 사용되는 사용용어들에 대한 개념격자를 형성하여, 질의용어를 서로 관련 있게 사용되는 사용용어로 매핑하는 것을 지원하는 개념격자 기반 질의용어 매핑지원 방법에 있어서,
    (a) 각 사용주체에서 서로 관련 있게 사용되는 사용용어를 집합(이하 용어집합)으로 입력받는 단계;
    (b) 용어집합과 상기 용어집합을 공통적으로 사용하는 사용주체 집합을 노드(이하 개념노드)로 구성하는 개념격자 그래프를 형성하되, 하위 개념노드의 용어 집합은 상위 개념노드의 용어 집합을 모두 상속하는 단계; 및
    (c) 질의용어를 입력받아 상기 질의용어에 대하여 용어노드로 구성하는 관련용어 그래프를 생성하되, 상기 용어노드는 상기 개념격자 그래프의 개념노드 중 하나와 대응되고, 상기 용어노드와 대응되는 개념노드는 상기 질의용어를 포함하는 단계를 포함하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  13. 제12항에 있어서,
    상기 용어노드는 자신과 대응되는 개념노드의 의료기관 집합개수를 이용하여 계산된 지지도 또는 확신도를 가지는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  14. 제13항에 있어서,
    상기 (a)단계에서, 각 사용주체의 등급 가중치를 입력받고,
    상기 지지도 또는 확신도는 상기 등급 가중치를 반영하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  15. 제14항에 있어서,
    상기 등급가중치는 정수배이고, 각 사용주체에 대한 사용용어 집합을 상기 등급가중치 만큼 더 중복하여, 상기 지지도 또는 확신도가 구해지는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  16. 제14항에 있어서,
    상기 용어노드의 지지도 또는 확신도는 대응되는 개념노드의 사용주체들의 등급 가중치의 합의 비율로 계산되는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  17. 제15항 또는 제16항에 있어서,
    상기 개념격자 그래프가 적어도 2개의 그래프로 분리되는 경우, 상기 지지도 또는 확신도의 전체 사용주체는 상기 용어노드와 대응되는 개념노드를 포함하는 분리된 그래프에 포함되는 사용주체으로 한정되는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  18. 제15항 또는 제16항에 있어서, 상기 방법은,
    (d) 질의용어를 입력받고 상기 질의용어에 대한 관련 사용용어 간의 관계를 상기 질의용어의 관련용어 그래프로 표시하고, 각 용어노드의 상위 노드와의 관계 정도를 지지도 또는 확신도로 표시해주는 질의용어 단계를 더 포함하는 것을 특징으로 하는 개념격자 기반 질의용어 매핑지원 방법.
  19. 제12항 내지 제16항 중 어느 한 항의 개념격자 기반 질의용어 매핑지원 방법을 수행하는 프로그램을 기록한 컴퓨터로 읽을 수 있는 기록매체.
PCT/KR2011/002264 2010-04-01 2011-04-01 사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법 Ceased WO2011122897A2 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR10-2010-0030059 2010-04-01
KR1020100030059A KR101078862B1 (ko) 2010-04-01 2010-04-01 사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법

Publications (2)

Publication Number Publication Date
WO2011122897A2 true WO2011122897A2 (ko) 2011-10-06
WO2011122897A3 WO2011122897A3 (ko) 2012-01-12

Family

ID=44712789

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2011/002264 Ceased WO2011122897A2 (ko) 2010-04-01 2011-04-01 사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법

Country Status (2)

Country Link
KR (1) KR101078862B1 (ko)
WO (1) WO2011122897A2 (ko)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102071702B1 (ko) 2017-10-30 2020-01-30 서울대학교산학협력단 연관 규칙의 단계적 추출 장치 및 방법

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
AU777693B2 (en) * 1999-03-05 2004-10-28 Canon Kabushiki Kaisha Database annotation and retrieval
US20050149510A1 (en) * 2004-01-07 2005-07-07 Uri Shafrir Concept mining and concept discovery-semantic search tool for large digital databases

Also Published As

Publication number Publication date
KR101078862B1 (ko) 2011-11-02
WO2011122897A3 (ko) 2012-01-12
KR20110110628A (ko) 2011-10-07

Similar Documents

Publication Publication Date Title
WO2021060899A1 (ko) 인공지능 모델을 사용 기관에 특화시키는 학습 방법, 이를 수행하는 장치
Carroll et al. Visualization and analytics tools for infectious disease epidemiology: a systematic review
WO2010087566A1 (en) Document analysis system
US20080221846A1 (en) Context-Aware, Adaptive Approach to Information Selection for Interactive Information Analysis
WO2022055186A1 (ko) 의료 영상에서 이상 소견 탐지 및 판독문 생성 방법
WO2016190496A1 (ko) 의료용 메타 데이터베이스 관리 방법 및 그 장치
US20080077564A1 (en) Document-search supporting apparatus and computer program product therefor
WO2012046906A1 (ko) 다중 자원을 통합한 지식베이스를 이용하여 연구 주체간의 상관관계가 표시된 자원검색 정보 제공 장치 및 방법
WO2011136425A1 (ko) 개체명 사전 및 마이닝 규칙이 결합된 온톨로지 스키마를 이용한 리소스 기술 프레임워크 네트워크 구축 장치 및 방법
WO2020242086A1 (ko) 다중 지식의 비교 우위를 추론하는 서버, 방법 및 컴퓨터 프로그램
US20180101581A1 (en) System and method for data management
WO2020122546A1 (ko) 특허와 논문 데이터를 활용한 국가 및 기업들의 과학 기술력 진단 및 예측 방법
WO2023195768A1 (ko) 자연어 처리 기반의 유사도 판단을 통한 특허 문헌의 시각화 방법 및 이를 제공하는 장치
US8611675B2 (en) Confusion matrix for classification systems
WO2022050551A1 (ko) 법률 서비스 제공 시스템 및 그 방법
WO2011136491A2 (ko) 용어체계 기반의 데이터 개체 정의지원 시스템
WO2022231392A1 (ko) 자동 기계학습을 통한 자동 진화형 플랫폼 구현 방법 및 장치
JPWO2019239543A1 (ja) 質問応答装置、質問応答方法およびプログラム
JP2007140859A (ja) 情報処理システム
WO2011122897A2 (ko) 사용주체의 가중치를 이용한 개념격자 기반 질의용어 매핑지원 시스템 및 방법
WO2012060502A1 (ko) 연구주체간의 상관관계 추론을 위한 시스템 및 방법
Li et al. Study on unbiased interval grey number prediction model with new information priority
WO2017191877A1 (ko) 프로버넌스 관리를 위한 압축 장치 및 방법
US20050038812A1 (en) Method and apparatus for managing data
WO2025249797A1 (ko) 개인 지식 그래프를 관리하는 방법 및 이를 이용하는 사용자 디바이스

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 11763069

Country of ref document: EP

Kind code of ref document: A2

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 11763069

Country of ref document: EP

Kind code of ref document: A2