EP4165519A1 - Procede et systeme de fusion d'informations - Google Patents
Procede et systeme de fusion d'informationsInfo
- Publication number
- EP4165519A1 EP4165519A1 EP20731485.7A EP20731485A EP4165519A1 EP 4165519 A1 EP4165519 A1 EP 4165519A1 EP 20731485 A EP20731485 A EP 20731485A EP 4165519 A1 EP4165519 A1 EP 4165519A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- information
- individuals
- instances
- property
- evolution
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/36—Creation of semantic tools, e.g. ontology or thesauri
- G06F16/367—Ontology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F7/00—Methods or arrangements for processing data by operating upon the order or content of the data handled
- G06F7/06—Arrangements for sorting, selecting, merging, or comparing data on individual record carriers
- G06F7/14—Merging, i.e. combining at least two sets of record carriers each arranged in the same ordered sequence to produce a single set having the same ordered sequence
Definitions
- TITLE PROCESS AND SYSTEM FOR MERGING INFORMATION
- the technical field of the present invention relates to information fusion methods and systems.
- the technical field of the present invention is also that of situational awareness methods and systems which are used to detect abnormal behavior of individuals (vehicle, person, etc.) and which are based on such fusion methods and systems. information.
- the information to be processed to establish such an operational table can come from various sources. Two categories of information provided can be distinguished: so-called “hard” information and “flexible” information ("soft" in English).
- the hard information provides a quantitative evaluation of elements and comes from physical sensors (camera, microphone, radar ).
- the flexible information comes from an extraction of linguistic content (observer report, text, phone call %) allowing a qualitative assessment of elements and possible relationships between them.
- hard information is precise information that can most often be reduced to a numerical value
- flexible information is information that is often difficult to reduce to a numerical value, requiring knowledge of the context in which the information was acquired. to understand it and which is difficult to use when isolated from the environment in which said information was collected.
- Information fusion involves several steps, the two main ones being (1) a calculation of similarity distance between the different information available, although these information is of a varied nature, and (2) G association of this information, or not, depending on the result of the similarity calculation.
- the objective here is to detect whether various information received concerns the same individual or not.
- the term “individual” is understood in the broad sense in the field of information fusion, namely a separate unit (entity) in a domain of interpretation (person, vehicle, object, group, etc.
- the information fusion solutions in the literature make a strict comparison between the properties of individuals detected in the information received at a given point in time, regardless of the time difference between the points in time when the information in question was generated. For example, when a maritime surveillance system attempts to compare information relating to a vessel observed three days ago with information relating to a vessel observed more recently in order to determine whether it is the same individual or not, the identity of the captain is at that time more reliable information than the respective positions of these vessels.
- the approach used is therefore a brake on the automation of information fusion processes, which then need, from an operational point of view, human intervention to ensure that a similarity detected between information is effectively a matter of concern. a correlation and not a simple coincidence without reality on the ground.
- An object of the present invention is to provide a method of processing information which originates from various sources and from which instances of individuals are generated by ontology alignment, the method of processing information comprising a fusion of information aimed at merging the instances of individuals which correspond to the same individual, the method being implemented by a data processing system, characterized in that the method comprises the following steps: generating the instances of individuals using an ontology which defines, for each property of each instance of individual, an evolution model to be applied to said property, the evolution model represents the evolution of the reliability of said property over time in relation to the variability over time of said property; merge information by comparing in pairs the instances of individuals generated with instances of individuals stored in the knowledge base, by performing for each joint property a similarity distance calculation in application at least of the model of evolution defined for said property, so as to define a confidence coefficient for each property to decide whether or not to merge said instances of individuals; and updating the knowledge base with the instances of individuals resulting from the information fusion.
- the fusion of information is efficient, because it limits the taking into account of properties according to
- each evolution model is of one type among the following three possible types: constant, for the properties which do not change over time; predictive, for properties which can be estimated over a certain limited period of time or with a certain uncertainty which evolves over time; and circumstantial, for properties whose evolution over time depends on the occurrence of an event.
- the properties are associated with evolution models adapted to different types of property variability.
- the circumstantial model of evolution is exponentially decreasing.
- the circumstantial model of evolution is exponentially decreasing.
- each instance of an individual which results from the merger of two other instances of an individual retains only one value for each property among those available in said other instances of an individual and the value retained depends on the model evolution with which said property is associated.
- the information fusion is refined.
- the value kept is that having the best precision
- the retained value is the most recent
- the conserved value is that showing the highest confidence coefficient according to the following system of equations:
- index 72 ⁇ 2 where the index "1" represents the oldest information and the index "2" represents the most recent information, where l is the coefficient representative of a reliability of the source having carried out the capture of the information considered, t is an exponential decay accentuation time factor, and t represents the instant of capture of the information considered.
- the method further comprises the following step: exploiting the results obtained by merging information in a situation management system, and detecting abnormal behavior of individuals using a set of predefined rules , or to a situation ontology model, and to instances of individuals resulting from the fusion of information.
- a human operator in deciding whether the information presented to him is duplicate or whether said information does indeed relate to distinct individuals is limited.
- the similarity distance calculation by applying at least the evolution model is aggregated with at least one other similarity calculation.
- the information fusion is refined.
- the similarity calculations are weighted.
- the merging of information can be easily personalized for a specific use case (maritime surveillance, etc.).
- a said further calculation of similarity distance is a calculation of taxonomic similarity distance and said further calculation of domain similarity distance is a range domain similarity distance calculation.
- the calculation of the similarity distance in application at least of the evolution model applies a reliability coefficient of the sources having captured the information considered. Thus, more credit can easily be given to information from reliable sources.
- the information to be processed is flexible information and / or hard information.
- information fusion is effective regardless of the nature, hard or flexible, of the information collected.
- the invention also relates to a computer program, which can be stored on a medium and / or downloaded from a communication network, in order to be read by a processor.
- This computer program includes instructions for implementing the above-mentioned method in any of their embodiments, when said program is executed by the processor.
- the invention also relates to an information storage medium storing such a computer program.
- the invention also relates to an information processing system which originates from various sources and from which instances of individuals are generated by ontology alignment, the information processing system comprising electronic circuitry implementing a fusion of individuals.
- the electronic circuitry implements: means for generating the instances of individuals using an ontology which defines, for each property of each instance of individual, an evolution model to be applied to said property, the evolution model represents the evolution of the reliability of said property over time in relation to the variability over time of said property; means for performing the merging of information by comparing two by two instances of individuals generated with instances of individuals stored in the knowledge base, by performing for each common property a similarity distance calculation in application of at least the evolution model defined for said property, so as to define a confidence coefficient for each property to decide whether or not to merge said instances of individuals; and means for updating the knowledge base with instances of individuals resulting from the information fusion.
- FIG. 1 schematically illustrates an information processing method implementing the present invention
- FIG. 2 schematically illustrates an example of a hardware arrangement of an information processing system in which the present invention can be implemented
- FIG. 3 schematically illustrates an example of the hardware arrangement of a control unit used in the information processing system
- FIG. 4A schematically illustrates a first example of a model of the evolution over time of a coefficient of confidence of a property of an instance of an individual
- FIG. 4B schematically illustrates a second example of a model of the evolution over time of a coefficient of confidence of a property of an individual instance
- FIG. 5 schematically illustrates a mechanism for calculating the distance of similarity between two instances of individuals, in a particular embodiment.
- Fig. 1 schematically illustrates an information processing method implementing the present invention. The method is implemented by an information processing system, an example of a hardware arrangement of which is detailed below in relation to FIG. 2.
- a step S 101 the information processing system collects information.
- Data is collected from multiple sources and the information collected comes from sources of various types and capacities.
- Each information collected is either of the hard information type or of the flexible information type.
- Multi-source collection involves collecting information from sources relevant to the targeted use case of information fusion.
- the hard information is obtained from sources such as physical sensors. This information is then structured, by the nature of the sensors which produce this information, in a raw data format. Soft information is linked to a human activity (social media, websites, official reports from a community or organization, etc.), are usually very large and unstructured. The extraction of flexible information is then based on a linguistic and semantic analysis of the content. Soft information is therefore considered as subjective, while hard information is considered objective.
- Open source intelligence platforms can also provide information resulting from one or more processing (translation, transcription, extraction, etc.) applied to pre-collected information, which makes it possible to derive so-called information from it.
- processing transformation, transcription, extraction, etc.
- 'individuals of interest eg, person, place, organization, event, equipment.
- the information collected can thus come from intelligence of human origin (designated under the term HUMINT, for “Human Intelligence” in English), from intelligence of open source origin (designated under the term OSINT, for “Open Source Intelligence” in English ) a maritime website, RSS (“Really Simple Syndication”) type flow syndication, an automatic identification system AIS (“Automatic Identification System”) for ships, databases maritime, radar information (designated under the term RADINT, for "Radar Intelligence” in English) with potentially different types of radar, information of electromagnetic origin (designated under the term SIGINT, for "Signal Intelligence” in English) such as radar activity detections of vessels or analysis of telephony signals mobile, and image source information (designated under the term IMINT, for “Image Intelligence”) such as images captured by satellites or drones.
- HUMINT Human Intelligence
- OSINT Open Source Intelligence
- RSS Resource Simple Syndication
- AIS Automatic Identification System
- Collection therefore makes it possible to obtain a set of hard and / or flexible information that concerns individuals.
- Information about these individuals is extracted from data available from various sources.
- the extraction can be done at the level of the source itself, so that the information processing system obtains in step S 101 information already “digested” (eg, recognition of a shape of a vessel in a video image sequence).
- the extraction can, as a variant, be done at the level of the information processing system, which then receives raw data from the source in question to be digested.
- a step S 102 the information processing system performs an ontology matching ("onthology matching" in English).
- Ontology is a representation of the information of a system that defines the types of individuals of this system with their categories, properties and relationships between these individuals for a specific operational use case (maritime surveillance, for example).
- the ontology thus makes it possible to have the same representation of information which is compatible with both hard and soft sources.
- Any individual identified and extracted at the end of the information collection is instantiated, to then feed relevant information into a situation monitoring system.
- any property linked to this individual and extracted from the corresponding collected information is instantiated.
- a property is either a literal (also called an "attribute"), such as for example the length of a ship, or a relation of an individual with another individual, such as for example the relation between a ship and its captain.
- a literal also called an "attribute”
- the property in question is not instantiated.
- an individual extracted from collected information can be totally or partially instantiated.
- an ontology can define an individual of type "ship", with several properties (eg, name of the ship, owner, date of observation, size, position, speed, IMO number (" International Maritime Organization number ”in English) ).
- a first source eg, AIS automatic identification system
- an instance (also referred to as an object) of an individual representing this vessel can be created with a literal instance for IMO number, observation date, position and speed, but not for the name of the vessel, the owner and the size, which are not part of the information contained in the messages of the automatic identification systems AIS.
- an individual instance representing that vessel with a literal instance for IMO number, vessel name and the shipowner can be created from information from this other source of information, but without an instance of a literal for speed, position and date of observation.
- an individual instance does not include an instance of one or more particular literals can already be information in itself. .
- Ontology alignment therefore consists of a total or partial instantiation of all individuals, with their properties and relationships, detected in the information collected, by inheriting the definitions provided by the ontology considered.
- the information collected can already be assigned, at the time of collection, to an ontology or not.
- the information processing system can also use an existing ontology with the information collected, or use its own ontology adapted to the use case (e.g., maritime surveillance).
- a transcription of the ontology provided by said information source into an ontology adapted to the use case e.g., maritime surveillance
- the instantiation of detected individuals relies directly on the ontology appropriate to the use case.
- the ontology adapted to the use case comprises parameters necessary for the establishment of evolution models in association with the instantiated properties.
- an appropriate ontology To apply the appropriate evolution model to each instantiated property, an appropriate ontology must be used. This comes from an expertise making it possible to determine which model describes the evolution over time of each defined property and its variability, and in in particular, to correctly parameterize the evolution model accordingly (eg, time factor t as presented below). The more a property is subject to variations over time, the less reliable this property is considered in information fusion.
- Each property is then associated with: a value; to an evolution model accompanied by one or more configuration parameters of said evolution model; preferably, a piece of information on the reliability of the information source that allowed the instantiation of the property in question; and information representative of an observation instant (ie, the moment when the value of the property was obtained by the information source).
- a classical ontology describes a property only by its value and its observation time, as well as possibly by the reliability of the information source. But here, each property is completed by an evolution model which represents the evolution of the reliability of said property over time in relation to the variability over time of said property.
- the term “reliability” is understood to mean the degree of confidence that the information processing system may have in a property value to decide whether or not to merge instances of individuals, in view of its variability over the period between the instants of. captures information from which said instances of individuals are extracted.
- a step S103 the information processing system performs an update of a knowledge base KB 205.
- knowledge bases are distinguished from simple databases. An explanation is given in the document “Knowledge Base Support for Decision Making Using Fusion Techniques in a C2 Environment”, Amanda Vizedom et al, Proceedings of the 4th International Conference on Information Fusion, International Society of Information Fusion, 2001, where he is indicated that the distinction between knowledge bases and databases is based on the distinction between general knowledge and specific data.
- a knowledge base is optimized for storing general, potentially complex knowledge of the type that can be instantiated.
- a database usually does not have the means to represent general principles, but is optimized to store very specific data, such as lists of elements and attributes.
- step S102 The instances of individuals during the ontology alignment in step S102 are therefore stored in the knowledge base KB 205 structured according to the ontology used to describe the individuals instantiated from the various information collected in step S 101 (with the necessary parameters for setting up evolution models).
- a step S 104 the information processing system performs an information merging operation.
- Information fusion is based on calculations of similarity distance between instances of individuals, and more precisely of similarity distances between properties of these instances of individuals.
- the similarity distance between two instances of individuals is a metric defining to what extent the instantiated individuals are similar or different, and even defining to what extent it is possible to decide whether these individuals are similar or different.
- the information fusion operation performed here takes into account evolution models, associated with each possible property of individuals according to the ontology applied in step S102. These evolution models make it possible to take into account the temporal dimension of the properties of individuals and their respective variabilities in the information fusion operation.
- step S104 mainly comprises two sub-steps: a sub-step S 1041 where similarity distance calculations are performed by applying the evolution models, for each property of each instance of an individual to be considered; and a data association sub-step S 1042, where the instances of individuals corresponding to the same individuals are associated, or according to the terminology applicable in the field, merged.
- This weighting corresponds to the uncertainty inherent in said property with respect to its collection method and to an evolution model corresponding to the estimated evolution over time of the variability of said property.
- the resulting weighting should express the fact that the more uncertain a property, the less impact it should have on similarity distance calculations, since information merging cannot rely on this property to decide whether two instances of individuals considered correspond or not to the same individual. For example, in the field of maritime surveillance, if we compare the position of a ship observed ten minutes ago to another position of a ship observed 4 days ago, it is not possible to know whether these two ships are one and the same or not, because in 4 days, the possibilities of changing the position of a ship are too vast for this to be a reliable criterion for comparison. Conversely, as the length of a vessel does not change, comparing a vessel length observation from a year ago with an observation from a day ago is reliable in trying to determine if it is the same ship or not.
- each property of an individual does not necessarily evolve in the same way as another property of that individual.
- the length of a ship is not likely to change, while its position is.
- Separate evolution models therefore represent these differences in the evolution of properties over time and therefore of the confidence to be given to these properties for the fusion of information as a function of the times of observation of the property in question.
- g r represents a confidence coefficient defined as follows: where l r is an optional coefficient representative of the reliability of the information source that made it possible to obtain the instance of the property p considered and m r is the evolution model applicable to the property p considered.
- l r is preferably equal to 1 - e s , where e s is the error rate of the information source.
- l r is preferably equal to the F-measure, also called F-score.
- a weight (or score) equal to "1" is considered a very reliable property to perform a similarity distance calculation and, conversely, a confidence coefficient (or weight or score) of zero means the property is too uncertain to be taken. taken into account in the calculation of similarity distance.
- the models of evolution are preferably of three possible types: constant; predictive; and circumstantial.
- the constant evolution model is associated with p properties which do not change over time, such as the length of a ship.
- a representation of a particular embodiment is provided in FIG. 4A, where it appears that the confidence coefficient g r is equal to the coefficient l r (m r being here equal to “1”).
- the predictive evolution model evolves over time and is therefore associated with p properties which evolve over time.
- properties p which correspond to the predictive evolution model are, for example, the speed of a ship, its position and its direction of navigation.
- the values of these properties p can be estimated (ie, predicted) over a certain period of time (over a limited period of time, beyond which the variability of the property p considered is such that its reliability is zero) or with a certain uncertainty that evolves over time. For example, knowing the position of a ship and the direction of its movement, it is easy to predict the area the ship will be in in the near future (eg, a few minutes later).
- the evolution is predictable, in particular thanks to mathematical tools.
- Kalman filters or particulate filters are preferred examples.
- predictive evolution models incorporate a notion of a confidence coefficient, often in the form of a covariance matrix.
- it is the comparison of the properties according to the predictive evolution model which directly integrates not only a predicted value but also the possible error on the prediction. This is the case, for example, with the Mahalanobis distance.
- circumstantial evolution model is associated with p properties, the evolution of which over time depends on the occurrence of an event.
- p properties the evolution of which over time depends on the occurrence of an event.
- the p properties associated with the circumstantial evolution model are therefore subject to modification following a specific unforeseeable event.
- circumstantial properties are the identity of the master or the flag of a vessel, which may change when the vessel in question changes owners.
- Another example is the location of the vessel, which can change a lot over time. Localization is here to be distinguished from position. Position is a set of geographic coordinates, while a vessel's location is the name of the place (e.g., Mediterranean Sea) where the vessel is located.
- the difficulty in circumstantial evolution models is to define the probability of such an event occurring and to find an adequate way to represent it. While other models could be used, exponential decay models appear to be a suitable approach.
- the similarity distance DS ⁇ l j , / fc ) between two instances of individuals I j and I k is then an average sum of the weighted similarity distances of each property p common to the two instances of individuals I j and I k and can then be calculated in the sub-step S 1041 as follows:
- a similarity distance calculation of a textual property can be obtained using the Levenshtein distance (also called "edit distance"), which is a metric for measuring the difference between two sequences of text.
- Levenshtein distance represents the minimum number of character change operations to be carried out in order to transform a first word, or a first sequence of words, to correspond to a second word, or respectively a second sequence of words .
- the Hamming distance (which is an upper bound of the Levenshtein distance) is used. The Hamming distance makes it possible to quantify the differences between two sequences of symbols or characters of the same length.
- Other digital calculations of similarity distances can be used to compare, for example, two speeds or two values of any other physical property.
- Normalization aims to ensure that the results of similarity distance calculations can then be used and compared together despite their heterogeneity and despite being based on different distance calculations.
- the purpose of normalization is to allow the result to be bounded by a distance, usually between 0 and 1. Typically, the results of distance calculations are close to 0 when there is no difference. For example, to normalize the Levenshtein or Hamming distance, it suffices to divide the result of the similarity distance calculation by the sum of the character length of the first sequence and the length of the second sequence
- the normalization can be transposed between -1 and 1.
- the normalization is then made between 0 and 1, then the result of this normalization is subtracted from 1.
- 1 represents the similarity
- -1 represents the dissimilarity.
- This similarity distance calculation by property p common to the instances of individuals considered can be aggregated with other similarity distance calculations, as detailed below in relation to FIG. 5, in order to obtain an aggregated similarity distance which is then used to decide whether or not to merge the instances of individuals I j and I k .
- substep S 1042 the information processing system performs a data association operation from the similarity distances calculated in substep S 1041.
- Data association is a heuristic for deciding whether two instances of individuals must be merged or not, given the similarity distance value (score) between these two instances of individuals.
- the instances of individuals following the collection of information and at least a subset of those already present in the KB 205 knowledge base are analyzed in pairs to determine if they correspond to the same individual and if they must therefore be merged.
- step S 104 therefore consists, as far as possible, of merging instances of individuals who represent the same individual.
- the individual instance which results from the merger of two original individual instances retains only one value for each property among those available in said original individual instances. The retained value depends on the evolution model with which the considered property is associated.
- the conserved value is that described by the source (eg, sensor) of the information from which is extracted the individual instance considered which has the best precision (which is known to the fact that the ontology has the information on the accuracy of the source which observed the property).
- index 72 ⁇ 2 where the index "1" represents the oldest information and the index "2" represents the most recent information, where l is the optional coefficient representative of the reliability of the source that performed the capture (or observation) of the information considered, t is the time factor of the predictive evolution model as defined above, and / represents the instant of capture (or observation) of the information considered.
- a step S105 the information processing system performs a new update of the knowledge base KB 205.
- each new individual instance resulting from the information merging is stored in the KB 205 knowledge base. Since the similarity distance was sufficiently small to allow the association of data between at least one pair of instances of individuals, the instances of individuals (and therefore their properties ) can be merged to generate an "augmented" instance for this individual. This new instance can then in turn be associated with one or more other instances during a new iteration of the information fusion operation.
- the instances of individuals which have allowed the fusion of information and the instance of individuals generated by the fusion of information are therefore all kept in the knowledge base KB 205 and are linked to each other therein. As a variant, the instances of individuals that were used to create a merged individual instance are not kept in KB 205 knowledge base.
- a situational awareness system uses the results obtained during the information merging operations carried out in step S 105 and represents these results in the form of synthetic views, in order to facilitate the detection of abnormal behavior.
- Such situational awareness systems are well known in the field of maritime surveillance and / or civil security, and are generally operated by regional, national or international organizations responsible for monitoring a given geographical area.
- the situation monitoring system is integrated, or connected, to the information processing system.
- Such situational awareness systems implement sets of predefined rules exploiting the results obtained in step S 105 to detect individuals (ship, etc.) with abnormal behavior compared to a behavior defined as standard in view of the type. of the individual considered, and to generate an alert if necessary, which is for example displayed to the operator.
- Such rule-based mechanisms are well known in the literature through expert systems.
- situational ontology models are used to characterize types of behavior.
- situation ontology is described in the document "Improving Maritime Situational Awareness by Fusing Sensor Information and Intelligence", van den Broek et al., International Conference on Information Fusion, 2011 ..
- Such situational awareness systems generally include one or more common operational views (or "Common Operational Picture, COP" in English) made up of synthetic graphical or / and tabular views presenting the results of the information fusion with those obtained by d other biases.
- the situational awareness system comprises, in a graphical interface, a geographical view of the monitored area with a background map or an aerial image or both superimposed. Vessels in the monitored area are superimposed in the geographic view by an icon and a label giving the vessel's identification information. A displacement vector, or a trajectory, can also be presented for each vessel on the geographic view.
- the situation monitoring system can also include a tabular or graphical view presenting the alerts generated following the exploitation of the results of the information fusion. These alerts can be presented to a human operator according to a color code according to the severity and / or the urgency of the situation, potentially accompanied by a visual and / or audible warning signal.
- One of the advantages obtained by using the results of the fusion of information resulting from the method of the invention in a situational awareness system is therefore to offer a correlation space between information much larger than that which a human operator is able to apprehend manually, that is to say by his only cognitive capacities with or without the help of the methods of fusion of information of the state of the art, this in order to eliminate the duplicates before display and offer improved and more automated situational awareness.
- This allows the human operator to focus on situational interpretation and situational decision making, rather than residual and manual correlation operations.
- the graphical interface also presents means of representing the history of information mergers carried out automatically at the during the implementation of the method and saved as and when in the knowledge base KB 205.
- step S106 uses the results of the fusion of information as described in step S106 to the examples of situation management and to the examples of modes of representation mentioned above.
- Fig. 2 schematically illustrates an example of a hardware arrangement of an information processing system in which the present invention can be implemented.
- the information processing system is for example a maritime surveillance system MSS (“Maritime Surveillance System” in English) 250.
- MSS Maritime Surveillance System
- the information collected concerns any vessel present at sea in an area. predefined geographic area (eg, all seas and oceans around the world). Sources have recovered partial or redundant information on ships. This information must be correlated so that it can be completed and merged in order to better understand the behavior of all these ships.
- the result of the information fusion is a descriptive list of vessels containing more complete and non-redundant information, which allows efficient work on the information retrieved, which is impossible without precise correlation of the information collected.
- Evolution models provide this precision by taking into account the temporal evolution of the properties of the instances of individuals following the collection of information and more particularly the variability of these properties over time.
- the units (or modules) shown in the example arrangement of Fig. 2 achieve this result.
- the information processing system comprises a DC (“Data Collector”) collection unit 201, in charge of recovering information from a set 200 of various information sources SI, S2, S3, S4, independently. whether the sources in question provide hard or soft information.
- the collection unit DC 201 has the behavior already described in relation to step S 101.
- the DC collection unit 201 can also include direct access to existing databases containing hard and / or flexible information which comes from various sources and which has been previously collected by another means.
- the information processing system is capable of interconnecting with a distributed database system originating from distinct actors and authorities.
- the information processing system further comprises an OM (“Ontology Matching”) ontology alignment unit 202, which has the behavior already described in relation to step S 102.
- the information processing system further comprises an input-output unit KIO (“Knowledge Input / Output” in English) in charge of ensuring the access, in input and output, of the knowledge base KB 205.
- KIO Knowledge Input / Output
- the input-output unit KIO 203 provides access to the knowledge base KB 205.
- the information processing system further comprises an information fusion unit IF ("Information Fusion" in English) 204, which has the behavior already described in relation to step S 104.
- information fusion unit IF Information Fusion
- the information processing system preferably further comprises a situation monitoring system.
- the situation monitoring system then comprises a trigger unit TRIGG (“Trigger” in English) 207 and a graphical user interface GUI (“Graphical User Interface” in English) 208.
- the trigger unit TRIGG 207 is in charge of lifting alerts on abnormal behavior detected as a result of data fusion.
- the GUI 208 graphical interface is configured to graphically represent alerts on abnormal behavior detected as a result of information merging, as well as individuals related to these alerts.
- the information processing system further comprises a CTRL control unit 206 in charge of coordinating, for example by means of a data bus 310, the various units of the information processing system, so as to implement the behavior already described. in relation to FIG. 1.
- each of the DC 201 collection units, OM 202 ontology alignment, KIO input / output 203 and IF information fusion units 204 can be implemented in hardware form, for example using an electronic component (“ chip ”) or a set of electronic components (“ chipset ”in English); or else be produced in software form and implemented by a processor executing the corresponding computer program instructions.
- chip electronic component
- chipset set of electronic components
- Fig. 3 schematically illustrates an example of a hardware arrangement of the control unit CTRL 206 of the information processing system.
- the example of the hardware architecture presented comprises, connected by a communication bus 310: a processor CPU 301; a random access memory RAM (“Random Access Memory” in English) 302; a ROM (“Read Only Memory”) 303 or a Flash memory; a storage unit or a storage medium drive, such as an SD ("Secure Digital”) card reader or an HDD (“Hard Disk Drive”) 304; and at least one 305 I / O interface.
- a communication bus 310 a processor CPU 301; a random access memory RAM (“Random Access Memory” in English) 302; a ROM (“Read Only Memory”) 303 or a Flash memory; a storage unit or a storage medium drive, such as an SD (“Secure Digital”) card reader or an HDD (“Hard Disk Drive”) 304; and at least one 305 I / O interface.
- CPU 301 is capable of executing instructions loaded into RAM 302 from ROM 303, external memory (such as an SD card), storage media (such as disk hard HDD), or a communication network. Upon power-up, the CPU 301 is able to read instructions from RAM 302 and execute them. These instructions form a computer program causing the CPU 301 to implement some or all of the algorithms and steps described here.
- all or part of the algorithms and steps described here can be implemented in software form by executing a set of instructions by a programmable machine, such as a DSP (“Digital Signal Processor”) or a microcontroller or a processor. All or part of the algorithms and steps described here can also be implemented in hardware form by a machine or a dedicated component, such as an FPGA (“Field-Programmable Gâte Array”) or an AS IC (“Application-Specific Integrated Circuit ”in English).
- the information processing system comprises electronic circuitry adapted and configured to implement the algorithms and steps described here.
- Fig. 5 schematically illustrates a mechanism for calculating the distance of similarity between two instances of individuals, in a particular embodiment in which a calculation of distance of similarity based on the evolution models is aggregated with at least one other calculation of distance of similarity.
- the instances of individuals are compared in pairs, eg, instances of individuals 01 and 02 are injected as input (I) of the similarity distance calculation.
- a first similarity distance is calculated using a taxonomic similarity distance calculation module TS (“Taxonomy Similarity”) 501.
- the instances of individuals 01 and 02 are instances of class in the ontology considered.
- the taxonomic similarity distance calculation compares the positions of the classes of instances of individuals 01 and 02.
- the classes and properties are hierarchical and this hierarchy can be represented by a graph.
- a class (node) "Submarine” and a class (node) "Boat” both inherit from a class (node) "Boat” which itself inherits from a class (node) "Vehicle” , and from the “Vehicle” class (node) also inherit from the “Aircraft” and “Land Vehicles” classes (nodes), and so on.
- a distance between two graph nodes can be calculated by counting the number of edges of the shortest path between the nodes considered in the graph.
- the taxonomic similarity measure also takes into account another criterion to represent depth in the ontological hierarchy.
- the taxonomic similarity distance TS (01; 02) is here defined from the distance which separates the two classes Cl and C2 of the instances of individuals 01 and 02 from the root R of the hierarchy and from the distance which separates their lowest common sub-denominator CO with respect to the root R of the hierarchy, according to the following formula: where d (R; CO) is the distance which separates the class CO from the root R of the hierarchy, d (R; CO; Cl) is the distance which separates the class Cl from the root R passing through the class CO and d (R; CO; C 2) is the distance separating class C2 from the root R passing through class CO.
- d (R; CO; C 2 is the distance separating class C2 from the root R passing through class CO.
- a second similarity distance is calculated using a domain and range similarity distance calculation module DRS ("Domain and Range Similarity" in English) 502.
- DRS Domain and Range Similarity
- the calculation of the domain similarity distance and DRS range compares the number of fields (properties) shared by the two classes C1 and C2 to which the two instances of individuals C1 and 02 belong respectively, normalized by their total number of fields.
- Ontology is in fact preferentially not limited to the hierarchical structure of concepts in the form of classes, but also includes domain and range definitions within the properties, as shown by the following system of equations.
- the calculation of distance of similarity between classes involves the comparison of properties which appear in common in the considered instances of these classes.
- OPR (C) represents the set of relation-type properties that have class C in the range definition of a second subject, and ⁇ OPR (C) ⁇ represents the cardinality of that set
- DPD ⁇ C) represents the set of literal type properties that have class C in their range definition, and
- a third similarity distance is calculated using a similarity distance calculation module based on the evolution models MoES (“Model of Evolution-based Similarity”) 503.
- the similarity distance based on models of evolution MoES between instances of individuals 01 and 02 is an average sum of the weighted similarity distances of each property p common to the two instances of individuals 01 and 02, as follows:
- the first, second and third similarity distances are then combined by an aggregator module AGG 504, in order to produce at the output (O) of the calculation of the similarity distance a similarity distance SD ("Similarity Distance" in English) between instances of individuals 01 and 02.
- the aggregator module AGG 504 applies respective weights to the first, second and third similarity distances, in order to give more or less importance to each of them and to standardize the result.
- the weights respectively assigned to the first, second and third similarity distances are defined as a function of the application framework considered. Ontology can thus, for example, give greater weight to the taxonomic similarity distance TS compared to the similarity distance based on the MoES evolution models and to the domain and range similarity distance DRS.
- the mechanism for calculating the distance of similarity between two instances of individuals has been presented in Fig. 5 in modular form.
- the modules in question can be hardware modules or software modules.
- the similarity distance calculation mechanism shown in FIG. 5 is also representative of a method including steps of calculating the first, second and third similarity distances, and the corresponding aggregation, as described above.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Animal Behavior & Ethology (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR1906376A FR3097346B1 (fr) | 2019-06-14 | 2019-06-14 | ProcéDé et système de fusion d’informations |
| PCT/EP2020/066282 WO2020249719A1 (fr) | 2019-06-14 | 2020-06-12 | Procede et systeme de fusion d'informations |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4165519A1 true EP4165519A1 (fr) | 2023-04-19 |
Family
ID=68581875
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP20731485.7A Pending EP4165519A1 (fr) | 2019-06-14 | 2020-06-12 | Procede et systeme de fusion d'informations |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US12353468B2 (fr) |
| EP (1) | EP4165519A1 (fr) |
| FR (1) | FR3097346B1 (fr) |
| WO (1) | WO2020249719A1 (fr) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114445776B (zh) * | 2022-01-18 | 2025-05-30 | 武汉理工大学 | 面向内河海事视频监控的船舶名称主动式识别方法 |
| KR20250165029A (ko) * | 2024-05-17 | 2025-11-25 | 주식회사 Lg 경영개발원 | 복수의 모델을 포함하는 인공지능 모델 기반의 모델 특정 방법 및 그 시스템 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7756800B2 (en) * | 2006-12-14 | 2010-07-13 | Xerox Corporation | Method for transforming data elements within a classification system based in part on input from a human annotator/expert |
| US7783586B2 (en) * | 2007-02-26 | 2010-08-24 | International Business Machines Corporation | System and method for deriving a hierarchical event based database optimized for analysis of biological systems |
| US7958155B2 (en) * | 2007-04-17 | 2011-06-07 | Semandex Networks, Inc. | Systems and methods for the management of information to enable the rapid dissemination of actionable information |
| US8244769B2 (en) * | 2007-05-31 | 2012-08-14 | Nec Corporation | System and method for judging properties of an ontology and updating same |
| US20120078595A1 (en) * | 2010-09-24 | 2012-03-29 | Nokia Corporation | Method and apparatus for ontology matching |
| US8856156B1 (en) * | 2011-10-07 | 2014-10-07 | Cerner Innovation, Inc. | Ontology mapper |
| US10019516B2 (en) * | 2014-04-04 | 2018-07-10 | University Of Southern California | System and method for fuzzy ontology matching and search across ontologies |
| AU2015258752A1 (en) * | 2014-05-12 | 2017-01-05 | Semantic Technologies Pty Ltd | Putative ontology generating method and apparatus |
| US10509814B2 (en) * | 2014-12-19 | 2019-12-17 | Universidad Nacional De Educacion A Distancia (Uned) | System and method for the indexing and retrieval of semantically annotated data using an ontology-based information retrieval model |
| US10824662B2 (en) * | 2015-10-13 | 2020-11-03 | Nuance Communications, Inc. | Methods and system for iteratively aligning data sources |
| US10302769B2 (en) * | 2017-01-17 | 2019-05-28 | Harris Corporation | System for monitoring marine vessels using fractal processing of aerial imagery and related methods |
| US11397851B2 (en) * | 2018-04-13 | 2022-07-26 | International Business Machines Corporation | Classifying text to determine a goal type used to select machine learning algorithm outcomes |
-
2019
- 2019-06-14 FR FR1906376A patent/FR3097346B1/fr active Active
-
2020
- 2020-06-12 WO PCT/EP2020/066282 patent/WO2020249719A1/fr not_active Ceased
- 2020-06-12 EP EP20731485.7A patent/EP4165519A1/fr active Pending
- 2020-06-12 US US17/619,087 patent/US12353468B2/en active Active
Also Published As
| Publication number | Publication date |
|---|---|
| FR3097346B1 (fr) | 2021-06-25 |
| US20220374464A1 (en) | 2022-11-24 |
| FR3097346A1 (fr) | 2020-12-18 |
| WO2020249719A1 (fr) | 2020-12-17 |
| US12353468B2 (en) | 2025-07-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11494648B2 (en) | Method and system for detecting fake news based on multi-task learning model | |
| EP2370938A2 (fr) | Procede et systeme pour la fusion de donnees ou d'information | |
| WO2025188382A2 (fr) | Atténuation de menaces pour la sécurité | |
| US20240223607A1 (en) | Systems and methods for intelligent identification and automated disposal of non-malicious electronic communications | |
| US20210303937A1 (en) | Ensemble weak support vector machines | |
| Hakak et al. | Propagation of fake news on social media: challenges and opportunities | |
| CN111160959A (zh) | 一种用户点击转化预估方法及装置 | |
| Demertzis et al. | A machine hearing framework for real-time streaming analytics using Lambda architecture | |
| Han et al. | A comprehensive framework incorporating deep learning for analyzing fishing vessel activity using automatic identification system data | |
| EP4165519A1 (fr) | Procede et systeme de fusion d'informations | |
| CA2370693C (fr) | Systeme et methode de pilotage d'un processus decisionnel lors de la poursuite d'un but globale dans un domaine d'application determine | |
| Martin et al. | Embracing firefly flash pattern variability with data-driven species classification | |
| CA2940380A1 (fr) | Determiner la severite d'une perturbation geomagnetique sur un reseau electrique a l'aide de mesures de similarite | |
| CN115035347B (zh) | 图片识别方法、装置及电子设备 | |
| FR2929426A1 (fr) | Procede et systeme d'attribution de score | |
| Saaya et al. | The Development of Trust Matrix for Recognizing Reliable Content in Social Media. | |
| Rachel et al. | A novel DLDRM: Deep learning–based flood disaster risk management framework by multimodal social media data | |
| Athif et al. | Association Rule Mining for Identifying High-Risk Drug Combinations in Overdose Fatalities: A Comparative Analysis of Apriori and FP-Growth Algorithms | |
| Yuan et al. | RoadFed: A Multimodal Federated Learning System for Improving Road Safety | |
| US12505650B1 (en) | Method, program, and apparatus for automated analysis of criminal evidence | |
| Swapnika et al. | Event detection and classification using deep compressed convolutional neural network | |
| CN119854102B (zh) | 网络训练及确定故障的方法、装置、设备、介质和产品 | |
| Lau Kheng Seng et al. | Oversampling Social Media-Sourced Image Datasets for Better Deep Learning Classification of Natural Disaster Damage Levels. | |
| Ahli et al. | A Survey on the Development of Machine Learning and Artificial Intelligence Techniques in Digital Forensics | |
| Patel et al. | Comparative Analysis of Fake News Identification Using Machine Learning Methods |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20211210 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20241120 |