EP4674087A1 - Training a machine learning model - Google Patents
Training a machine learning modelInfo
- Publication number
- EP4674087A1 EP4674087A1 EP23710486.4A EP23710486A EP4674087A1 EP 4674087 A1 EP4674087 A1 EP 4674087A1 EP 23710486 A EP23710486 A EP 23710486A EP 4674087 A1 EP4674087 A1 EP 4674087A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- vulnerability
- vulnerabilities
- occurring
- subset
- machine learning
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L63/00—Network architectures or network communication protocols for network security
- H04L63/20—Network architectures or network communication protocols for network security for managing network security; network security policies in general
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/50—Monitoring users, programs or devices to maintain the integrity of platforms, e.g. of processors, firmware or operating systems
- G06F21/57—Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
- G06F21/577—Assessing vulnerabilities and evaluating computer system security
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/0895—Weakly supervised learning, e.g. semi-supervised or self-supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/01—Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
Definitions
- This disclosure relates to a method, apparatus, a computer program, a carrier and a computer program product. More particularly but non-exclusively, the disclosure relates to training a machine learning model to predict vulnerabilities in a communications network.
- the disclosure herein relates to security in computer networks, hosts and servers.
- Current processes used for vulnerability detection, scanning, analysis, and patching solutions are largely based around the following activities:
- Vulnerabilities are often given a severity rating based on how easy the weakness is to exploit, and the impact exploiting it can have on the program, server or data.
- the prevailing philosophy in existing systems is to develop a process and plan to continuously assess, track and mitigate vulnerabilities and security weaknesses in servers within a system infrastructure using vulnerability scanning tools in order to minimize exposure to exploitable vulnerabilities in the system.
- each identified vulnerability is classified individually based on its CVSS value which attempts to assign severity scores to single vulnerabilities, allowing responders to prioritize responses and resources according to threat.
- CVSS value attempts to assign severity scores to single vulnerabilities, allowing responders to prioritize responses and resources according to threat.
- it is often complex to chain vulnerabilities to identify and evaluate the impact of complex threats which may be comprised of several simultaneous vulnerabilities.
- security patches are generally created on an individual basis resulting in a time consuming and potentially troublesome patching process.
- a method performed by a node in a communications network comprises: i) obtaining vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) matching a first vulnerability in the plurality of vulnerabilities to a first subset of cooccurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) training a machine learning model using the first vulnerability as an example input and the first subset of cooccurring vulnerabilities as an example output of the machine learning model.
- a method performed by a node in a communication system comprising: detecting a second vulnerability in a host server, and using a machine learning model to predict a second subset of co-occurring vulnerabilities that are predicted to occur at an overlapping time with the first vulnerability, wherein the machine learning model was trained using training data comprising example vulnerabilities and corresponding example subsets of co-occurring vulnerabilities.
- a node in a communications network comprising: a memory comprising instruction data representing a set of instructions; and a processor configured to communicate with the memory and to execute the set of instructions, wherein the set of instructions, when executed by the processor, cause the processor to: i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
- a node in a communications network configured to: i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
- a fifth aspect there is a computer program comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out a method according to the first or second aspects.
- a carrier containing a computer program according to the fifth aspect wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.
- a seventh aspect there is a computer program product comprising non transitory computer readable media having stored thereon a computer program according to the fifth aspect.
- a method, apparatus, a computer program, a carrier and a computer program product that introduce co-occurring vulnerabilities as metrics to enable deeper vulnerability analysis beyond traditional risk-based ranking.
- vulnerabilities output from a vulnerability scan can be linked together in an automated manner using a data-driven and machine learning approach. This reduces the burden on human engineers manually reviewing large scan reports. It further allows for improved patching, since patches may be better targeted at groups of co-occurring vulnerabilities, rather than merely individual vulnerabilities with the highest CVSS values. Linking vulnerabilities in this manner thus provides data of increased actionable value which can be used to increase the efficiency of security hardening processes.
- Fig. 1 illustrates an example node according to embodiments herein;
- Fig. 2 shows an example method according to embodiments herein;
- Fig. 3 shows an example method according to embodiments herein;
- Fig. 4 shows an example method according to embodiments herein;
- Fig. 5 shows an example method according to embodiments herein;
- Fig. 6 shows an example system according to some embodiments herein;
- Fig. 7 shows an example carrier containing a computer program;
- Fig. 8 shows an example computer program product.
- a communications network or telecommunications network may comprise any one, or any combination of: a wired link (e.g. ASDL) or a wireless link such as Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), New Radio (NR), WiFi, Bluetooth or future wireless technologies.
- GSM Global System for Mobile Communications
- WCDMA Wideband Code Division Multiple Access
- LTE Long Term Evolution
- NR New Radio
- WiFi Bluetooth
- Bluetooth Bluetooth
- a wireless network may implement communication standards, such as Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), and/or other suitable 2G, 3G, 4G, 5G, 6G or future generation of standards; wireless local area network (WLAN) standards, such as the IEEE 802.11 standards; and/or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave and/or ZigBee standards.
- GSM Global System for Mobile Communications
- UMTS Universal Mobile Telecommunications System
- LTE Long Term Evolution
- WLAN wireless local area network
- WiMax Worldwide Interoperability for Microwave Access
- Bluetooth Z-Wave and/or ZigBee standards.
- Fig. 1 shows an example node 100 in a communications network 107 according to some embodiments herein.
- the node 100 may be otherwise referred to herein as a network node.
- the node 100 is configured (e.g. adapted, operative, or programmed) to perform any of the embodiments of the method 200 or 400 as described below.
- the node 100 comprises a processor (e.g. processing circuitry or logic) 102.
- the processor 102 may control the operation of the node 100 in the manner described herein.
- the processor 102 can comprise one or more processors, processing units, multi- core processors or modules that are configured or programmed to control the node 100 in the manner described herein.
- the processor 102 can comprise a plurality of computer programs and/or hardware modules that are each configured to perform, or are for performing, individual or multiple steps of the functionality of the node 100 as described herein.
- the node 100 comprises a memory 104.
- the memory 104 of the node 100 can be configured to store a computer program 106 with program code or instructions that can be executed by the processor 102 of the node 100 to perform the functionality described herein.
- the memory 104 can be configured to store any requests, resources, information, data, signals, or similar that are described herein.
- the processor 102 may be configured to control the memory 104 to store any requests, resources, information, data, signals, or similar that are described herein.
- the node 100 may comprise one or more virtual machines running different software and/or processes.
- the node 100 may therefore comprise one or more servers, switches and/or storage devices and/or may comprise cloud computing infrastructure or infrastructure configured to perform in a distributed manner, that runs the software and/or processes.
- the node 100 may comprise other components in addition or alternatively to those indicated in Fig. 1.
- the node 100 may comprise a communications interface.
- a communications interface may be for use in communicating with other nodes e.g. via a communications network.
- the communications interface may be configured to transmit to and/or receive from nodes or network functions requests, resources, information, data, signals, or similar.
- the processor 102 may be configured to control such a communications interface to make/receive such transmissions.
- the node 100 may be implemented in (e.g. form part of) a communications network 107.
- the node 100 may be implemented in a management layer of a communications network.
- the node 100 may comprise a node in a network management system, security management system or a vulnerability management platform.
- the node 100 may generally be any node, network function host, server, server farm, distributed computer system or network device in a communications network.
- the node 100 may comprise any component or network function (e.g. any hardware) in a communications network suitable for performing the functions described herein.
- Examples of the node 100 include, but are not limited to, core network function hosts such as, for example, hosts for core network functions in a Fifth Generation Core network (5GC). It is realized that the node 100 may equally be a node/device in any future network, such as a future 3GPP (3 rd Generation Partnership Project) sixth generation communication network, irrespective of whether the node 100 would there be placed in a core network or outside of the core network.
- 3GPP 3 rd Generation Partnership Project
- the node 100 may be positioned in a part of the communications network which is outside of the future 3GPP communication network, but connected to the future 3GPP communication network, e.g. to provide a cloud computing platform to offload the 3GPP communication network.
- the node 100 is for use in security patching of host servers in a communications network.
- current vulnerability scans tend to produce many hundreds, or thousands of lines of vulnerability data that are manually reviewed by human security experts who have to determine which vulnerabilities to patch and how to go about patching them.
- Embodiments herein aim to aid such human engineers by using machine learning to predict for an input vulnerability, a subset of co-occurring vulnerabilities that are predicted to occur at the same or overlapping time as the input vulnerability.
- vulnerabilities are linked, for example, they may arise or become exploitable when particular conditions are found in a network.
- a human engineer may be able to select a better security patch in order to patch an entire group of vulnerabilities in one go (rather than patching individual vulnerabilities one at a time in a more ad-hoc manner), thus improving and simplifying the security patching process.
- the node 100 is configured to i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
- a host server in this context is a server that is connected to the internet (e.g. via a communications network as described above).
- a host server may run services accessible via the internet.
- Such host servers may be scanned using vulnerability scanners (e.g. software) to detect security vulnerabilities.
- Vulnerability scanners are tools that monitor applications and networks to identify security vulnerabilities. Vulnerability scanners conduct scans to identify potential exploits or vulnerabilities by comparing e.g. configuration parameters, network traffic patterns or similar, to a database of known vulnerabilities, or by comparing scanning results against a database of known vulnerabilities and security weaknesses. Vulnerability scanners are used by companies to test applications and networks against known vulnerabilities and to identify new vulnerabilities.
- the output of a scan is a report which, as described above in the background, may contain thousands of lines of identified vulnerabilities at any given time. Each vulnerability may be graded according to severity. Vulnerabilities may be time dependent, and thus a scan report may comprise time information, for example, indicating a time at which a vulnerability was detected and/or a length of time in which the vulnerability was detected. A scan report may detail the state of an application or host server and/or provide recommendations to remedy known issues.
- Fig. 2 shows a method 200 in a node 100 in a communications network according to some embodiments herein.
- the method 200 is computer implemented in that it is performed by the node 100.
- the method comprises obtaining vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan.
- the method comprises matching a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability.
- the method comprises training a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
- vulnerability data is obtained from (e.g. output from) a vulnerability scan.
- the vulnerability data comprises a plurality of vulnerabilities, e.g. in a list.
- Each vulnerability may be time stamped.
- the time stamp may indicate the time at which the vulnerability was first detected.
- the time stamp may be in the form of a time window (e.g. a time at which each vulnerability was first detected and a time at which it ceased to be detected).
- a first vulnerability of the plurality of vulnerabilities is matched with a first subset of co-occurring vulnerabilities.
- the first subset of co-occurring vulnerabilities are one or more other vulnerabilities in the plurality of vulnerabilities that co-occurred with the first vulnerability.
- Co-occurred in this sense may mean overlapping in time, in any manner.
- the first vulnerability may occur in an overlapping, or partially overlapping time window with each co-occurring vulnerability in the plurality of co-occurring vulnerabilities.
- the first vulnerability may occur simultaneously, e.g. at the same time as each co-occurring vulnerability in the plurality of co-occurring vulnerabilities.
- the first vulnerability may overlap in time in a different manner to each co-occurring vulnerability, e.g. it may start and stop at (approximately) the same time as some of the first subset of cooccurring vulnerabilities, while merely overlapping in time with other vulnerabilities in the first subset of co-occurring vulnerabilities. So long as a vulnerability overlaps in time in some manner, then it may be matched to the first vulnerability and added to the first subset of cooccurring vulnerabilities.
- the output of step 204 is an indication of the first vulnerability and the first subset of cooccurring vulnerabilities.
- an indication is a name, label or any other reference that may be used to identify a vulnerability.
- the output may be in the form of a tuple:
- the output may be encoded, e.g. using a cipher to convert each vulnerability into a numerical or vectoral identifier. This is described in more detail below.
- Steps 202 and 204 may be repeated for each vulnerability in the plurality of vulnerabilities, e.g. each unique vulnerability (e.g. the steps may not be repeated for duplicated vulnerabilities in the plurality of vulnerabilities).
- an indication of a total number of co-occurrences of the co-occurring vulnerability with the first vulnerability may be obtained.
- the total number of co-occurrences of each (other) vulnerability with the first vulnerability may be obtained in step 204.
- step 31 An embodiment of steps 202 and 204 are illustrated in Fig 3 which illustrates an example data processing function.
- step 202 is performed and vulnerability data from a vulnerability scan is obtained.
- the vulnerability data (otherwise referred to as input data) is the output of one or more NessusTM scans, although any other vulnerability scan reports could equally be used.
- the following fields of a vulnerability scan may be obtained:
- CVSS Common Vulnerability Scoring System
- step 31 the input data is read and a list of unique vulnerabilities is determined (e.g. duplicated vulnerabilities are removed).
- step 32 a first vulnerability is selected from the plurality of vulnerabilities output from step 31.
- step 33 coupled vulnerabilities that occur in overlapping time windows with the first vulnerability, and within the same host are identified. For example, if vulnerability X occurs in host Y, then in step 33, the method may comprise counting simultaneous, or overlapping occurrences of vulnerabilities A, B, C,... This may be repeated for all hosts in a given data set.
- the coupled vulnerabilities may be cached and stored e.g. in the form:
- vulnerability X was found to co-occur with vulnerability A 7 time, and vulnerability B twice, while Vulnerability A was found to co-occur with vulnerability B once.
- step 35 the process is repeated for second and subsequent vulnerabilities in a looped manner until each unique vulnerability in the plurality of vulnerabilities has been assessed.
- the method may further be repeated for vulnerability data obtained from a plurality of different host servers.
- step 36 the cached coupled vulnerabilities are stored, e.g. in long term storage.
- the input data to the process illustrated in Fig. 3 is static (e.g. the input parameter types and/or structure is always the same), and may be found from vulnerability scan reports.
- the process in Fig. 3 produces coupled vulnerability information that depicts how many times a vulnerability has co-occurred at the same time as each other vulnerability. Once every vulnerability has been processed (e.g. once the loop in step 35 has been exhausted), then the coverage of each vulnerability may also be determined. Coverage indicates the number of occurrences of each vulnerability and may be expressed in terms of vulnerability X occurred x times, or in y percentage of host servers scanned.
- the method 200 comprises training a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
- the output from step 204 is used as (a piece of) training data with which to train a machine learning model
- the method 200 may be looped as described above, on each vulnerability in the plurality of vulnerabilities to build up a training data set of [vulnerability, co-occurring vulnerabilities] pairs.
- a machine learning model may be trained in this way to output, for an input vulnerability, one or more other vulnerabilities that are predicted to co-occur with an input vulnerability.
- a machine learning process may comprise a procedure that is run on data to create a model that may be referred to as a “machine learning model”.
- Machine learning process comprise procedures and/or instructions through which training data, may be processed or used in a training process to generate a machine learning model. Examples of machine learning processes include but are not limited to processes such as back-propagation and gradient descent. Machine learning processes learn from training data comprising example input and output pairs. The example outputs in a training data example are the ground truth or “correct” outputs for the respective example input.
- the machine learning model is multi-class classifier.
- a decision tree see paper by Quinlan (1986) entitled: “Induction of decision trees” Machine Learning volume 1 , pages 81-106 (1986)
- a random forest-based classifier see paper by Breiman (2001) entitled “Random Forests”', Mach Learn 45(1): 5-32).
- the first machine learning model is a deep neural network, see paper by Schmidhuber (2015) entitled: “Deep learning in neural networks: An overview” Neural Networks Volume 61, January 2015, Pages 85-117.
- the machine learning model (otherwise referred to herein as the model) can be any type of classification or regression model, trained using a supervised or semi-supervised learning process.
- the input to the model may be an indication of a vulnerability.
- An indication in sense is e.g. a name, label, or any other reference that can be used to identify a vulnerability.
- the below italic gives an example model input.
- the first row provides column names
- the second row depicts a first vulnerability that occurs within host xx.yyy.cf.d
- a Telnet server is listening on the remote port.
- the remote host is running a Telnet server, a remote terminal server
- Unencrypted Telnet Server Description: The remote Telnet server transmits traffic in cleartext
- SSH is preferred over Telnet since it protects credentials from eavesdropping and can tunnel additional data streams such as an X11 session. Disable the Telnet service and use SSH instead.
- the output of the machine learning model is an indication of a subset of co-occurring vulnerabilities that are predicted to co-occur with a respective input vulnerability.
- the training data may be in the form of [example input vulnerability, example subset of co-occurring vulnerabilities predicted for the example input vulnerability] pairs.
- the output of the machine learning model comprises an indication of a coverage of the input vulnerability.
- the output may indicate the number of comparable host servers on which the input vulnerability has occurred. E.g. the vulnerability has occurred on / percent of similar host servers.
- the training data may be in the form of [example input vulnerability, [example subset of cooccurring vulnerabilities predicted for the example input vulnerability, coverage]] pairs.
- the output may further comprise a risk score for each of the predicted co-occurring vulnerabilities.
- the inputs and outputs may be presented in a human-readable form e.g. each may be given a name or other reference.
- these names or references may be converted into a numerical/vectoral form which may be more readily processed or read by a machine.
- the first vulnerability and the first subset of co-occurring vulnerabilities may be transformed using a mapping into a unique numerical form.
- a cipher (or transformation cipher) may be used for the transformation.
- a cipher refers to a mapping function that can be used to encode and decode vulnerability names into numbers and vice-versa, e.g. “vulnerability X” ⁇ -> 1.
- a cipher may be a simple enumeration in which each unique vulnerability is assigned a unique number.
- a cipher may be a more complex transformation to convert each vulnerability in the plurality of vulnerabilities into a form more easily processed by a machine learning model, e.g. into a numerical or vector form.
- Fig. 5 shows a method of training a machine learning model and creation of a transformation cipher, according to some embodiments herein.
- the input data is read (e.g. obtained or received) and unique vulnerabilities are found in the input data.
- the input data is the coupled vulnerability data output from step 204 above.
- the input data may be in the following format: Vulnerability X : ⁇ Vulnerability A: 7 Vulnerability B: 2 ⁇ Vulnerability A: ⁇ Vulnerability B: 1 ⁇ •
- vulnerability description information that considers the following for each vulnerability may also be used: vulnerability:
- Training data may be created from this input.
- the numbers (of co-occurrences) are used to determine the number of samples in the model training.
- this sample may be added 7 times.
- the above input data example can be translated to a training set that considers Vulnerability X as the vulnerability of interest is composed from 7 samples of Vulnerability A, and 2 samples of Vulnerability B. In this way, the ratios/distribution of the coupled vulnerabilities are preserved because the machine learning model will be more likely to predict those vulnerabilities that have been co-occurring more often.
- the vulnerability labels e.g. names or references
- a cipher is created which can be used to encode and decode human-understandable vulnerability information into/from numbers and vectors.
- the data may also be split into training and testing/validation sets. As an example, 80 percent of the data may be used in the training phase and 20 percent may be used for testing/validation of the trained machine learning model.
- the machine learning model is trained using the training data.
- the machine learning model could be a decision tree or random forest classifier that supports non-binary (e.g. multi-class) classification.
- step 55 the model performance is evaluated using the test data.
- One evaluation factor may the coverage of the machine learning model’s prediction e.g. if vulnerability X is predicted to be coupled with vulnerability A and B, then this can be compared with the ground truth.
- step 56 the evaluation result is checked. If the model does not satisfy a quality (or reliability) criteria, then the method moves to step 54 and hyper-parameter tuning is performed for the machine learning model. As an example, if the machine learning model is a random forest classifier, then the depth and number of estimators may be incrementally changed and the model re-trained using the new hyper-parameters (as in step 53).
- step 57 the model is stored for use in predicting co-occurrent vulnerabilities in real-time.
- Python library for machine learning https ://scikit- learn.org/stable/modules/generated/sklearn. ensemble. RandomForestClassifier.html
- predict and classifier. predict_proba can be applied in model performance evaluation using y_test where y_test are coupled vulnerabilities (Class), and their probabilities of occurrence (Float). For instance, during evaluation, the system takes a configurable parameter of requiring model accuracy of 80% of labels being predicted correctly, then compares classifier. predict(X_test) with y_test, to see if this requirement is fulfilled or not. If requirement is not fulfilled, then conduct hyper parameter tuning as depicted in following bullet.
- increase/decrease parameters e.g., n_estimators from 20 to 21 and change criterion from "gini” to "entropy”, and repeat increases/decreases until a set of parameters are found that satisfy the model performance criterion.
- the machine learning model may be used to predict co-occurring vulnerabilities in real time using a method such as the method 400 shown in Fig. 4.
- the method 400 may be performed by a node such as the node 100 described above.
- the method 400 comprises, in a first step 402, detecting a second vulnerability in a host server.
- the method 400 comprises using a machine learning model to predict a second subset of co-occurring vulnerabilities that are predicted to occur in an overlapping time frame with the first vulnerability, the machine learning model having been trained using training data comprising example vulnerabilities and corresponding example subsets of co-occurring vulnerabilities.
- the method 400 may be performed subsequent to the method 200, or as a stand-alone method, e.g. using a machine learning model trained using the method 200.
- the second vulnerability may be a new vulnerability, e.g. detected as part of a new scan in a new vulnerability scan report.
- the new vulnerability may be a vulnerability of interest to a user.
- the second vulnerability is input to a machine learning model obtained, for example, according to the method 200.
- the machine learning model may provide as output, a prediction of a second subset of co-occurring vulnerabilities that are predicted to co-occur with the second vulnerability.
- the second vulnerability and the second subset of co-occurring vulnerabilities may be used to determine a patch (e.g. a security patch) for the second vulnerability.
- a patch is a remedy or action to perform in order to neutralise (or minimise the effects/exploitability of) the second vulnerability.
- the patch may patch both the second vulnerability and one or more of the co-occurring vulnerabilities in the second subset of cooccurring vulnerabilities. Because the method 400 may be used to obtain predictions of cooccurring vulnerabilities, patches may be chosen that better address both the second vulnerability and its co-occurring vulnerabilities in one go.
- a vulnerability apparatus 600 receives a request from user, machine or other apparatus in step 61 where, the request indicates a vulnerability of interest (e.g a new vulnerability) and the user expects to gain knowledge of the coverage and the coupled vulnerabilities that are related to the vulnerability of interest.
- a vulnerability of interest e.g a new vulnerability
- An indication of a vulnerability may be e.g. a name, label or any other reference to the vulnerability of interest.
- the request may indicate, e.g. that the user has found a vulnerability X with risk score of low.
- the output of the vulnerability apparatus 600 may indicate that the vulnerability of interest, which in this example may be labelled “X” has occurred 100 times on comparable assets, and in 80% of the occurrences, vulnerabilities A and B have also occurred, where A has medium risk score and B has a high risk score.
- the user can compare if their system has also reported A and B and the user can consider vulnerabilities X, A and B in security hardening, rather than considering vulnerability X in isolation.
- the number of occurrences (100) may guide the user in data-driven decision making in security patching and hardening.
- the proposed solutions herein use available vulnerability report data as an input and, in machine learning, all information is included in the trained model. Therefore, we do not need to query the report data each time the user makes a request rather we can do the predictions using machine learning model in utilization of deep vulnerability intelligence.
- the proposed solutions enable advanced security vulnerability analytics where one can create knowledge of the most vicious vulnerabilities based on risk scores, coverage and other coupled vulnerabilities.
- an analytics functionality could return the most interesting/vicious vulnerabilities also considering preferences in risk score, coverage and number of other coupled vulnerabilities and their risk scores.
- machine learning can thus be used to gain vulnerability intelligence from multiple vulnerability scan results. Even a single scan result may consist of tens of thousands of lines which can be challenging to be processed by human. Using the method herein, it is no longer necessary to transfer or read collections of sensible vulnerability scan reports each time that a vulnerability of interest is processed. All information from these reports can be included in the machine learning model.
- Embodiments herein have increased ranking dimensions by introducing coverage and additional coupled vulnerabilities which lead to a great impact in actionable data.
- embodiments herein provide intelligent vulnerability assessment based on coverage and coupling.
- a computer program product comprising a computer readable medium, the computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform the method or methods described herein, such as the methods 200 and 400.
- the disclosure also applies to computer programs, particularly computer programs on or in a carrier, adapted to put embodiments into practice.
- the program may be in the form of a source code, an object code, a code intermediate source and an object code such as in a partially compiled form, or in any other form suitable for use in the implementation of the method according to the embodiments described herein.
- a program code implementing the functionality of the method or system may be sub-divided into one or more sub-routines.
- the sub-routines may be stored together in one executable file to form a self-contained program.
- Such an executable file may comprise computer-executable instructions, for example, processor instructions and/or interpreter instructions (e.g. Java interpreter instructions).
- one or more or all of the sub-routines may be stored in at least one external library file and linked with a main program either statically or dynamically, e.g. at run-time.
- the main program contains at least one call to at least one of the sub-routines.
- the sub-routines may also comprise function calls to each other.
- Fig. 7 shows a carrier 700 containing a computer program 106.
- a carrier may be an electronic signal, optical signal, radio signal or computer readable storage medium.
- the carrier of a computer program may be any entity or device capable of carrying the program.
- the carrier may be or include a computer readable storage medium, such as a ROM, for example, a CD ROM or a semiconductor ROM, or a magnetic recording medium, for example, a hard disk.
- the carrier may be a transmissible carrier such as an electric or optical signal, which may be conveyed via electric or optical cable or by radio or other means.
- the carrier When the program is embodied in such a signal, the carrier may be constituted by such a cable or other device or means.
- the carrier may be an integrated circuit in which the program is embedded, the integrated circuit being adapted to perform, or used in the performance of, the relevant method.
- Fig. 8 shows a computer program product 800 comprising non transitory computer readable media 802 having stored thereon a computer program 106 as described above.
- a computer program may be stored/distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Computing Systems (AREA)
- Computer Security & Cryptography (AREA)
- Software Systems (AREA)
- Computer Hardware Design (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Data Mining & Analysis (AREA)
- Computational Linguistics (AREA)
- Mathematical Physics (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- General Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Debugging And Monitoring (AREA)
- Data Exchanges In Wide-Area Networks (AREA)
Abstract
A method performed by a node (100) in a communications network (107). The method comprises: obtaining (202) vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, matching (204) a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and training (206) a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
Description
TRAINING A MACHINE LEARNING MODEL
TECHNICAL FIELD
This disclosure relates to a method, apparatus, a computer program, a carrier and a computer program product. More particularly but non-exclusively, the disclosure relates to training a machine learning model to predict vulnerabilities in a communications network.
BACKGROUND
The disclosure herein relates to security in computer networks, hosts and servers. Current processes used for vulnerability detection, scanning, analysis, and patching solutions are largely based around the following activities:
Conducting vulnerability scanning activities with vulnerability scanning tools that attempt to identify the operating systems in a target server and the software installed on them, together with other attributes such as open ports and system accounts.
Maintaining a database of known vulnerabilities, checking scanning results against the vulnerability database, and highlighting the known vulnerabilities.
Having a severity rating system of vulnerabilities based on, e.g. the National Vulnerability Database (NVD) which uniquely lists Common Vulnerabilities and Exposures (CVE) which is integrated with Common Vulnerability Scoring System (CVSS) for assessing the severity of security vulnerabilities that are found.
- Vulnerabilities are often given a severity rating based on how easy the weakness is to exploit, and the impact exploiting it can have on the program, server or data.
Evaluating system, software and infrastructure for unpatched holes and gaps based on scan results, where the evaluation is performed by a security expert using his/her prior knowledge and experience.
Remediating known vulnerabilities by installing security patches to the systems in order to prevent incidents where known vulnerabilities are exploited, from taking place.
The prevailing philosophy in existing systems is to develop a process and plan to continuously assess, track and mitigate vulnerabilities and security weaknesses in servers within a system infrastructure using vulnerability scanning tools in order to minimize exposure to exploitable vulnerabilities in the system.
The document NIST International Report 8409 “Measuring the Common Vulnerability Scoring System Base Score Equation”, November 2022,
https://doi.org/10.6028/NIST.IR.8409, retrieved on 25 February 2023 discloses an evaluation of the CVSS Version 3 base score equation.
SUMMARY
There are various issues with the current vulnerability detection, scanning, analysis, and patching solutions. For example, a single vulnerability scan for large systems can result in tens of thousands of lines of findings, which require deep and complex manual analysis by a (human) security expert. Each new scan creates similar massive amounts of findings with slightly different results. Often a vulnerability scanning result is transported as a report in excel or other format for a security expert to review. The report may contain sensitive information that requires obfuscation or pseudonymization of the sensitive data items.
Furthermore, each identified vulnerability is classified individually based on its CVSS value which attempts to assign severity scores to single vulnerabilities, allowing responders to prioritize responses and resources according to threat. However, it is often complex to chain vulnerabilities to identify and evaluate the impact of complex threats which may be comprised of several simultaneous vulnerabilities. Lastly, security patches are generally created on an individual basis resulting in a time consuming and potentially troublesome patching process.
It is an object of the invention herein to address some of these issues to provide improved security processes.
Thus, according to a first aspect herein there is a method performed by a node in a communications network. The method comprises: i) obtaining vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) matching a first vulnerability in the plurality of vulnerabilities to a first subset of cooccurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) training a machine learning model using the first vulnerability as an example input and the first subset of cooccurring vulnerabilities as an example output of the machine learning model.
According to a second aspect herein there is a method performed by a node in a communication system, the method comprising: detecting a second vulnerability in a host server, and using a machine learning model to predict a second subset of co-occurring vulnerabilities that are predicted to occur at an overlapping time with the first vulnerability, wherein the machine learning model was trained using training data comprising example vulnerabilities and corresponding example subsets of co-occurring vulnerabilities.
According to a third aspect herein there is a node in a communications network, the node comprising: a memory comprising instruction data representing a set of instructions;
and a processor configured to communicate with the memory and to execute the set of instructions, wherein the set of instructions, when executed by the processor, cause the processor to: i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
According to a fourth aspect there is a node in a communications network configured to: i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
According to a fifth aspect there is a computer program comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out a method according to the first or second aspects.
According to a sixth aspect there is a carrier containing a computer program according to the fifth aspect wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.
According to a seventh aspect there is a computer program product comprising non transitory computer readable media having stored thereon a computer program according to the fifth aspect.
Thus, described herein is a method, apparatus, a computer program, a carrier and a computer program product that introduce co-occurring vulnerabilities as metrics to enable deeper vulnerability analysis beyond traditional risk-based ranking. In this way, vulnerabilities output from a vulnerability scan can be linked together in an automated manner using a data-driven and machine learning approach. This reduces the burden on human engineers manually reviewing large scan reports. It further allows for improved patching, since patches may be better targeted at groups of co-occurring vulnerabilities, rather than merely individual vulnerabilities with the highest CVSS values. Linking vulnerabilities in this manner thus provides data of increased actionable value which can be used to increase the efficiency of security hardening processes.
BRIEF DESCRIPTION OF THE DRAWINGS
For a better understanding and to show more clearly how embodiments herein may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings, in which:
Fig. 1 illustrates an example node according to embodiments herein;
Fig. 2 shows an example method according to embodiments herein; Fig. 3 shows an example method according to embodiments herein; Fig. 4 shows an example method according to embodiments herein; Fig. 5 shows an example method according to embodiments herein; Fig. 6 shows an example system according to some embodiments herein; Fig. 7 shows an example carrier containing a computer program; and Fig. 8 shows an example computer program product.
DETAILED DESCRIPTION
The disclosure herein relates to a method, apparatus, a computer program, a carrier and a computer program product in a communications network. A communications network or telecommunications network may comprise any one, or any combination of: a wired link (e.g. ASDL) or a wireless link such as Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), New Radio (NR), WiFi, Bluetooth or future wireless technologies. The skilled person will appreciate that these are merely examples and that a communications network may comprise other types of links. A wireless network may be configured to operate according to specific standards or other types of predefined rules or procedures. Thus, a wireless network may implement communication standards, such as Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), and/or other suitable 2G, 3G, 4G, 5G, 6G or future generation of standards; wireless local area network (WLAN) standards, such as the IEEE 802.11 standards; and/or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave and/or ZigBee standards.
Fig. 1 shows an example node 100 in a communications network 107 according to some embodiments herein. The node 100 may be otherwise referred to herein as a network node. The node 100 is configured (e.g. adapted, operative, or programmed) to perform any of the embodiments of the method 200 or 400 as described below.
The node 100 comprises a processor (e.g. processing circuitry or logic) 102. The processor 102 may control the operation of the node 100 in the manner described herein. The processor 102 can comprise one or more processors, processing units, multi-
core processors or modules that are configured or programmed to control the node 100 in the manner described herein. In particular implementations, the processor 102 can comprise a plurality of computer programs and/or hardware modules that are each configured to perform, or are for performing, individual or multiple steps of the functionality of the node 100 as described herein.
The node 100 comprises a memory 104. In some embodiments, the memory 104 of the node 100 can be configured to store a computer program 106 with program code or instructions that can be executed by the processor 102 of the node 100 to perform the functionality described herein. Alternatively, or in addition, the memory 104 can be configured to store any requests, resources, information, data, signals, or similar that are described herein. The processor 102 may be configured to control the memory 104 to store any requests, resources, information, data, signals, or similar that are described herein.
It will be appreciated that the node 100 may comprise one or more virtual machines running different software and/or processes. The node 100 may therefore comprise one or more servers, switches and/or storage devices and/or may comprise cloud computing infrastructure or infrastructure configured to perform in a distributed manner, that runs the software and/or processes.
It will be appreciated that the node 100 may comprise other components in addition or alternatively to those indicated in Fig. 1. For example, in some embodiments, the node 100 may comprise a communications interface. A communications interface may be for use in communicating with other nodes e.g. via a communications network. For example, the communications interface may be configured to transmit to and/or receive from nodes or network functions requests, resources, information, data, signals, or similar. The processor 102 may be configured to control such a communications interface to make/receive such transmissions.
The node 100 may be implemented in (e.g. form part of) a communications network 107. In some embodiments herein, the node 100, may be implemented in a management layer of a communications network. For example, the node 100 may comprise a node in a network management system, security management system or a vulnerability management platform.
These are merely examples however, and the node 100 may generally be any node, network function host, server, server farm, distributed computer system or network device in a communications network. For example, the node 100 may comprise any component or network function (e.g. any hardware) in a communications network suitable for performing the functions described herein. Examples of the node 100 include, but are not limited to, core network function hosts such as, for example, hosts for core network functions in a Fifth Generation Core network (5GC). It is realized that the node 100 may equally be a
node/device in any future network, such as a future 3GPP (3rd Generation Partnership Project) sixth generation communication network, irrespective of whether the node 100 would there be placed in a core network or outside of the core network. It is also realized that the node 100 may be positioned in a part of the communications network which is outside of the future 3GPP communication network, but connected to the future 3GPP communication network, e.g. to provide a cloud computing platform to offload the 3GPP communication network..
Generally (as will be described in more detail below), the node 100 is for use in security patching of host servers in a communications network. As described in the background and summary sections, current vulnerability scans tend to produce many hundreds, or thousands of lines of vulnerability data that are manually reviewed by human security experts who have to determine which vulnerabilities to patch and how to go about patching them. Embodiments herein aim to aid such human engineers by using machine learning to predict for an input vulnerability, a subset of co-occurring vulnerabilities that are predicted to occur at the same or overlapping time as the input vulnerability. Often, vulnerabilities are linked, for example, they may arise or become exploitable when particular conditions are found in a network. Thus, by predicting groups of co-occurring vulnerabilities in this manner, a human engineer may be able to select a better security patch in order to patch an entire group of vulnerabilities in one go (rather than patching individual vulnerabilities one at a time in a more ad-hoc manner), thus improving and simplifying the security patching process.
To this end, briefly, the node 100 is configured to i) obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan, ii) match a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability, and iii) train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
A host server in this context is a server that is connected to the internet (e.g. via a communications network as described above). A host server may run services accessible via the internet. Such host servers may be scanned using vulnerability scanners (e.g. software) to detect security vulnerabilities. Vulnerability scanners are tools that monitor applications and networks to identify security vulnerabilities. Vulnerability scanners conduct scans to identify potential exploits or vulnerabilities by comparing e.g. configuration parameters, network traffic patterns or similar, to a database of known vulnerabilities, or by comparing scanning results against a database of known
vulnerabilities and security weaknesses. Vulnerability scanners are used by companies to test applications and networks against known vulnerabilities and to identify new vulnerabilities. The output of a scan is a report which, as described above in the background, may contain thousands of lines of identified vulnerabilities at any given time. Each vulnerability may be graded according to severity. Vulnerabilities may be time dependent, and thus a scan report may comprise time information, for example, indicating a time at which a vulnerability was detected and/or a length of time in which the vulnerability was detected. A scan report may detail the state of an application or host server and/or provide recommendations to remedy known issues.
Fig. 2 shows a method 200 in a node 100 in a communications network according to some embodiments herein. The method 200 is computer implemented in that it is performed by the node 100. Briefly, in a first step 202 the method comprises obtaining vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan. In a second step 204 the method comprises matching a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability. In a third step 206 the method comprises training a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
In more detail, in step 202, vulnerability data is obtained from (e.g. output from) a vulnerability scan. As noted above, the vulnerability data comprises a plurality of vulnerabilities, e.g. in a list. Each vulnerability may be time stamped. In some examples, the time stamp may indicate the time at which the vulnerability was first detected. In other examples, the time stamp may be in the form of a time window (e.g. a time at which each vulnerability was first detected and a time at which it ceased to be detected).
In step 204, a first vulnerability of the plurality of vulnerabilities is matched with a first subset of co-occurring vulnerabilities. The first subset of co-occurring vulnerabilities are one or more other vulnerabilities in the plurality of vulnerabilities that co-occurred with the first vulnerability. Co-occurred in this sense may mean overlapping in time, in any manner. For example, the first vulnerability may occur in an overlapping, or partially overlapping time window with each co-occurring vulnerability in the plurality of co-occurring vulnerabilities. As another example, the first vulnerability may occur simultaneously, e.g. at the same time as each co-occurring vulnerability in the plurality of co-occurring vulnerabilities. Generally, the first vulnerability may overlap in time in a different manner to each co-occurring vulnerability, e.g. it may start and stop at (approximately) the same time as some of the first subset of cooccurring vulnerabilities, while merely overlapping in time with other vulnerabilities in the first
subset of co-occurring vulnerabilities. So long as a vulnerability overlaps in time in some manner, then it may be matched to the first vulnerability and added to the first subset of cooccurring vulnerabilities.
The output of step 204 is an indication of the first vulnerability and the first subset of cooccurring vulnerabilities. In this sense, an indication is a name, label or any other reference that may be used to identify a vulnerability. The output may be in the form of a tuple:
[first vulnerability; first subset of co-occurring vulnerabilities].
It will be appreciated that the output may be encoded, e.g. using a cipher to convert each vulnerability into a numerical or vectoral identifier. This is described in more detail below.
Steps 202 and 204 may be repeated for each vulnerability in the plurality of vulnerabilities, e.g. each unique vulnerability (e.g. the steps may not be repeated for duplicated vulnerabilities in the plurality of vulnerabilities). In some embodiments, for each co-occurring vulnerability in the first subset of co-occurring vulnerabilities, an indication of a total number of co-occurrences of the co-occurring vulnerability with the first vulnerability may be obtained. In other words, the total number of co-occurrences of each (other) vulnerability with the first vulnerability may be obtained in step 204.
An embodiment of steps 202 and 204 are illustrated in Fig 3 which illustrates an example data processing function. In Fig. 3, in step 31 , step 202 is performed and vulnerability data from a vulnerability scan is obtained. In this example the vulnerability data (otherwise referred to as input data) is the output of one or more Nessus™ scans, although any other vulnerability scan reports could equally be used. As an example, the following fields of a vulnerability scan may be obtained:
Host Information
IP
Port
Asset label
Vulnerability Information
Name
Synopsis
Solution
Plugin ID
Risk Description
Low
High
Critical
CVE (Common Vulnerabilities and Exposures)
Common Vulnerability Scoring System (CVSS)
It will be appreciated that these fields are merely an example and that other fields could also be used, alternatively or additionally to those described above.
In step 31 the input data is read and a list of unique vulnerabilities is determined (e.g. duplicated vulnerabilities are removed).
In step 32, a first vulnerability is selected from the plurality of vulnerabilities output from step 31.
In step 33, coupled vulnerabilities that occur in overlapping time windows with the first vulnerability, and within the same host are identified. For example, if vulnerability X occurs in host Y, then in step 33, the method may comprise counting simultaneous, or overlapping occurrences of vulnerabilities A, B, C,... This may be repeated for all hosts in a given data set.
In step 34, the coupled vulnerabilities may be cached and stored e.g. in the form:
Vulnerability X : {
Vulnerability A: 7
Vulnerability B: 2
}
Vulnerability A: {
Vulnerability B: 1
}•
In this example, vulnerability X was found to co-occur with vulnerability A 7 time, and vulnerability B twice, while Vulnerability A was found to co-occur with vulnerability B once.
In this example in step 35, the process is repeated for second and subsequent vulnerabilities in a looped manner until each unique vulnerability in the plurality of vulnerabilities has been assessed. The method may further be repeated for vulnerability data obtained from a plurality of different host servers.
In step 36, the cached coupled vulnerabilities are stored, e.g. in long term storage.
As noted above, the input data to the process illustrated in Fig. 3 is static (e.g. the input parameter types and/or structure is always the same), and may be found from vulnerability scan reports. As an output, the process in Fig. 3 produces coupled vulnerability information that depicts how many times a vulnerability has co-occurred at the same time as each other vulnerability. Once every vulnerability has been processed (e.g. once the loop in step 35 has been exhausted), then the coverage of each vulnerability may also be
determined. Coverage indicates the number of occurrences of each vulnerability and may be expressed in terms of vulnerability X occurred x times, or in y percentage of host servers scanned.
Turning back to the method 200, in step 206, the method 200 comprises training a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model. In other words, the output from step 204 is used as (a piece of) training data with which to train a machine learning model, the method 200 may be looped as described above, on each vulnerability in the plurality of vulnerabilities to build up a training data set of [vulnerability, co-occurring vulnerabilities] pairs. A machine learning model may be trained in this way to output, for an input vulnerability, one or more other vulnerabilities that are predicted to co-occur with an input vulnerability.
The skilled person will be familiar with machine learning and models that can be trained using machine learning processes, but briefly, a machine learning process may comprise a procedure that is run on data to create a model that may be referred to as a “machine learning model”. Machine learning process comprise procedures and/or instructions through which training data, may be processed or used in a training process to generate a machine learning model. Examples of machine learning processes include but are not limited to processes such as back-propagation and gradient descent. Machine learning processes learn from training data comprising example input and output pairs. The example outputs in a training data example are the ground truth or “correct” outputs for the respective example input.
In some embodiments herein, the machine learning model is multi-class classifier. For example, a decision tree (see paper by Quinlan (1986) entitled: “Induction of decision trees" Machine Learning volume 1 , pages 81-106 (1986)) or a random forest-based classifier (see paper by Breiman (2001) entitled “Random Forests"', Mach Learn 45(1): 5-32).
In other embodiments, the first machine learning model is a deep neural network, see paper by Schmidhuber (2015) entitled: “Deep learning in neural networks: An overview” Neural Networks Volume 61, January 2015, Pages 85-117.
Although examples have been given above, more generally, the machine learning model (otherwise referred to herein as the model) can be any type of classification or regression model, trained using a supervised or semi-supervised learning process.
Example Model Inputs
In some embodiments, the input to the model may be an indication of a vulnerability. An indication in sense is e.g. a name, label, or any other reference that can be used to identify a vulnerability.
There may be other input parameters, such as for example, data describing a host server (e.g. size, location of server, configuration parameters of server) for which the input vulnerability has been detected.
The below italic gives an example model input. In this example the first row provides column names, the second row depicts a first vulnerability that occurs within host xx.yyy.cf.d and third row depicts a second vulnerability that is found within same host (=coupled). It will be appreciated that this is a realistic example but that it utilises dummy inputs, e.g., host IP address is given as xx.yyy.cf.d.
The below italic gives an example model input. For the shake of clarity some input from real model input is omitted from the illustration below. The Illustration is also formatted in human readable format.
Plugin ID: 1
CVE: CVE-2020-10188
NVD Published Date: 03/06/2020
NVD Last Modified: 11/30/2021
CVSS v2.0 Base Score: Medium
Host: xx.yyy.cf.d
Protocol. TCP
Port: 23
Name: Telnet Server Detection
Description: A Telnet server is listening on the remote port. The remote host is running a Telnet server, a remote terminal server
Solution: Disable this service if you do not use it.
Plugin ID: 2
CVE: CVE-2020-10182
NVD Published Date: 15/09/2017
NVD Last Modified: 01/23/2020
CVSS v2.0 Base Score: Medium
Host: xx.yyy.cf.d
Protocol. TCP
Port: 23
Name: Unencrypted Telnet Server
Description: The remote Telnet server transmits traffic in cleartext
Solution: SSH is preferred over Telnet since it protects credentials from eavesdropping and can tunnel additional data streams such as an X11 session. Disable the Telnet service and use SSH instead.
Example Model Outputs
In some embodiments, the output of the machine learning model is an indication of a subset of co-occurring vulnerabilities that are predicted to co-occur with a respective input vulnerability. In such an example, the training data may be in the form of [example input vulnerability, example subset of co-occurring vulnerabilities predicted for the example input vulnerability] pairs.
In some embodiments, the output of the machine learning model comprises an indication of a coverage of the input vulnerability. For example the output may indicate the number of comparable host servers on which the input vulnerability has occurred. E.g. the vulnerability has occurred on / percent of similar host servers. In such an example, the training data may be in the form of [example input vulnerability, [example subset of cooccurring vulnerabilities predicted for the example input vulnerability, coverage]] pairs.
In some embodiments, the output may further comprise a risk score for each of the predicted co-occurring vulnerabilities.
As noted above, the inputs and outputs may be presented in a human-readable form e.g. each may be given a name or other reference. In order to optimise and simplify the training of the machine learning model, these names or references may be converted into a numerical/vectoral form which may be more readily processed or read by a machine. Thus, in step 206, the first vulnerability and the first subset of co-occurring vulnerabilities (e.g. the names or references associated with each vulnerability) may be transformed using a mapping into a unique numerical form.
As an example, a cipher (or transformation cipher) may be used for the transformation. A cipher refers to a mapping function that can be used to encode and decode vulnerability names into numbers and vice-versa, e.g. “vulnerability X”<-> 1. A cipher may be a simple enumeration in which each unique vulnerability is assigned a unique number. In other examples, a cipher may be a more complex transformation to convert each vulnerability in the plurality of vulnerabilities into a form more easily processed by a machine learning model, e.g. into a numerical or vector form.
Turning now to Fig. 5, which shows a method of training a machine learning model and creation of a transformation cipher, according to some embodiments herein. In this embodiment in step 51 , the input data is read (e.g. obtained or received) and unique
vulnerabilities are found in the input data. In this embodiment, the input data is the coupled vulnerability data output from step 204 above.
As an example, the input data may be in the following format: Vulnerability X : { Vulnerability A: 7 Vulnerability B: 2 } Vulnerability A: { Vulnerability B: 1 }• Furthermore, vulnerability description information that considers the following for each vulnerability may also be used: vulnerability:
Host Information
IP
Port
Asset label
- Vulnerability Information
Name
Synopsis Solution Plugin ID Risk description Low High critical
- CVE
Training data may be created from this input. In this example, the numbers (of co-occurrences) are used to determine the number of samples in the model training. As an example, if Vulnerability A has co-occurred with Vulnerability X 7 times, then this sample may be added 7 times. As such, the above input data example, can be translated to a training set that considers Vulnerability X as the vulnerability of interest is composed from 7 samples of Vulnerability A, and 2 samples of Vulnerability B. In this way, the ratios/distribution of the coupled vulnerabilities are preserved because the machine learning model will be more likely to predict those vulnerabilities that have been co-occurring more often. Thus, in this example, post-training, the predictive phase could indicate that Vulnerability X is coupled with Vulnerability A and Vulnerability B, where probability score is higher for A (=due to higher
number of occurrences) and probability score is lower for B (=yet, the Vulnerability B still has a positive probability to be occurring if Vulnerability X has occurred).
In step 52, the vulnerability labels (e.g. names or references) are transformed into numbers and vectors. Here a cipher is created which can be used to encode and decode human-understandable vulnerability information into/from numbers and vectors. The data may also be split into training and testing/validation sets. As an example, 80 percent of the data may be used in the training phase and 20 percent may be used for testing/validation of the trained machine learning model.
In step 53, the machine learning model is trained using the training data. As noted above, the machine learning model could be a decision tree or random forest classifier that supports non-binary (e.g. multi-class) classification.
In step 55, the model performance is evaluated using the test data. One evaluation factor may the coverage of the machine learning model’s prediction e.g. if vulnerability X is predicted to be coupled with vulnerability A and B, then this can be compared with the ground truth.
In step 56, the evaluation result is checked. If the model does not satisfy a quality (or reliability) criteria, then the method moves to step 54 and hyper-parameter tuning is performed for the machine learning model. As an example, if the machine learning model is a random forest classifier, then the depth and number of estimators may be incrementally changed and the model re-trained using the new hyper-parameters (as in step 53).
If the model performs well, then in step 57, the model is stored for use in predicting co-occurrent vulnerabilities in real-time.
Pseudo-code for the example in Fig. 5, is given below. The pseudo-code is based on the use of Open-Source Machine Learning Library SciKit-Learn described in the paper: “Scikit-learn: Machine Learning in Python”, Pedregosa et al., JMLR 12, pp. 2825- 2830, 2011.
Python library for machine learning : https ://scikit- learn.org/stable/modules/generated/sklearn. ensemble. RandomForestClassifier.html
»> classifies RandomForestClassifier(n_estimators = 20, criterion = {“gini”, “entropy”,
“logjoss”}, default-’gini”, max_depth=2, random_state=0, maxJeaf_nodes=None)
»> classifier.fit(XJrain, yjrain), where X rain are vulnerabilities of interest, and yjrain are coupled vulnerabilities
»> classifier. predict(X est), where XJest are vulnerabilities of interest that are not introduced during model training. Returns coupled vulnerabilities (Class)
»> classifier. predict_proba(X_test), where X_test are vulnerabilities of interest that are not introduced during model training. Returns probability of coupled vulnerabilities (Float)
»> The results from classifier. predict and classifier. predict_proba can be applied in model performance evaluation using y_test where y_test are coupled vulnerabilities (Class), and their probabilities of occurrence (Float). For instance, during evaluation, the system takes a configurable parameter of requiring model accuracy of 80% of labels being predicted correctly, then compares classifier. predict(X_test) with y_test, to see if this requirement is fulfilled or not. If requirement is not fulfilled, then conduct hyper parameter tuning as depicted in following bullet.
»> During hyper parameter tuning, increase/decrease parameters, e.g., n_estimators from 20 to 21 and change criterion from "gini" to "entropy", and repeat increases/decreases until a set of parameters are found that satisfy the model performance criterion.
The machine learning model may be used to predict co-occurring vulnerabilities in real time using a method such as the method 400 shown in Fig. 4. the method 400 may be performed by a node such as the node 100 described above.
The method 400 comprises, in a first step 402, detecting a second vulnerability in a host server. In a second step 404, the method 400 comprises using a machine learning model to predict a second subset of co-occurring vulnerabilities that are predicted to occur in an overlapping time frame with the first vulnerability, the machine learning model having been trained using training data comprising example vulnerabilities and corresponding example subsets of co-occurring vulnerabilities. The method 400 may be performed subsequent to the method 200, or as a stand-alone method, e.g. using a machine learning model trained using the method 200.
In step 402, the second vulnerability may be a new vulnerability, e.g. detected as part of a new scan in a new vulnerability scan report. Alternatively, the new vulnerability may be a vulnerability of interest to a user.
In step 404, the second vulnerability is input to a machine learning model obtained, for example, according to the method 200. The machine learning model may provide as output, a prediction of a second subset of co-occurring vulnerabilities that are predicted to co-occur with the second vulnerability.
The second vulnerability and the second subset of co-occurring vulnerabilities may be used to determine a patch (e.g. a security patch) for the second vulnerability. In this sense, a patch is a remedy or action to perform in order to neutralise (or minimise the effects/exploitability of) the second vulnerability. The patch may patch both the second
vulnerability and one or more of the co-occurring vulnerabilities in the second subset of cooccurring vulnerabilities. Because the method 400 may be used to obtain predictions of cooccurring vulnerabilities, patches may be chosen that better address both the second vulnerability and its co-occurring vulnerabilities in one go.
Turning now to Fig. 6 which shows a method of using a machine learning model according to an embodiment herein. In this example, in real-time, a vulnerability apparatus 600 receives a request from user, machine or other apparatus in step 61 where, the request indicates a vulnerability of interest (e.g a new vulnerability) and the user expects to gain knowledge of the coverage and the coupled vulnerabilities that are related to the vulnerability of interest. An indication of a vulnerability may be e.g. a name, label or any other reference to the vulnerability of interest. The request may indicate, e.g. that the user has found a vulnerability X with risk score of low.
In step 64, a machine learning model, trained according to the method 200 described above, may be obtained in step 63a from a database. A cipher may also be obtained in step 63b, to convert the indication of the vulnerability of interest into a machine- readable number or vector. In step 64, the vulnerability of interest is mapped (or encoded) using the cipher and the resulting encoded label for the vulnerability of interest is provided to the machine learning model in step 65. In step 66 and 67, in response to the request, the vulnerability apparatus 600 returns one or more co-occurring vulnerabilities that are predicted to occur at an overlapping time period with the vulnerability of interest. These may be decoded in step 66, e.g. from machine-readable numbers or vectors as output by the machine learning model, back into human-readable names/labels.
As an example, the output of the vulnerability apparatus 600 may indicate that the vulnerability of interest, which in this example may be labelled “X” has occurred 100 times on comparable assets, and in 80% of the occurrences, vulnerabilities A and B have also occurred, where A has medium risk score and B has a high risk score. With this output, the user can compare if their system has also reported A and B and the user can consider vulnerabilities X, A and B in security hardening, rather than considering vulnerability X in isolation. Also the number of occurrences (100) may guide the user in data-driven decision making in security patching and hardening.
The proposed solutions herein use available vulnerability report data as an input and, in machine learning, all information is included in the trained model. Therefore, we do not need to query the report data each time the user makes a request rather we can do the predictions using machine learning model in utilization of deep vulnerability intelligence.
Moreover, the proposed solutions enable advanced security vulnerability analytics where one can create knowledge of the most vicious vulnerabilities based on risk scores, coverage and other coupled vulnerabilities. For example, an analytics functionality
could return the most interesting/vicious vulnerabilities also considering preferences in risk score, coverage and number of other coupled vulnerabilities and their risk scores.
The use of machine learning according to the methods described herein, can thus be used to gain vulnerability intelligence from multiple vulnerability scan results. Even a single scan result may consist of tens of thousands of lines which can be challenging to be processed by human. Using the method herein, it is no longer necessary to transfer or read collections of sensible vulnerability scan reports each time that a vulnerability of interest is processed. All information from these reports can be included in the machine learning model.
In traditional systems, vulnerabilities are ranked based on their risk rating. Embodiments herein have increased ranking dimensions by introducing coverage and additional coupled vulnerabilities which lead to a great impact in actionable data.
In addition to patching a single vulnerability, the coupled vulnerabilities can be used together with machine learning to predict other vulnerabilities that often occur simultaneously. In this way, security engineers can conduct efficient security patching where patching process can consider the other vulnerability information and thus increasing security coverage of a single patch. Thus, embodiments herein provide intelligent vulnerability assessment based on coverage and coupling.
Turning now to other embodiments, there is also provided a computer program product comprising a computer readable medium, the computer readable medium having computer readable code embodied therein, the computer readable code being configured such that, on execution by a suitable computer or processor, the computer or processor is caused to perform the method or methods described herein, such as the methods 200 and 400.
Thus, it will be appreciated that the disclosure also applies to computer programs, particularly computer programs on or in a carrier, adapted to put embodiments into practice. The program may be in the form of a source code, an object code, a code intermediate source and an object code such as in a partially compiled form, or in any other form suitable for use in the implementation of the method according to the embodiments described herein.
It will also be appreciated that such a program may have many different architectural designs. For example, a program code implementing the functionality of the method or system may be sub-divided into one or more sub-routines. Many different ways of distributing the functionality among these sub-routines will be apparent to the skilled person. The sub-routines may be stored together in one executable file to form a self-contained program. Such an executable file may comprise computer-executable instructions, for example, processor instructions and/or interpreter instructions (e.g. Java interpreter instructions). Alternatively, one or more or all of the sub-routines may be stored in at least
one external library file and linked with a main program either statically or dynamically, e.g. at run-time. The main program contains at least one call to at least one of the sub-routines. The sub-routines may also comprise function calls to each other.
Fig. 7 shows a carrier 700 containing a computer program 106. A carrier may be an electronic signal, optical signal, radio signal or computer readable storage medium. The carrier of a computer program may be any entity or device capable of carrying the program. For example, the carrier may be or include a computer readable storage medium, such as a ROM, for example, a CD ROM or a semiconductor ROM, or a magnetic recording medium, for example, a hard disk. Furthermore, the carrier may be a transmissible carrier such as an electric or optical signal, which may be conveyed via electric or optical cable or by radio or other means. When the program is embodied in such a signal, the carrier may be constituted by such a cable or other device or means. Alternatively, the carrier may be an integrated circuit in which the program is embedded, the integrated circuit being adapted to perform, or used in the performance of, the relevant method.
Fig. 8 shows a computer program product 800 comprising non transitory computer readable media 802 having stored thereon a computer program 106 as described above.
Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfil the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be stored/distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method performed by a node (100) in a communications network (107), the method comprising: obtaining (202) vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan; matching (204) a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of cooccurring vulnerabilities overlapping in time with the first vulnerability; and training (206) a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
2. A method as in claim 1 further comprising: repeating the matching (204) and training for each vulnerability in the plurality of vulnerabilities.
3. A method as in claim 1 or 2 wherein the matching further comprises: for each co-occurring vulnerability in the first subset of co-occurring vulnerabilities, obtaining an indication of a total number of co-occurrences of the co-occurring vulnerability with the first vulnerability; and wherein in the training (206) the total counts are further provided as an example output of the machine learning model.
4. A method as in claim 1 or 2 wherein the obtaining (202), matching (204) and training (206) are repeated for vulnerability data obtained from a plurality of different host servers.
5. A method as in any one of claims 1 to 4 further comprising: transforming (52) the vulnerability data using a mapping to transform each vulnerability in the plurality of vulnerabilities into a unique numerical form.
6. A method as in claim 5 wherein in the training (206), the first vulnerability and the first subset of co-occurring vulnerabilities are converted into their respective unique numerical forms before performing the training.
7. A method as in any one of the preceding claims further comprising: inputting a second vulnerability to the machine learning model; and receiving as output, a prediction of a second subset of co-occurring vulnerabilities that are predicted to co-occur with the second vulnerability.
8. A method as in claim 7 further comprising: using the second vulnerability and the second subset of co-occurring vulnerabilities to determine a patch for the second vulnerability.
9. A method as in any one of the preceding claims wherein the machine learning model is a multi-class classification model.
10. A method as in claim 9 wherein the multi-class classification model is a decision tree or a random forest classifier.
11. A method performed by a node (100) in a communication system (107), the method comprising: detecting (402) a second vulnerability in a host server; and using (404) a machine learning model to predict a second subset of cooccurring vulnerabilities that are predicted to occur at an overlapping time with the first vulnerability; wherein the machine learning model was trained using training data comprising example vulnerabilities and corresponding example subsets of co-occurring vulnerabilities.
12. A method as in claim 11 further comprising: using the second vulnerability and the second subset of co-occurring vulnerabilities to determine a patch for the second vulnerability.
13. A node (100) in a communications network (107), the node comprising: a memory (104) comprising instruction data representing a set of instructions; and a processor (102) configured to communicate with the memory and to execute the set of instructions (106), wherein the set of instructions, when executed by the processor, cause the processor to: obtain (202) vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan;
match (204) a first vulnerability in the plurality of vulnerabilities to a first subset of co-occurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability; and train (206) a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
14. A node (100) as in claim 13 further configured to perform the method of any one of claims 2 to 12.
15. A node (100) in a communications network (107) configured to: obtain vulnerability data comprising a plurality of vulnerabilities detected on a host server during a vulnerability scan; match a first vulnerability in the plurality of vulnerabilities to a first subset of cooccurring vulnerabilities in the plurality of vulnerabilities, the first subset of co-occurring vulnerabilities overlapping in time with the first vulnerability; and train a machine learning model using the first vulnerability as an example input and the first subset of co-occurring vulnerabilities as an example output of the machine learning model.
16. A node (100) as in claim 15 further configured to perform the method of any one of claims 2 to 12.
17. A computer program (106) comprising instructions which, when executed on at least one processor, cause the at least one processor to carry out a method according to any of claims 1 to 12.
18. A carrier (700) containing a computer program (106) according to claim 17, wherein the carrier comprises one of an electronic signal, optical signal, radio signal or computer readable storage medium.
19. A computer program product (800) comprising non transitory computer readable media (802) having stored thereon a computer program (106) according to claim 17.
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/SE2023/050172 WO2024181895A1 (en) | 2023-02-27 | 2023-02-27 | Training a machine learning model |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4674087A1 true EP4674087A1 (en) | 2026-01-07 |
Family
ID=85569613
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23710486.4A Pending EP4674087A1 (en) | 2023-02-27 | 2023-02-27 | Training a machine learning model |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4674087A1 (en) |
| CN (1) | CN121014186A (en) |
| WO (1) | WO2024181895A1 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10114954B1 (en) * | 2017-11-30 | 2018-10-30 | Kenna Security, Inc. | Exploit prediction based on machine learning |
| CN111949994A (en) * | 2020-08-19 | 2020-11-17 | 北京紫光展锐通信技术有限公司 | Vulnerability analysis method and system, electronic device and storage medium |
-
2023
- 2023-02-27 WO PCT/SE2023/050172 patent/WO2024181895A1/en not_active Ceased
- 2023-02-27 CN CN202380097510.2A patent/CN121014186A/en active Pending
- 2023-02-27 EP EP23710486.4A patent/EP4674087A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN121014186A (en) | 2025-11-25 |
| WO2024181895A1 (en) | 2024-09-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20200336507A1 (en) | Generative attack instrumentation for penetration testing | |
| US12476994B2 (en) | Automated cybersecurity vulnerability prioritization | |
| Koroniotis et al. | A deep learning-based penetration testing framework for vulnerability identification in internet of things environments | |
| CN118710224A (en) | Enterprise platform security management method and system based on artificial intelligence | |
| Hasan et al. | Root cause analysis of anomalies in 5g ran using graph neural network and transformer | |
| Gholian et al. | DeExp: Revealing model vulnerabilities for spatio-temporal mobile traffic forecasting with explainable AI | |
| Bibers et al. | A comprehensive comparative study of individual ML models and ensemble strategies for network intrusion detection systems | |
| Teixeira et al. | Beyond performance comparing the costs of applying Deep and Shallow Learning | |
| Chatzimiltis et al. | AI-on-RAN for cyber defense: An XAI-LLM framework for interpretable anomaly detection | |
| WO2024181895A1 (en) | Training a machine learning model | |
| Chatzimiltis et al. | Interpretable Anomaly-Based DDoS Detection in AI-RAN with XAI and LLMs | |
| Song et al. | Toward Automatically connecting IoT devices with vulnerabilities in the wild | |
| Vu et al. | Enhancing network attack detection across infrastructures: An automatic labeling method and deep learning model with an attention mechanism | |
| CN119402222A (en) | A network security defense method and device based on AI big model and ATT&CK framework | |
| Yang et al. | 5GR-DTAD: a domain and data-driven framework for diagnosing abnormal downlink throughput in 5G RAN | |
| Mogilicharla et al. | Edge-Deployable ML Agent for Real-Time Tactic and Technique Attribution in Microgrid Security | |
| Omer et al. | Hidden Markov models for predicting cell-level mobile networks performance degradation | |
| Odarchenko et al. | Development of the Testbed for Testing Deep Learning Based IDS System for 5G Network. | |
| Kumar et al. | Machine Learning-Based Web Application Firewall for Real-Time Threat Detection | |
| Sodhro et al. | 5G beyond for healthcare: Leveraging AI/ML and diverse datasets for cybersecurity | |
| Zhang et al. | A hybrid machine learning intrusion detection system for wireless sensor networks | |
| Termos et al. | ADAP-GNN: Adaptive property-aware graph neural network for intrusion detection in IoT networks | |
| US20260100964A1 (en) | Generative systems and methods for adaptive vulnerability management | |
| Rahman et al. | Explaining Network Intrusion Detection System with SHAP and LIME | |
| CN119066506B (en) | Data processing method and system applied to data center station construction |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250929 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |