WO2024178581A1 - 一种数据处理方法、装置、存储介质及程序产品 - Google Patents

一种数据处理方法、装置、存储介质及程序产品 Download PDF

Info

Publication number
WO2024178581A1
WO2024178581A1 PCT/CN2023/078600 CN2023078600W WO2024178581A1 WO 2024178581 A1 WO2024178581 A1 WO 2024178581A1 CN 2023078600 W CN2023078600 W CN 2023078600W WO 2024178581 A1 WO2024178581 A1 WO 2024178581A1
Authority
WO
WIPO (PCT)
Prior art keywords
attack
intrusion detection
samples
type
sample set
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/078600
Other languages
English (en)
French (fr)
Inventor
侯硕
岳青伦
李廷森
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to CN202380070132.9A priority Critical patent/CN119968810A/zh
Priority to PCT/CN2023/078600 priority patent/WO2024178581A1/zh
Publication of WO2024178581A1 publication Critical patent/WO2024178581A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L9/00Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
    • H04L9/40Network security protocols

Definitions

  • the present application relates to the field of network security technology and provides a data processing method, device, storage medium and program product.
  • intrusion detection samples In order to improve the network security of automobiles, the industry has built intrusion detection samples for some typical intrusion scenarios, and used these intrusion detection samples to conduct attack tests on automobiles before they leave the factory to evaluate the anti-attack performance of the automobiles. Although this method can ensure that only automobiles with good anti-attack performance leave the factory, since the intrusion detection samples are only built for some typical intrusion scenarios, the intrusion detection samples are obviously not comprehensive and are not conducive to improving the accuracy of network security assessment of automobiles.
  • the present application provides a data processing method, device, storage medium and program product for constructing a more comprehensive intrusion detection sample set to improve the accuracy of network security assessment of devices to be detected (such as vehicles).
  • the present application provides a data processing method applicable to a data processing device, which may be any device with processing capabilities, such as a server or a server cluster composed of multiple servers.
  • the method includes: the data processing device obtains a first type of attack sample by attacking a device to be detected, obtains a second type of attack sample by applying noise to the first type of attack sample, and then constructs an intrusion detection sample set based on the first type of attack sample and the second type of attack sample.
  • the first type of attack samples are obtained by attacking the device to be detected, and belong to attack samples under known intrusion scenarios
  • the second type of attack samples are obtained by adding noise to the first type of attack samples, and can be considered as attack samples under unknown intrusion scenarios obtained by performing some deformation on the attack samples under known intrusion scenarios.
  • an intrusion detection sample set is constructed by combining the first type of attack samples and the second type of attack samples, so that the intrusion detection sample set can cover both attack samples under known intrusion scenarios and attack samples under unknown intrusion scenarios, so that the intrusion detection samples in the intrusion detection sample set are more comprehensive.
  • using a more comprehensive intrusion detection sample set to evaluate the device to be detected can also improve the accuracy of network security assessment of the device to be detected.
  • the first type of attack samples may include real attack samples and simulated attack samples, wherein the real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected by an attack tool.
  • the human attack method can be used to construct attack behaviors that are highly concealed or rely on business logic, so as to obtain attack samples that are easy to mark in a real attack environment, while the attack tool attack method can obtain attack samples that are not easy to mark in a real attack environment by simulating the generation of attack traffic.
  • the human attack method and the attack tool attack method to comprehensively construct the first type of attack samples, the first type of attack samples can fully cover the attack samples of various attack types that may exist in the real attack environment, thereby improving the comprehensiveness of the first type of attack samples.
  • the real attack samples can correspond to one or more of the following attack types: identity document (ID) non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit (ECU) attack, identical ID attack, controller area network (CAN) scanning attack, unified diagnostic services (UDS) execution of sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown pop connection attack, unknown push connection attack.
  • ID identity document
  • replay attack tampering attack
  • data length error attack data length error attack
  • signal out of defined range attack context error attack
  • ID source non-specified electronic control unit (ECU) attack identical ID attack
  • controller area network (CAN) scanning attack controller area network (CAN) scanning attack
  • UDS unified diagnostic services
  • the simulated attack samples can correspond to one or more of the following attack types: ID fuzzy attack, data fuzzy attack, CAN denial-of-service (Dos) attack, Ethernet (ethnic, ETH) Dos attack, malformed packet injection attack, port scanning attack.
  • a data processing device obtains a first type of attack sample by attacking a device to be detected, including: the data processing device traverses each of a plurality of preset attack types, and when traversing each attack type: executes an attack behavior corresponding to the attack type on the device to be detected, and obtains traffic data generated by the device to be detected for the attack behavior, and then, when it is determined that the traffic data is attack traffic, marks the traffic data as a first type of attack sample.
  • the preset multiple attack types can be, for example, the multiple attack types that are easy to mark in the real attack environment and the multiple attack types that are not easy to mark in the real attack environment given in the aforementioned design.
  • the first type of attack samples can fully cover various known attack types, thereby improving the richness and comprehensiveness of the first type of attack samples.
  • the traffic data processing device obtains the traffic data generated by the device to be detected for any attack behavior, if it is determined that the traffic data is not attack traffic but normal traffic, the traffic data can be marked as a non-attack sample. Then, after traversing all attack types, an intrusion detection sample set is constructed based on the first type of attack samples, the second type of attack samples and the non-attack samples.
  • the intrusion detection sample set includes both attack samples (including the aforementioned first type attack samples and second type attack samples) and non-attack samples.
  • attack samples including the aforementioned first type attack samples and second type attack samples
  • non-attack samples In this way, not only can the sample information in the intrusion detection sample set be more complete, but also when the intrusion detection sample set is used to perform an attack test on the device to be detected, the anti-attack effect of the device to be detected can be more accurately defined based on whether the device to be detected can intercept attack samples and whether it can not intercept non-attack samples.
  • unlabeled samples are unknown samples, that is, samples that cannot be identified as attack samples or non-attack samples under current technical means.
  • unlabeled samples are abnormal samples. For example, when the attack test of the data processing device causes the software and hardware system of the device to be detected to fail, the device to be detected itself may generate some abnormal data. These abnormal data neither meet the characteristics of normal traffic nor the characteristics of attack traffic, but will still be collected by the data processing device.
  • the data processing device can directly The intrusion detection sample set can be discarded directly to save the data volume of the intrusion detection sample set.
  • an intrusion detection sample set can be constructed together according to attack samples, unlabeled samples and non-attack samples, so as to add all samples that actually exist when attacking the device to be detected to the intrusion detection sample set, thereby improving the sample richness in the intrusion detection sample set and facilitating the subsequent marking of unlabeled samples or performing other operations through other analyses.
  • the attack samples and non-attack samples in the intrusion detection sample set occupy the same proportion, for example, the attack samples and non-attack samples each account for 50% of all samples. In this way, by equally dividing the attack samples and non-attack samples, the attack samples and non-attack samples in the intrusion detection sample set can be balanced, which is convenient for subsequent extraction of the same proportion of data for attack testing of the device to be detected, thereby improving the credibility of the attack test results.
  • the samples with a larger proportion can be trimmed so that the proportion of attack samples after trimming is the same as the proportion of non-attack samples.
  • the samples with a larger proportion are usually non-attack samples, also called context data. In this way, by trimming non-attack samples, the balance of attack samples and non-attack samples in the intrusion detection sample set can be maintained.
  • a data processing device obtains a second type of attack sample by applying noise to a first type of attack sample, including: the data processing device first applies noise to the first type of attack sample to obtain a perturbation sample, then inputs the perturbation sample into an attack recognition model, and obtains a recognition result output by the attack recognition model; when the recognition result indicates that it is impossible to determine whether the perturbation sample is an attack sample, the perturbation sample is determined to be a second type of attack sample; otherwise, the perturbation sample is adjusted according to the recognition result, and the adjusted perturbation sample is input into the attack recognition model again, and the above process is repeated until the recognition result corresponding to the adjusted perturbation sample indicates that it is impossible to determine whether the adjusted perturbation sample is an attack sample, and the adjusted perturbation sample is determined to be a second type of attack sample.
  • the second type of attack sample is an unrecognizable sample obtained by adding noise on the basis of the first type of attack sample of known attack type. It can be considered as an attack sample whose attack type cannot be determined under current technical means, that is, an attack sample of unknown attack type. In this way, by adding the attack sample to the intrusion detection sample set, the phenomenon of misreporting these attack samples of unknown attack type as non-attack samples can be avoided in real attack test scenarios.
  • the data processing device constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also extract features from the intrusion detection sample set to obtain an offline detection sample set, which is used to input into the intrusion detection model to evaluate the detection performance of the intrusion detection model according to the detection results output by the intrusion detection model. For example, when more offline detection samples belonging to attack samples in the offline detection sample set are detected as attack samples by the intrusion detection model, and more offline detection samples belonging to non-attack samples are detected as non-attack samples by the intrusion detection model, it means that the detection effect of the intrusion detection model is better.
  • the data processing device can also adjust the parameters of the intrusion detection model, and use the adjusted intrusion detection model to re-detect the offline detection sample set, and repeat the above process until an intrusion detection model with a good detection effect is obtained.
  • the intrusion detection sample set can support the evaluation of the detection effect of the intrusion detection model in an offline state (referred to as offline evaluation), so as to continuously optimize the intrusion detection model according to the detection effect, obtain an intrusion detection model with better detection effect, and improve the implementation effect of the intrusion detection model on the device to be detected.
  • offline evaluation the evaluation of the detection effect of the intrusion detection model in an offline state
  • the data processing device extracts features from the intrusion detection sample set to obtain offline detection samples.
  • the invention discloses a set of intrusion detection samples, comprising: a data processing device first determines the message type of each intrusion detection sample in the intrusion detection sample set, and then, for each intrusion detection sample whose message type is a transmission control protocol (TCP) message, a feature extraction is performed on all intrusion detection samples belonging to the same TCP connection to obtain an offline detection sample, and for each intrusion detection sample whose message type is a CAN (FD) message or a UDP message with a variable data rate, a feature extraction is performed on the intrusion detection sample of each CAN (FD) message or UDP message to obtain an offline detection sample.
  • TCP transmission control protocol
  • the features extracted by offline evaluation may include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.
  • the extracted features may include all the features shown here, while for intrusion detection samples belonging to CAN (FD) messages or UDP messages, they may include timestamp, frequency feature, protocol type, content feature, packet loss rate, and number of error packets.
  • the data processing device can extract all the aforementioned features for each intrusion detection sample. Then, when a certain feature of a certain intrusion detection sample does not exist, the feature of the intrusion detection sample set is configured as a preset character.
  • the preset character can be, for example, a number, letter, symbol, or a combination of one or more of them.
  • the data processing device constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and then input the online detection sample set into the device to be detected through the test tool.
  • the online detection sample set is used to evaluate the detection performance of the device to be detected that is deployed with an intrusion detection model.
  • the intrusion detection sample set can support the online evaluation of the detection effect of the device to be detected that is deployed with the intrusion detection model (referred to as online evaluation), so as to determine the anti-attack performance of the device to be detected based on the detection effect, and ensure that only the device to be detected with good anti-attack effect is shipped out of the factory.
  • online evaluation the online evaluation of the detection effect of the device to be detected that is deployed with the intrusion detection model
  • test tool can be CANoe, PCAN, Technica or other tools that can realize online evaluation
  • intrusion detection sample set after format conversion can be .PCAP, .ASC, .BLF or other formats corresponding to other online test tools.
  • the data processing device constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also determine the evaluation value corresponding to the intrusion detection sample set based on the values of the intrusion detection sample set under various preset indicators, and adjust the intrusion detection sample set when the evaluation value is lower than the preset threshold.
  • the preset indicators can be set according to the characteristics of the system architecture to which the device to be detected belongs.
  • the intrusion detection sample set can be effectively optimized according to the evaluation results, so that the optimized intrusion detection sample set is more suitable for the system architecture to which the device to be detected belongs.
  • the preset indicators may include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, business coverage indicator, data marking indicator, balance indicator, feature independence indicator, and ease of use indicator.
  • data redundancy indicator, attack coverage indicator, protocol coverage indicator, business coverage indicator, data marking indicator and balance indicator are quantitative indicators
  • feature independence indicator and ease of use indicator are qualitative indicators.
  • the value ranges of various preset indicators can be configured to be the same range, such as [0,1].
  • the evaluation value corresponding to the intrusion detection sample set can be the weighted average of the values of the intrusion detection sample set under each preset indicator, wherein the weights corresponding to each preset indicator can be the same or different.
  • each quantitative preset indicator can be configured to correspond to a first weight
  • each qualitative preset indicator can correspond to a second weight
  • the first weight is greater than the second weight.
  • the first type of attack sample can be obtained by attacking any of the following areas: the entire device to be detected; one or more physical areas of the device to be detected; or one or more functional areas of the device to be detected.
  • the one or more physical areas may include one or more of the vehicle trunk area, the left front body area, the right front body area, the left rear body area, or the right rear body area
  • the one or more functional areas may include one or more of the vehicle central control gateway area, the body control area, the cockpit control area, the power control area, the chassis control area, or the infotainment area.
  • the data processing device can construct an intrusion detection sample set for the entire device to be detected, or construct an intrusion detection sample set for one or more physical areas in the device to be detected, or construct an intrusion detection sample set for one or more functional areas in the device to be detected, or, it can also construct an intrusion detection sample set for a combination of one or more physical areas and one or more functional areas. It can be seen that the method for constructing the intrusion detection sample set can be applicable to various different construction scenarios, which helps to improve the flexibility, versatility and ease of use of the intrusion detection sample set.
  • the present application provides a data processing device, which can be any device with processing capabilities, such as a server or a server cluster composed of servers.
  • the data processing device includes: an attack unit, which is used to obtain a first type of attack sample by attacking a device to be detected; a perturbation unit, which is used to obtain a second type of attack sample by applying noise to the first type of attack sample; and a construction unit, which is used to construct an intrusion detection sample set based on the first type of attack sample and the second type of attack sample.
  • the first type of attack samples may include real attack samples and simulated attack samples.
  • the real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.
  • the real attack sample can correspond to one or more of the following attack types: ID non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, etc. Attack, ID source non-specified ECU attack, identical ID attack, CAN scanning attack, UDS sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown out-of-stack connection attack, unknown in-stack connection attack.
  • ID non-existence attack ID non-specified ECU attack
  • identical ID attack identical ID attack
  • CAN scanning attack UDS sensitive operation attack
  • message authentication error attack ECU identity spoofing attack
  • man-in-the-middle attack man-in-the-middle attack
  • ECU authentication error attack brute force attack
  • application layer protocol error attack unknown out-of-stack connection attack, unknown in-stack connection attack.
  • the simulated attack samples may correspond to one or more of the following attack types: ID Fuzz attack, data Fuzz attack, CAN Dos attack, ETH Dos attack, malformed packet injection attack, and port scanning attack.
  • the attack unit is specifically used to: traverse each attack type among a plurality of preset attack types, and when traversing each attack type: execute the attack behavior corresponding to the attack type on the device to be detected, and obtain the traffic data generated by the device to be detected for the attack behavior; if the traffic data is attack traffic, mark the traffic data as a first type attack sample.
  • the attack unit after obtaining the traffic data generated by the device to be detected in response to the attack behavior, the attack unit is also used to: if the traffic data is normal traffic, mark the traffic data as a non-attack sample; correspondingly, the construction unit is specifically used to: construct an intrusion detection sample set based on the first type of attack samples, the second type of attack samples and the non-attack samples.
  • the perturbation unit is specifically used to: apply noise to the first type of attack sample to obtain a perturbation sample, input the perturbation sample into the attack recognition model, obtain the recognition result output by the attack recognition model, adjust the perturbation sample according to the recognition result, until the recognition result corresponding to the adjusted perturbation sample indicates that it is impossible to determine whether the perturbation sample is an attack sample, and then determine the adjusted perturbation sample as a second type of attack sample.
  • the recognition result output by the attack recognition model is used to indicate whether the perturbation sample is an attack sample.
  • the data processing device may further include a feature extraction unit, which is used to: extract features from the intrusion detection sample set to obtain an offline detection sample set, and the offline detection sample set is used to evaluate the detection effect of the intrusion detection model.
  • a feature extraction unit which is used to: extract features from the intrusion detection sample set to obtain an offline detection sample set, and the offline detection sample set is used to evaluate the detection effect of the intrusion detection model.
  • the feature extraction unit is specifically used to: determine the message type of each intrusion detection sample in the intrusion detection sample set, for each intrusion detection sample whose message type is a TCP message, obtain an offline detection sample by performing feature extraction on all intrusion detection samples belonging to the same TCP connection, and for each intrusion detection sample whose message type is a CAN (FD) message or a UDP message, obtain an offline detection sample by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.
  • the features extracted from the aforementioned features include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.
  • the data processing device may also include a format conversion unit, which is used to: convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and input the online detection sample set into the device to be detected through the test tool, wherein the online detection sample set is used to evaluate the detection performance of the device to be detected that is deployed with an intrusion detection model.
  • a format conversion unit which is used to: convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and input the online detection sample set into the device to be detected through the test tool, wherein the online detection sample set is used to evaluate the detection performance of the device to be detected that is deployed with an intrusion detection model.
  • the data processing device may further include an adjustment unit, which is used to determine an evaluation value corresponding to the intrusion detection sample set based on the value of the intrusion detection sample set under each preset indicator, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.
  • an adjustment unit which is used to determine an evaluation value corresponding to the intrusion detection sample set based on the value of the intrusion detection sample set under each preset indicator, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.
  • the preset indicators include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, business coverage indicator, data labeling indicator, balance indicator, feature independence indicator, and ease of use indicator.
  • the first type of attack sample may be obtained by attacking any of the following areas: the entire device to be detected; one or more physical areas of the device to be detected; or one or more functional areas of the device to be detected.
  • the present application provides a data processing device, including a processor, the processor is connected to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the data processing device performs a method as described in any one of the designs of the first aspect above.
  • the present application provides a data processing device, including a processor and a memory, the memory storing computer program instructions, and the processor executing the computer program instructions to implement a method as described in any one of the designs of the first aspect above.
  • the present application provides a data processing device, including a processor, a memory and a transceiver, the memory stores computer program instructions, and the processor runs the computer program instructions to call the transceiver to implement a method as described in any one of the designs in the first aspect above.
  • the present application provides a chip, which may include a processor and an interface, wherein the processor is used to read instructions through the interface to execute a method as described in any one of the designs in the first aspect above.
  • the present application provides a data processing system, which may include a device to be detected and a data processing device, and the data processing device is used to execute the method described in any one of the designs in the first aspect above.
  • the present application provides a computer-readable storage medium storing a computer program.
  • the computer program When the computer program is executed, the method described in any one of the above-mentioned first aspects is implemented.
  • the present application provides a computer program product, which, when executed on a processor, implements a method as described in any one of the above-mentioned first aspects.
  • FIG1 exemplarily shows a possible system architecture diagram provided by an embodiment of the present application
  • FIG2 exemplarily shows a schematic diagram of a partition architecture of a vehicle provided in an embodiment of the present application
  • FIG3 exemplarily shows a flow chart of a data processing method provided in an embodiment of the present application
  • FIG4 exemplarily shows a schematic diagram of a process for obtaining a first type of attack sample provided in an embodiment of the present application
  • FIG5 exemplarily shows a schematic diagram of an application scenario of an intrusion detection sample set provided in an embodiment of the present application
  • FIG6 exemplarily shows a flow chart of evaluating an intrusion detection sample set provided by an embodiment of the present application
  • FIG. 7 exemplarily illustrates a design architecture diagram of a development data processing solution provided in an embodiment of the present application
  • FIG8 exemplarily shows a schematic structural diagram of a data processing device provided in an embodiment of the present application.
  • FIG9 exemplarily shows a schematic structural diagram of another data processing device provided in an embodiment of the present application.
  • the data processing method disclosed in the present application can be used to construct an intrusion detection sample set, which can be used to train an intrusion detection model, and can also be used to detect the detection effect of the intrusion detection model or the device to be detected deployed with the intrusion detection model.
  • the device to be detected can be any terminal device with communication capabilities, and in particular, it can be a terminal device that has certain requirements for network security.
  • the terminal device may include but is not limited to: intelligent transportation equipment, such as cars, ships, drones, trains, vans, trucks, flying cars, etc.; smart home devices, such as TVs, sweeping robots, smart desk lamps, audio systems, smart lighting systems, electrical control systems, home background music, home theater systems, intercom systems, video surveillance, etc.; intelligent manufacturing equipment, Such as robots, industrial equipment, industrial computers, smart logistics, smart factories, etc.
  • the terminal device can also be a computer device, such as a desktop, personal computer, server, etc.
  • the terminal device can also be a portable electronic device, such as a mobile phone, tablet computer, PDA, headset, speaker, wearable device (such as smart watch), vehicle-mounted device, virtual reality device, augmented reality device, etc.
  • portable electronic devices include but are not limited to devices equipped with Or a portable electronic device with other operating systems.
  • the portable electronic device may also be a laptop computer (Laptop) with a touch-sensitive surface (eg, a touch panel).
  • system and “network” in the embodiments of the present application can be used interchangeably.
  • “At least one” means one or more, and “plurality” means two or more.
  • “And/or” describes the association relationship of associated objects, indicating that three relationships may exist.
  • a and/or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.
  • the character “/” generally indicates that the associated objects before and after are in an “or” relationship.
  • At least one of the following” or similar expressions refers to any combination of these items, including any combination of single items or plural items.
  • At least one of a, b, or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, c can be single or multiple.
  • ordinal numbers such as “first” and “second” mentioned in the embodiments of the present application are used to distinguish multiple objects, and are not used to limit the priority or importance of multiple objects.
  • first type of attack sample and the second type of attack sample are only used to distinguish different types of attack samples, and do not indicate the difference in priority or importance of the two types of attack samples.
  • connection can be understood as electrical connection, and the connection between two electrical components can be a direct or indirect connection between the two electrical components.
  • a and B are connected, which can be either A and B directly connected, or A and B indirectly connected through one or more other electrical components, such as A and B are connected, or A and C are directly connected, C and B are directly connected, and A and B are connected through C.
  • connection can also be understood as coupling, such as electromagnetic coupling between two inductors. In short, the connection between A and B enables the transmission of electrical energy between A and B.
  • FIG. 1 exemplarily shows a possible system architecture diagram provided by an embodiment of the present application.
  • the system architecture includes a device to be detected 100 and a data processing device 200.
  • the device to be detected 100 may be a device with certain requirements for network security, such as a vehicle.
  • the data processing device 200 may be any device with data processing capabilities, such as a server or a server cluster composed of multiple servers, or may also be a chip or circuit, such as a chip or circuit arranged in a server or a server cluster.
  • the data processing device 200 may be a cloud server, and the cloud server may be connected to the device to be detected 100 by wireless.
  • a data processing device 200 may be connected to only one device to be detected 100 as shown in FIG. 1, or may be connected to multiple devices to be detected 100 at the same time.
  • the data processing device 200 in the embodiment of the present application may integrate all functions on an independent physical device, or may distribute the functions on multiple independent physical devices, which is not specifically limited in the embodiment of the present application.
  • the system architecture may also include one or more of a database 300, a model training device 400, and a test tool 500, or may also include other devices, such as routing devices, wireless relay devices, wireless backhaul devices, and operation management and maintenance devices.
  • the database 300 may The database 300 is used to store the intrusion detection sample set constructed by the data processing device 200.
  • the database 300 can be a storage unit independent of the data processing device 200, such as a database server, or an internal storage unit of the data processing device 200, such as a cache memory, a random access memory, a register, a main memory or a read-only memory.
  • the model training device 400 can be any device with algorithm development and verification capabilities, such as a model training server.
  • the IDS is deployed in the model training device 400.
  • the IDS is a test system released by the automotive open system architecture (AUTOSAR) for automotive network security. It is one of the most mainstream vehicle-mounted test systems at present, and can train a high-accuracy, high-timeliness and high-robustness intrusion detection model based on limited vehicle-mounted hardware resources and vehicle-mounted storage resources in combination with rules and machine learning.
  • the test tool 500 refers to a tool that can inject test traffic into the device to be detected 100, such as a CAN test tool, a TCP test tool or a UDP test tool.
  • the test tool 500 can present a user interface to the outside. By clicking on the corresponding test command (such as traffic injection type, traffic injection quantity and traffic injection frequency, etc.) on the user interface, the user can drive the test tool 500 to automatically inject test traffic into the device to be tested 100 according to the corresponding test command.
  • the corresponding test command such as traffic injection type, traffic injection quantity and traffic injection frequency, etc.
  • the data processing device 200 when including the device to be detected 100, the data processing device 200, the database 300, the model training device 400 and the test tool 500, can be connected to the device to be detected 100, the database 300, the model training device 400 and the test tool 500 respectively, and the device to be detected 100 can also be connected to the model training device 400 and the test tool 500.
  • the data processing device 200 can construct an intrusion detection sample set for the device to be detected 100, and store the intrusion detection sample set in the database 300.
  • the data processing device 200 can convert the intrusion detection sample set in the database 300 into an offline detection sample set, and can select a part of the offline detection samples to send to the model training device 400, and after the model training device 400 uses the part of the offline detection samples to train the intrusion detection model, the intrusion detection model is deployed in the device to be detected 100, so that the device to be detected 100 can use the intrusion detection model to identify attack traffic.
  • the data processing device 200 can also send another part of the converted offline detection samples to the model training device 400, and the model training device 400 uses the previously trained intrusion detection model to detect the part of the offline detection samples to obtain offline evaluation information, and the data processing device 200 evaluates the detection effect of the intrusion detection model trained by the model training device 400 according to the offline evaluation information.
  • the data processing device 200 can convert the intrusion detection sample set in the database 300 into an online detection sample set, and then input part or all of the online detection samples into the device to be detected 100 through an online tool, and obtain the online evaluation information generated by the device to be detected 100 for the online detection samples, and evaluate the detection effect of the device to be detected 100 deployed with the intrusion detection model according to the online evaluation information.
  • the intrusion detection sample set will not only be used to train the intrusion detection model, but also be used as a test sample to test the quality of the intrusion detection model and the quality of the equipment to be detected with the intrusion detection model deployed.
  • the sample types in the intrusion detection sample set are sufficient, the detection effect of the intrusion detection model trained based on sufficient intrusion detection samples will be better, and then the anti-attack performance of the equipment to be detected with the intrusion detection model deployed will also be better.
  • the reliability of the evaluation results obtained by evaluating the intrusion detection model or the equipment to be detected based on sufficient intrusion detection samples will also be better. Therefore, how to construct an intrusion detection sample set with sufficient sample types is crucial to improving the detection effect of the intrusion detection model, improving the anti-attack effect of the equipment to be detected, improving the reliability of the intrusion detection model, and improving the reliability of the equipment to be detected.
  • intrusion detection sample sets when constructing intrusion detection sample sets, the industry only uses attack tools to simulate some typical intrusion scenarios to attack the detection equipment, resulting in the intrusion detection sample sets only containing typical intrusion scenarios.
  • the intrusion detection samples corresponding to the scene have very limited sample types in the intrusion detection sample set, which is not conducive to improving the detection effect of the intrusion detection model, improving the anti-attack effect of the equipment to be detected, improving the reliability of the intrusion detection model, and improving the reliability of the equipment to be detected, which is also not conducive to the implementation of intrusion detection technology on the equipment to be detected.
  • an embodiment of the present application provides a data processing method for constructing an intrusion detection sample set with richer sample types, so that it can cover both intrusion detection samples in known intrusion scenarios and intrusion detection samples in unknown intrusion scenarios, so as to improve the detection effect of the intrusion detection model trained using the intrusion detection sample set, improve the anti-attack effect of the device to be detected deployed with the intrusion detection model, improve the reliability of evaluating the intrusion detection model using the intrusion detection sample set, and improve the reliability of evaluating the device to be detected using the intrusion detection sample set.
  • FIG2 shows a schematic diagram of a partition architecture of a vehicle provided in the embodiment of the present application, wherein:
  • FIG2 (A) shows a physical partition architecture diagram of a vehicle, which divides the entire vehicle into a vehicle-mounted trunk area, a right front body area, a left front body area, a right rear body area, and a left rear body area according to different physical areas.
  • the right front body area, the left front body area, the right rear body area, and the right rear body area exist as branch areas of the vehicle-mounted trunk area.
  • a vehicle control unit (VCU) is deployed in the vehicle-mounted trunk area, and each branch area is deployed with its own control node (i.e., Z 1 , Z 2 , Z 3 , Z 4 ) and several ECUs connected to the control node. All ECUs in each branch area are connected to the control node via CAN (FD), and the control node in each branch area is connected to the VCU in the vehicle-mounted trunk area via ETH.
  • VCU vehicle control unit
  • FIG2 (B) shows a functional partition architecture diagram of a vehicle.
  • the architecture divides the entire vehicle into a vehicle-mounted central control gateway functional area and K other functional areas according to different functions to be implemented.
  • the K other functional areas exist as branch areas of the vehicle-mounted central control gateway functional area, each of which implements different functions.
  • the K other functional areas may include one or more of the body control domain, cockpit control domain, power control domain, chassis control domain or infotainment domain, etc., and K is a positive integer.
  • a GateWay gateway is deployed in the vehicle-mounted central control gateway functional area, and each other functional area is deployed with its own domain controller (i.e., D 1 , D 2 , ..., D 4K ) and several ECUs connected to the domain controller. All ECUs in each other functional area are connected to the domain controller via CAN (FD), and the domain controller in each other functional area is connected to the GateWay in the vehicle-mounted central control gateway functional area via ETH.
  • D 1 , D 2 , ..., D 4K domain controller
  • All ECUs in each other functional area are connected to the domain controller via CAN (FD), and the domain controller in each other functional area is connected to the GateWay in the vehicle-mounted central control gateway functional area via ETH.
  • an intrusion detection sample set can be constructed for the entire vehicle, an intrusion detection sample set can be constructed for one or more physical areas in the vehicle, an intrusion detection sample set can be constructed for one or more functional areas in the vehicle, or an intrusion detection sample set can be constructed for a combination of one or more physical areas and one or more functional areas, etc.
  • the physical area can be the vehicle trunk area, the right front body area, the left front body area, the right rear body area or the left rear body area
  • the functional area can be the vehicle central control gateway functional area, or other functional areas, without specific limitation.
  • the specific areas for which intrusion detection sample sets are constructed can be set according to the processing capacity of the data processing device and the actual needs of the user. For example, when the processing capacity of the data processing device is strong, you can choose to construct an intrusion detection sample set for the entire vehicle, so as to use the efficient processing capacity to construct a relatively complete global intrusion detection sample set. Intrusion detection sample set. Among them, the global intrusion detection sample set can be applicable to global attack scenarios, and can also be applicable to local attack scenarios, and has good versatility. On the contrary, when the processing capacity of the data processing device is not strong, these areas can be selected specifically according to the areas involved in the actual business to construct local intrusion detection sample sets.
  • the local intrusion detection sample set is a subset of the global intrusion detection sample set. Although it can only be applied to local attack scenarios, the scale of the area it needs to process becomes smaller, so it can better reduce the difficulty of constructing the intrusion detection sample set and reduce the complexity of the intrusion detection sample set.
  • the data processing device can be the data processing device 200 shown in FIG1, or it can be other communication nodes, communication devices or communication systems that can support the data processing device to implement the required functions, such as chips, chip systems, circuits or circuit systems, without specific limitation.
  • FIG3 exemplarily shows a flow chart of a data processing method provided in an embodiment of the present application, and the data processing method can be executed by a data processing device, such as the data processing device 200 shown in FIG1 .
  • the method includes:
  • Step 301 The data processing device obtains a first type of attack sample by attacking a device to be detected.
  • the data processing device can initiate an attack behavior for the entire device to be detected to obtain the first type of attack sample corresponding to the entire device to be detected.
  • the data processing device can only initiate an attack behavior for the one or more regions to obtain the first type of attack sample corresponding to the one or more regions.
  • the device to be detected is a vehicle, which may be a vehicle under a distributed electrical and electronic architecture, a domain centralized electrical and electronic architecture, a vehicle centralized electrical and electronic architecture, or any vehicle-mounted architecture that may appear in the future.
  • the data processing device can construct an intrusion detection sample set corresponding to each vehicle-mounted architecture by attacking vehicles belonging to each vehicle-mounted architecture, thereby facilitating users to select one or more intrusion detection sample sets corresponding to the vehicle-mounted architectures for use according to actual needs, thereby improving the user experience.
  • the data processing device can also attack multiple vehicles belonging to the vehicle-mounted architecture together, so that the first type of attack sample can cover multiple vehicles under the vehicle-mounted architecture, avoiding the problem of inaccurate sample collection due to problems with the vehicle when only one vehicle is attacked.
  • the data processing device when attacking multiple vehicles belonging to a vehicle-mounted architecture, after obtaining a large number of first-type attack samples corresponding to the multiple vehicles, the data processing device can also filter these first-type attack samples by means of clustering or model recognition, for example, filtering out first-type attack samples that are significantly different from other first-type attack samples, and only retaining relatively similar first-type attack samples. In this way, by cleaning up first-type attack samples that are obviously problematic in advance, it is possible to avoid subsequent meaningless processing of first-type attack samples that are obviously problematic, and effectively save the computing resources of the data processing device.
  • the first type of attack samples may include real attack samples and simulated attack samples.
  • the real attack samples are obtained by manually attacking the device to be detected, for example, they can be collected during the penetration test process.
  • the penetration test process refers to the attacker standing from the perspective of a hacker to actually attack the vehicle for the purpose of breaking through the device to be detected.
  • This attack method can construct highly concealed attack behaviors and attack behaviors that rely on business logic, and can obtain attack samples that are easy to mark in a real attack environment.
  • the simulated attack samples are obtained by attacking the device to be detected with an attack tool.
  • the devices can be collected from the device to be detected after simulating the generation of attack traffic using an attack tool and automatically injecting the attack traffic into the device to be detected.
  • This attack method can obtain data that is not easy to mark. Attack samples marked in a real attack environment.
  • the human attack method and the attack tool attack method to comprehensively construct the first type of attack samples, the first type of attack samples can fully cover the attack samples of various attack types that may exist in the real attack environment, thereby improving the comprehensiveness of the first type of attack samples.
  • the first type of attack sample can be obtained by attacking the device to be detected according to various known attack types in the intrusion scenarios known at this stage.
  • the various known attack types include attack types that are easy to collect and mark in a real attack environment (referred to as the first attack type) and attack types that are not easy to collect and mark in a real attack environment (referred to as the second attack type). Therefore, by using artificial methods to actually attack the device to be detected according to the first attack type, the above-mentioned real attack samples can be collected, and by using attack tools to simulate the attack on the device to be detected according to the second attack type, the above-mentioned simulated attack samples can be collected in a centralized manner.
  • each ECU in the vehicle communicates through a CAN (FD) bus and an ETH bus
  • Table 1 shows a corresponding relationship table of possible attack types in the device to be detected and the attack type corresponding to each attack sample.
  • the first attack type may include one or more of the following attack types:
  • ID non-existence attack refers to attacking the vehicle by changing the ID in the CAN message to a non-existent ID
  • Replay attack is to deceive the vehicle by sending CAN messages that the vehicle has received before;
  • Tampering attack refers to attacking the vehicle by tampering with the data carried in the CAN message
  • Data length error attack refers to attacking the vehicle by modifying the data length of the CAN message
  • Signal out-of-range attack means attacking the vehicle by modifying the signal value in the CAN message to a value greater than the maximum specified value or less than the minimum specified value;
  • Context error attack refers to attacking the vehicle by publishing a specific message or signal that is not suitable for a certain state on the CAN network, such as sending an acceleration signal when braking;
  • ID source non-specified ECU attack means attacking the vehicle by publishing CAN messages that should be published by the designated ECU through the non-specified ECU;
  • the same ID attack occurs when the same ID is carried in the CAN messages published by two ECUs to attack the vehicle;
  • CAN scanning attack refers to hacking into the vehicle by sending CAN scanning messages
  • UDS performs sensitive operation attacks, which means attacking vehicles by carrying sensitive information in UDS messages;
  • Message authentication error attack refers to modifying the message authentication process of the CAN bus so that the vehicle cannot successfully authenticate the message
  • ECU identity spoofing attack refers to deceiving the vehicle by changing the ECU identity in the ETH message
  • a man-in-the-middle attack is a method of attacking a vehicle by virtually placing a device controlled by an intruder between two ECUs connected via ETH.
  • ECU authentication error attack refers to modifying the ETH ECU authentication process so that the vehicle cannot successfully ECU;
  • Brute force attack refers to attacking the vehicle by deciphering the ETH key in an exhaustive manner
  • Application layer protocol error attack refers to attacking the vehicle by modifying the application layer protocol so that the vehicle cannot obtain the correct application layer protocol for message interaction;
  • Unknown outbound connection attack refers to attacking the vehicle by giving a fake outbound connection
  • Unknown push connection attack refers to attacking the vehicle by giving a false push connection.
  • ID does not exist attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified ECU attack, identical ID attack, CAN scanning attack, UDS execution sensitive operation attack and message authentication error attack belong to the attack types existing in CAN (FD) communication mode
  • ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack and unknown stack connection attack belong to the attack types existing in ETH communication mode.
  • the second attack type may include one or more of the following attack types:
  • ID fuzzy attack refers to attacking vehicles by fuzzing the ID in the CAN message
  • Data fuzz attack refers to attacking vehicles by fuzzing the data in CAN messages
  • CAN DoS attack refers to attacking the vehicle by stopping the sending and receiving services of a certain part of the CAN bus
  • the DoS attack of ETH refers to attacking the vehicle by stopping the sending and receiving services of a part of the ETH bus;
  • Malformed packet injection attack refers to attacking vehicles by injecting malformed packets
  • a port scan attack is an attempt to break into a vehicle by sending port scan messages.
  • ID fuzzy attack, data fuzz attack and CAN Dos attack belong to the attack types existing in the CAN (FD) communication mode
  • ETH Dos attack, malformed packet injection attack and port scanning attack belong to the attack types existing in the ETH communication mode.
  • the data processing device may attack the device to be detected in a variety of ways to obtain the first type of attack sample.
  • FIG. 4 shows a specific flow diagram of obtaining the first type of attack sample provided in an embodiment of the present application.
  • the flow includes:
  • Step 401 The data processing device obtains multiple preset attack types.
  • the preset multiple attack types may exemplarily include all the attack types shown in Table 1 above, so that the first type attack samples can fully cover various known attack types, thereby improving the richness and comprehensiveness of the first type attack samples.
  • step 402 the data processing device determines whether there is an attack type that has not been traversed among the preset multiple attack types. If so, step 403 is executed; if not, step 409 is executed.
  • Step 403 the data processing device executes an attack behavior corresponding to an attack type that has not been traversed on the device to be detected, and obtains traffic data generated by the device to be detected for the attack behavior.
  • the required attack code can be manually written in advance, and the attack code and the corresponding first attack type can be mapped and stored in the data processing device.
  • the data processing device obtains an attack type that has not been traversed, if it is determined that the attack type belongs to the first attack type, it can directly obtain the attack code corresponding to the first attack type from the local, and automatically generate the corresponding attack behavior according to the attack code and then attack the device to be detected.
  • the data processing device can call the attack tool, and use the attack tool to generate the attack behavior corresponding to the second attack type and then attack the device to be detected.
  • the entire attack process can be automatically implemented by the data processing device, which helps to improve the unified management of the entire attack process, and there is no need to wait for manual on-site programming, which helps to save the delay of sample collection.
  • the data processing device can access the on-board diagnostics (OBD) interface of the device to be detected through a data line, and the OBD interface can be exemplarily an interface of a type II on-board diagnostic system, i.e., an OBD-II interface.
  • OBD on-board diagnostics
  • the data processing device can select an attack type from all attack types in a random manner, in a sequential manner, or in other ways, and then attack the device to be detected according to the attack type, and obtain the flow data generated by the device to be detected for the attack type through the OBD interface.
  • the data processing device can select an attack type from the attack types that have not been traversed to continue attacking the device to be detected, and repeat the above process until all attack types have been traversed.
  • the traffic collection operation of the entire attack process can be automatically implemented by a data processing device.
  • the collection duration can be configured in the data processing device. After the data processing device starts to execute the attack behavior, it can start timing. During the timing time, the data continuously collects the traffic data output by the device to be detected from the OBD interface of the device to be detected until the configured collection duration is reached, and the collection is terminated.
  • the collection duration can be several hours, several days, several weeks or even several months, and can be specifically configured by those skilled in the art according to actual needs. For example, when the number of preset attack types is large, the time required to attack the device to be detected is longer, and the collection time can be configured to be longer, such as several weeks. Conversely, when the number of preset attack types is small, the time required to attack the device to be detected is shorter, and the collection time can be configured to be shorter, such as several days.
  • step 404 the data processing device determines whether the traffic data is attack traffic. If so, step 405 is executed; if not, step 406 is executed.
  • step 405 the data processing device marks the traffic data as a first type attack sample, and then executes step 402 .
  • the data processing device can mark the attack traffic as a real attack sample; conversely, if the traffic data is collected under the second attack type, the data processing device can mark the attack traffic as a simulated attack sample.
  • Step 406 the data processing device determines whether the flow data is normal flow, if so, executes step 407 , if not, executes step 408 .
  • step 407 the data processing device marks the traffic data as a non-attack sample, and then executes step 402 .
  • non-attack samples are also called context data or context samples.
  • step 408 the data processing device determines that the traffic data is an unlabeled sample, and then executes step 402 .
  • the data processing device may not mark the traffic data, but treat it as an unmarked sample.
  • the unmarked sample is an abnormal sample.
  • the attack test of the data processing device causes the hardware and software system failure of the device to be tested, the device to be tested itself may generate some abnormal data. These abnormal data do not meet the characteristics of normal traffic or attack traffic, but will still be collected by the data processing device.
  • Step 409 the data processing device ends the attack process.
  • the data processing device can obtain attack samples, such as real attack samples and simulated attack samples, as well as non-attack samples and even unlabeled samples by attacking the device to be detected.
  • This acquisition method can obtain multiple types of samples, which is convenient for improving the sample richness of the subsequent construction of the intrusion detection sample set.
  • FIG. 4 is only an exemplary introduction to a possible way to obtain the first type of attack sample, and the embodiments of the present application are not limited to only using this method to obtain the first type of attack sample.
  • the data processing device can also combine at least two of the preset multiple attack types, execute the attack behaviors corresponding to at least two attack types on the device to be detected at one time, and then obtain the flow data generated by the device to be detected, separate the flow data corresponding to each of the at least two attack types from the flow data, and obtain the first type of attack sample based on the at least two flow data. It should be understood that there are many possible acquisition methods, which will not be listed one by one here.
  • Step 302 The data processing device obtains a second type of attack sample by applying noise to the first type of attack sample.
  • the data processing device may convert the format of the first type of attack sample to obtain the first type of attack sample in text format or binary format, and store the first type of attack sample in text format or binary format in the original database.
  • the data processing device may traverse all first-type attack samples in the original database, and when traversing each first-type attack sample: applying noise to the first-type attack sample to obtain a perturbation sample, then inputting the perturbation sample into the attack recognition model, and obtaining a recognition result output by the attack recognition model; when the recognition result indicates that it is impossible to determine whether the perturbation sample is an attack sample, the perturbation sample is determined to be a second-type attack sample; otherwise, the perturbation sample is adjusted according to the recognition result, and the adjusted perturbation sample is input into the attack recognition model again, and the above process is repeated continuously. The process is repeated until the identification result corresponding to the adjusted disturbance sample indicates that it is impossible to determine whether the adjusted disturbance sample is an attack sample, and then the adjusted disturbance sample is determined as a second type attack sample.
  • the attack recognition model may be any model with recognition capability, and specifically, may be a neural network model trained by an artificial intelligence (AI) algorithm, such as a generative adversarial network (GAN) algorithm.
  • AI artificial intelligence
  • GAN generative adversarial network
  • the GAN algorithm may learn the features of some known attack samples and non-attack samples, and construct an attack recognition model based on the learned features, so that the attack recognition model can identify the probability that the input sample belongs to the attack sample.
  • the attack recognition model may include a generator and a discriminator, and any first-type attack sample input to the attack recognition model is first received by the generator.
  • the generator may generate a perturbation sample by applying an N-dimensional noise vector corresponding to the N-dimensional feature space to the first-type attack sample, and then input the perturbation sample into the discriminator. After the discriminator identifies the probability that the perturbation sample belongs to the attack sample, if the probability is greater than 50%, the generator is notified to adjust the N-dimensional noise vector in the first direction, and if the probability is less than 50%, the generator is notified to adjust the N-dimensional noise vector in the second direction.
  • the generator uses the adjusted N-dimensional noise vector to re-scramble the first type of attack sample to generate a new perturbation sample, and sends it to the discriminator, which re-identifies the probability that the new perturbation sample belongs to the attack sample. If the probability is 50%, the new perturbation sample can be used as a second type of attack sample. Otherwise, the above process is repeated until a perturbation sample with a probability of 50% being an attack sample is obtained.
  • the attack identification model can generate samples whose sample types cannot be identified.
  • the sample is obtained by scrambling the attack sample (i.e., the real attack sample and the simulated attack sample) obtained by the real attack on the device to be detected, it is highly likely to be an attack sample. Therefore, the sample can be considered as an attack sample whose attack type cannot be determined under the current technical means, and the attack sample is easily misreported in the real identification process.
  • the data processing device can treat the sample as a second type of attack sample and subsequently store it in the intrusion detection sample set, so as to avoid the phenomenon of misreporting these attack samples of unknown attack types as non-attack samples in the real attack test scenario.
  • the second type of attack sample is generated by the confrontation between the generator and the discriminator in the AI algorithm
  • the second type of sample can also be called an AI adversarial sample, or can also have other names, which is not specifically limited in the embodiments of the present application.
  • Step 303 The data processing device constructs an intrusion detection sample set according to the first type attack samples and the second type attack samples.
  • the data processing device obtains the first type of attack samples (including real attack samples and simulated attack samples) and non-attack samples according to the above step 301, and obtains the second type of attack samples according to the above step 302, these samples can be uniformly saved in .pcap format, and then the intrusion detection sample set is constructed based on all samples in .pcap format.
  • the .pcap format is a format that is connected to the existing IDS system. If it is connected to other systems, the data processing device can also save these samples in the format that other systems are connected to. The embodiment of the present application does not specifically limit this.
  • the data processing device can directly discard them to save the amount of data in the intrusion detection sample set, or after traversing all attack types, construct an intrusion detection sample set based on attack samples, unlabeled samples and non-attack samples, so as to add all samples that actually exist when attacking the device to be detected to the intrusion detection sample set, thereby improving the richness of samples in the intrusion detection sample set, and facilitating the subsequent marking of unlabeled samples through other analyses or performing other operations, or in some special cases, they can also be marked as abnormal samples and added to the intrusion detection sample set, without specific limitation.
  • the samples with a larger proportion can be trimmed so that the proportion of the trimmed attack samples and non-attack samples are the same, such as attack samples and non-attack samples each accounting for 50% of all samples.
  • samples with a larger proportion are usually non-attack samples, and may also be attack samples in some special cases.
  • the intrusion detection sample set can cover not only the first type of attack samples in known intrusion scenarios, but also the second type of attack samples in unknown intrusion scenarios, and non-attack samples.
  • the anti-attack effect of the device to be detected can be more accurately defined according to whether the device to be detected can intercept the attack samples and whether it can not intercept the non-attack samples.
  • the above content introduces the specific construction process of the intrusion detection sample set.
  • the following is a detailed introduction to the application of the constructed intrusion detection sample set.
  • FIG5 exemplarily shows a schematic diagram of an application scenario of an intrusion detection sample set provided by an embodiment of the present application.
  • the intrusion detection sample set can be applied to one or more scenarios in a model training scenario, an offline evaluation scenario, or an online evaluation scenario.
  • the solid line in the figure shows the application process of the model training scenario
  • the dotted line in the figure shows the application process of the offline evaluation scenario
  • the double-node line in the figure shows the application process of the online evaluation scenario.
  • model training refers to the use of an intrusion detection sample set to train an intrusion detection model in the early algorithm development and verification. Since the training samples required for the intrusion detection model have their own feature format, and the intrusion detection samples in the intrusion detection sample set may not match the feature format, before training the intrusion detection model, the data processing device may also first extract features from the intrusion detection sample set, and construct an offline detection sample set based on the extracted features. In this way, when it is necessary to train the intrusion detection model, the data processing device can directly select some offline detection samples from the offline detection sample set as training set data, input them to the model training device, and the model training device uses the training set data to train the intrusion detection model.
  • the training set data may include offline detection samples of attack types and offline detection samples of normal types
  • the offline detection samples of attack types are obtained by extracting features from the first type of attack samples and/or the second type of attack samples in the intrusion detection sample set
  • the offline detection samples of normal types are obtained by extracting features from non-attack samples in the intrusion detection sample set
  • the offline detection samples of attack types and normal types can also maintain consistency in quantity, so that a better intrusion detection model can be obtained based on balanced data training.
  • each offline detection sample in the offline detection sample set can be saved in a text format, such as .csv.
  • the .csv format is a format that is connected to the model training device in the existing IDS system. If it is connected to other types of model training devices, it can be saved in a format adapted to other systems, without specific limitation.
  • the data processing device when performing feature extraction, can analyze each intrusion detection sample in isolation, that is, extract features from each intrusion detection sample to obtain an offline detection sample, or can combine multiple intrusion detection samples for centralized analysis, such as extracting features from multiple intrusion detection samples with associated relationships to obtain an offline detection sample.
  • multiple intrusion detection samples with associated relationships can be, for example, multiple intrusion detection samples belonging to the same connection, multiple intrusion detection samples whose traffic data comes from the same ECU, or multiple intrusion detection samples whose traffic data is different from each other. Multiple intrusion detection samples sent to the same ECU, etc.
  • the message types in the vehicle may generally include TCP messages, CAN (FD) messages, and UDP messages, wherein TCP messages are messages transmitted on the connection after a connection is established between at least two ECUs, while CAN (FD) messages and UDP messages are messages sent on the corresponding bus by broadcasting and then obtained by the required ECU node from the bus, and the contents of these three messages include both the source Internet Protocol (IP) address and the destination IP address.
  • TCP messages have the concept of connection
  • CAN (FD) messages and UDP messages do not have the concept of connection.
  • the data processing device can first determine the message type of each intrusion detection sample in the intrusion detection sample set, and then, for each intrusion detection sample whose message type is a TCP message, according to the source IP address and the destination IP address contained in the message content, obtain all intrusion detection samples belonging to the same TCP connection, and extract features from these intrusion detection samples to obtain an offline detection sample.
  • the data processing device can first determine the message type of each intrusion detection sample in the intrusion detection sample set, and then, for each intrusion detection sample whose message type is a TCP message, according to the source IP address and the destination IP address contained in the message content, obtain all intrusion detection samples belonging to the same TCP connection, and extract features from these intrusion detection samples to obtain an offline detection sample.
  • FD CAN
  • UDP UDP
  • the features extracted by the above feature extraction may include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.
  • timestamp refers to the time when the message is collected, including date and time.
  • the frequency feature refers to the communication frequency of sending and receiving messages between the source IP and the destination IP.
  • the protocol type refers to the protocol adapted by the message, such as any of the protocol types in Table 3 above.
  • the content feature refers to the substantive content carried in the message, such as data or instructions.
  • the packet loss rate refers to the proportion of message loss when sending and receiving messages between the source IP and the destination IP.
  • the number of error packets refers to the number of error messages that occur when sending and receiving messages between the source IP and the destination IP.
  • the connection duration refers to the time interval between the establishment and disconnection of a connection.
  • the connection initiator refers to the ECU that requests to establish a connection.
  • the connection receiver refers to the ECU that receives the request sent by the connection initiator.
  • the data processing device can also extract all the above features for each type of intrusion detection sample. Specifically, for intrusion detection samples belonging to TCP messages, since there is a concept of connection, all the above features can be extracted. For intrusion detection samples belonging to CAN (FND) messages or UDP messages, since there is no concept of connection, only the timestamp, frequency feature, protocol type, content feature, packet loss rate and number of error packets in the above features can be extracted, but the connection duration, connection initiator and connection receiver cannot be extracted. In this case, in order to maintain the format consistency of offline detection samples, the data processing device can also configure the features that cannot be extracted as preset characters, which can be, for example, numbers, letters, symbols, or a combination of one or more of them.
  • model training scenario by providing a variety of extractable features, users can choose one or more features according to the application layer requirements in the actual model training scenario to convert the intrusion detection sample set to the offline detection sample set, so as to adapt to different model training scenarios and improve the versatility of the intrusion detection sample set in the model training field.
  • offline evaluation refers to the use of an intrusion detection sample set to test the detection effect of an intrusion detection model trained in a model training scenario in early algorithm development and verification. Since an offline detection sample set has been extracted in the model training scenario, after the intrusion detection model is trained using part of the offline detection samples in the offline detection sample set, the data processing device can also select another part of the offline detection samples from the offline detection sample set as the test set data, input it to the model training device, and the model training device uses the test set data to test the intrusion detection model to obtain offline evaluation information, and then sends the offline evaluation information to the data processing device, which evaluates the detection effect of the intrusion detection model based on the offline evaluation information.
  • the test set data may also include offline detection samples of attack type and offline detection samples of normal type, and the offline detection samples of attack type and the offline detection samples of normal type may also maintain consistency in number.
  • the model training device obtains the detection result of the intrusion detection model for each offline detection sample
  • the total number of correctly identified offline detection samples is obtained by combining the number of offline detection samples of attack type identified as attack samples and the number of offline detection samples of normal type identified as non-attack samples
  • the total number of incorrectly identified offline detection samples is obtained by combining the number of offline detection samples of attack type identified as non-attack samples and the number of offline detection samples of normal type identified as attack samples
  • the total number of correctly identified offline detection samples and the total number of incorrectly identified offline detection samples are carried in the offline evaluation information and sent to the data processing device.
  • the data processing device can determine that the detection effect of the intrusion detection model is better, otherwise, the detection effect is
  • model training device can also directly send the detection result of each offline detection sample as offline evaluation information to the data processing device, and the data processing device will automatically count the total number of offline detection samples that are correctly identified and the total number of offline detection samples that are incorrectly identified to complete the offline evaluation.
  • the model training device can also directly send the detection result of each offline detection sample as offline evaluation information to the data processing device, and the data processing device will automatically count the total number of offline detection samples that are correctly identified and the total number of offline detection samples that are incorrectly identified to complete the offline evaluation.
  • implementation methods which will not be listed here one by one.
  • the intrusion detection sample set can support the evaluation of the detection effect of the intrusion detection model in an offline state. In this way, it is convenient for the model training device to continuously optimize the intrusion detection model according to the detection effect, obtain an intrusion detection model with better detection effect, and provide a basis for the implementation of the intrusion detection model on the device to be detected.
  • online evaluation refers to the use of an intrusion detection sample set to test the detection effect of the device to be detected that is deployed with the intrusion detection model in the later algorithm deployment.
  • the intrusion detection model can have a good detection effect, but the detection effect is only measured on the basis of being separated from the device to be detected, and cannot represent the actual effect after being actually applied to the device to be detected.
  • the intrusion detection model on the device to be detected, and after a real attack on the device to be detected that is deployed with the intrusion detection model, determine the actual detection effect of applying the intrusion detection model in the device to be detected based on the response of the device to be detected.
  • the data processing device can perform format conversion on the intrusion detection samples in the intrusion detection sample set to obtain online detection samples that match the format of the test tool, and then input the online detection samples into the device to be detected through the test tool, and obtain the online evaluation information generated by the intrusion detection model deployed in the device to be detected for the online detection samples, and evaluate the detection performance of the device to be detected deployed with the intrusion detection model according to the online evaluation information.
  • the format conversion can also be performed on a connection basis.
  • these intrusion detection samples are first aggregated to obtain a preliminary online detection sample, and then the format of the preliminary online detection sample is converted to a format that is adapted to the test tool in the current scenario to obtain an online detection sample.
  • the format of each intrusion detection sample can be directly converted to a format that is adapted to the test tool in the current scenario. Convert it into a format adapted by the test tool in the current scenario to obtain a corresponding online detection sample.
  • the test tool can inject different online test samples into the vehicle in real time through the OBD interface of the vehicle. For each online test sample, if the vehicle identifies it as an attack sample, it can alarm the security operations center (SOC) in the cloud through the VCU. Then, after all the online test samples are injected, the SOC combines the number of all online test samples and the alarm record information of the vehicle during this period to determine the total number of online test samples that are correctly identified and the total number of online test samples that are incorrectly identified, and then generates online evaluation information based on these two total numbers and sends it to the data processing device.
  • SOC security operations center
  • the data processing device can determine that the detection effect of the device to be tested with the intrusion detection model deployed is better, otherwise, the detection effect is worse.
  • the test tools in the embodiments of the present application may specifically be CANoe, PCAN, Technica or other tools that can implement online testing.
  • the online detection sample after format conversion may be a format corresponding to .PCAP, .ASC, .BLF or other tools that can implement online testing.
  • the intrusion detection sample set can support the online evaluation of the detection effect of the device to be tested that is deployed with the intrusion detection model. In this way, it is convenient for users to determine the anti-attack performance of the device to be tested based on the detection effect, ensuring that only the devices to be tested with good anti-attack effects are shipped out of the factory.
  • the data processing device constructs and applies the intrusion detection sample set.
  • the data processing device can also evaluate the quality of the constructed intrusion detection sample set. The following exemplarily introduces a specific evaluation process.
  • FIG. 6 is a schematic diagram of a process of evaluating an intrusion detection sample set provided in an embodiment of the present application.
  • the process includes:
  • Step 601 A data processing device obtains an intrusion detection sample set.
  • Step 602 The data processing device calculates the value of the intrusion detection sample set under each preset indicator, and calculates the evaluation value corresponding to the intrusion detection sample set according to the value of the intrusion detection sample set under each preset indicator.
  • each preset indicator can be set according to the characteristics of the system architecture to which the device to be tested belongs, and can exemplarily include quantitative indicators and qualitative indicators.
  • Quantitative indicators refer to evaluation indicators that can be defined by accurate quantities
  • qualitative indicators refer to evaluation indicators that cannot be directly quantified and need to be quantified through other means.
  • Table 2 shows a schematic table of possible preset indicators provided in an embodiment of the present application:
  • each preset indicator may include one or more of a data redundancy indicator, an attack coverage indicator, a protocol coverage indicator, a service coverage indicator, a balance indicator, a feature independence indicator, and an ease of use indicator.
  • the data redundancy indicator, the attack coverage indicator, the protocol coverage indicator, the service coverage indicator, and the balance indicator are quantitative indicators
  • the feature independence indicator and the ease of use indicator are qualitative indicators.
  • the value range of each preset indicator can also be consistent, for example, all are set to [0,1], that is, the value of each preset indicator can be any real number between 0 and 1, including 0 and 1.
  • the data redundancy index is used to indicate the non-redundancy degree of the intrusion detection sample set, and can be expressed as the ratio of the number of non-redundant intrusion detection samples in the intrusion detection sample set to the number of all intrusion detection samples. For example, when there are 100 intrusion detection samples in the intrusion detection sample set, if 5 of the intrusion detection samples correspond to the same ID, and the remaining 95 intrusion detection samples correspond to different IDs, then the value of the intrusion detection sample set under the data redundancy index is 95/100.
  • the data redundancy index can also be expressed in other forms, as long as it can ensure that it is negatively correlated with the ratio of the number of redundant intrusion detection samples to the number of all intrusion detection samples. In this way, when all intrusion detection samples are different, the value of the data redundancy index is the largest, the number of valid samples in the intrusion detection sample set is the largest, and the adequacy of the samples is the best. As the number of identical intrusion detection samples increases, the value of the data redundancy index gradually decreases, the number of valid samples in the intrusion detection sample set gradually decreases, and the adequacy of the samples gradually deteriorates. Until all intrusion detection samples are the same, the value of the data redundancy index is the smallest, the number of valid samples in the intrusion detection sample set is the smallest, and the adequacy of the samples is the worst.
  • the attack coverage index is used to indicate the coverage of the attack types included in the intrusion detection sample set, and can be expressed as the ratio of the number of attack types covered by the intrusion detection sample set to the number of all attack types that may exist in the device to be detected. For example, assuming that all the possible attack types in the device to be detected are the 25 attack types shown in Table 1, and the first type of attack sample in the intrusion detection sample set is obtained by attacking the device to be detected according to 10 of the attack types, then the value of the intrusion detection sample set under the attack coverage index is 10/25.
  • the attack coverage index can also be expressed in other forms, as long as it can ensure a positive correlation with the ratio of the number of covered attack types to the number of all possible attack types.
  • the value of the attack coverage index is the largest, the sample types in the intrusion detection sample set are the largest, and the sample diversity is the best.
  • the value of the attack coverage index gradually decreases, the sample types in the intrusion detection sample set gradually decreases, and the sample diversity gradually deteriorates.
  • the value of the attack coverage index is the smallest, the sample types in the intrusion detection sample set are the smallest, and the sample diversity is the worst.
  • the protocol coverage index is used to indicate the coverage degree of the communication protocols included in the intrusion detection sample set, and can be expressed as the ratio of the number of communication protocols covered by the intrusion detection sample set to the number of all communication protocols that may exist in the device to be detected.
  • Table 3 is a schematic table of communication protocols that may exist in the field of Internet of Vehicles provided in an embodiment of the present application. It should be understood that with the development of Internet of Vehicles technology, new communication protocols may appear in the future, so the communication protocols in Table 3 can also be updated accordingly, and the embodiment of the present application does not specifically limit this.
  • the device to be detected is a vehicle
  • all possible communication protocols in the vehicle are the eight protocols shown in Table 3, and the communication protocols covered by the intrusion detection sample set include CAN (FD), DoCAN, DDS, MQTT and HTTP (S), then the value of the intrusion detection sample set under the protocol coverage index is 5/8.
  • the protocol coverage index can also be expressed in other forms, as long as it can ensure that the ratio of the number of covered communication protocols to the number of all possible communication protocols is positively correlated.
  • the value of the protocol coverage index is the largest, and the samples in the intrusion detection sample set are obtained by attacking all communication protocol messages.
  • the intrusion detection sample set has the most sample sources and the richness of the samples is the best.
  • the value of the protocol coverage index gradually decreases, the source of samples in the intrusion detection sample set gradually decreases, and the richness of samples gradually deteriorates.
  • the value of the protocol coverage index is the smallest, the source of samples in the intrusion detection sample set is the least, and the richness of samples is the worst.
  • the service coverage index is used to indicate the coverage of the services included in the intrusion detection sample set, and can be expressed as the ratio of the number of services included in the intrusion detection sample set to the number of all services that may exist in the device to be detected. For example, assuming that all services that may exist in the device to be detected include remote control, log transmission, over-the-air (OTA) update, diagnostic service, video transmission and network management, and the services included in the intrusion detection sample set are remote control, OTA update and diagnostic service, then the value of the intrusion detection sample set under the service coverage index is 3/6.
  • OTA over-the-air
  • the service coverage index can also be expressed in other forms, as long as it can ensure a positive correlation with the ratio of the number of covered services to the number of all possible services.
  • the value of the service coverage index is the largest, and the service applicability of the sample is the best.
  • the value of the service coverage index gradually decreases, and the service applicability of the intrusion detection sample set gradually deteriorates, until no service is covered, the value of the protocol coverage index is the smallest, and the service applicability of the intrusion detection sample set is the worst.
  • the data labeling index is used to indicate the labeling degree of the intrusion detection samples in the intrusion detection sample set, and can be expressed exemplarily as the ratio of the number of labeled intrusion detection samples in the intrusion detection sample set to the number of all intrusion detection samples.
  • the labeled intrusion detection samples may include the real attack samples, simulated attack samples and non-attack samples obtained in the above step 301 and the second type of attack samples obtained in the above step 302, and the unlabeled intrusion detection samples include the unlabeled samples obtained in the above step 301.
  • the intrusion detection sample set when there are 100 intrusion detection samples in the intrusion detection sample set, if 20 of the intrusion detection samples are first type attack samples, 20 of the intrusion detection samples are second type attack samples, 55 of the intrusion detection samples are non-attack samples, and 5 of the intrusion detection samples are unlabeled. Samples, the number of labeled intrusion detection samples in the intrusion detection sample set is 95, so the value of the intrusion detection sample set under the data labeling indicator is 95/100.
  • the data labeling index can also be expressed in other forms, as long as it can ensure that the ratio of the number of labeled intrusion detection samples to the total number of all intrusion detection samples is positively correlated. In this way, when all intrusion detection samples in the intrusion detection sample set are labeled, it means that all intrusion detection samples have been clearly divided into non-attack samples and attack samples, and do not contain uncertain samples, and the samples in the intrusion detection sample set have the highest clarity.
  • the number of labeled intrusion detection samples decreases, the number of uncertain samples contained in the intrusion detection sample set gradually increases, and the clarity of the samples gradually deteriorates, until all intrusion detection samples are unlabeled, the intrusion detection sample set contains the most uncertain samples, and the clarity of the samples is the worst.
  • the balance index is used to indicate the degree of balance between attack samples and non-attack samples in the intrusion detection sample set, and can be expressed as the ratio of the difference between the total number of all intrusion detection samples in the intrusion detection sample set and the difference between the number of attack samples and non-attack samples to the total number of all intrusion detection samples. For example, when there are 100 intrusion detection samples in the intrusion detection sample set, if 45 of them are attack samples and 55 are non-attack samples, the difference between the number of attack samples and non-attack samples in the intrusion detection sample set is 10, so the value of the balance index of the intrusion detection sample set can be (100-10)/100.
  • the balance index can also be expressed in other forms, as long as it can ensure a negative correlation with the difference in the number of attack samples and non-attack samples.
  • the value of the balance index is the largest, the sample balance in the intrusion detection sample set is the best, and it is easier to take out equal amounts of attack samples and non-attack samples from the intrusion detection sample set for testing later.
  • the sample balance in the intrusion detection sample set gradually deteriorates, and it becomes more difficult to take out equal amounts of attack samples and non-attack samples from the intrusion detection sample set for testing.
  • the sample balance in the intrusion detection sample set is the worst, and it is impossible to take out equal amounts of attack samples and non-attack samples from the intrusion detection sample set for testing.
  • the feature independence index is used to indicate the independence of the features extracted when using the intrusion detection sample set for offline evaluation. The more independent features there are, the larger the value of the feature independence index is, and the fewer independent features there are, the smaller the value of the feature independence index is.
  • whether a feature is independent can be judged by technical personnel in this field based on experience, for example, it can be judged by at least two of engineers, experts or third-party organizations, so as to obtain a more accurate evaluation result by combining the experience of all parties.
  • the usability index is used to indicate the versatility of application scenarios that the intrusion detection sample set is compatible with.
  • the application scenario can be, for example, the format of the test tools that can be supported when the intrusion detection sample set is used for online evaluation.
  • the application scenarios that the intrusion detection sample set is compatible with can also be judged by technical personnel in this field based on experience, for example, by at least two of engineers, experts or third-party organizations, so as to obtain a more accurate evaluation result by combining the experience of all parties.
  • the intrusion detection sample set is calculated in each After calculating the values under the preset indicators, the data processing device can also perform weighted averaging on the values of the intrusion detection sample set under each preset indicator according to the weights corresponding to each preset indicator, and use the calculated weighted average as the evaluation value of the intrusion detection sample set.
  • the weights corresponding to each preset indicator can be the same or different.
  • each quantitative preset indicator can be configured to correspond to a first weight
  • each qualitative preset indicator can correspond to a second weight
  • the first weight is greater than the second weight.
  • step 603 the data processing device determines whether the evaluation value corresponding to the intrusion detection sample set is lower than a preset threshold value. If so, step 604 is executed; if not, step 605 is executed.
  • step 604 the data processing device adjusts the intrusion detection sample set, and then executes step 602 .
  • the data processing device can adjust the intrusion detection sample set according to the value of the intrusion detection sample set under each preset index, so that the value of the adjusted intrusion detection sample set under one or more preset indicators becomes larger, thereby increasing the evaluation value corresponding to the intrusion detection sample set. For example, when the value of the intrusion detection sample set under the balance index is low, the data processing device can increase the value of the intrusion detection sample set under the balance index by cutting a large number of samples so that the attack samples and the non-attack samples are close to each other.
  • the data processing device can reduce the ratio of the number of redundant intrusion detection samples to the number of all intrusion detection samples by deleting redundant intrusion detection samples, so as to increase the value of the intrusion detection sample set under the data redundancy index.
  • Step 605 The data processing device stores the intrusion detection sample set.
  • the data processing device can store the intrusion detection sample set in a database in a text format.
  • the text format can be .cpap or other formats supported by IDS.
  • the intrusion detection sample set can be continuously optimized according to the evaluation results, so that the optimized intrusion detection sample set is more suitable for the system architecture to which the device to be detected belongs. Moreover, by combining quantitative indicators and qualitative indicators to comprehensively judge the construction quality of the intrusion detection sample set, the evaluation results can also be made more comprehensive and more convincing.
  • evaluating the quality of the intrusion detection sample set by preset indicators is only an optional evaluation method. In actual operation, there may be other evaluation methods, such as direct evaluation through human experience, or indirect evaluation through a third-party organization, or evaluation by comparing historical intrusion detection sample sets, etc.
  • evaluation methods such as direct evaluation through human experience, or indirect evaluation through a third-party organization, or evaluation by comparing historical intrusion detection sample sets, etc.
  • the embodiments of the present application do not make specific limitations on this.
  • FIG7 exemplarily illustrates a design architecture diagram for developing the above data processing solution provided by an embodiment of the present application.
  • the design architecture diagram exemplarily can be an interface presented to development testers or operation and maintenance personnel, and the development testers or operation and maintenance personnel write corresponding program codes according to the various functions that need to be implemented in the interface.
  • the design architecture includes a hardware tool layer, a data generation layer, a call interface layer, and an application scenario layer. The content of each layer is described in detail below.
  • the hardware tool layer is responsible for providing hardware interface tools, which are used to access the device to be detected and provide support for collecting the flow data of the device to be detected.
  • the hardware tool layer can mainly include a signal-oriented CAN (FD) tool and a service-oriented ETH tool. These two tools can be connected to the OBD interface of the vehicle to support the data acquisition module to collect the flow data generated by the vehicle for the attack behavior.
  • FD signal-oriented CAN
  • ETH service-oriented ETH tool
  • the data generation layer is responsible for constructing and viewing the intrusion detection sample set, which mainly includes a data acquisition module, an AI generated sample module, a feature extraction module, a format conversion module, a sample set evaluation module and a data viewing module, and may also include a database.
  • the data acquisition module can collect the traffic data generated by the device to be detected for the attack behavior through the hardware tool layer, and by analyzing the traffic data, mark the first type of attack samples and non-attack samples, and store the first type of attack samples and non-attack samples in the database.
  • the AI generated sample module can add noise to the first type of attack samples marked by the data acquisition module, and after obtaining the second type of attack samples that are easily misreported through the AI adversarial algorithm, store the second type of attack samples in the database.
  • the first type of attack samples, the second type of attack samples and non-attack samples in the database constitute the intrusion detection sample set.
  • the feature extraction module can extract features from the intrusion detection sample set in the database to form an offline detection sample set
  • the format conversion module can convert the format of the intrusion detection sample set in the database to form an online detection sample set.
  • the sample set evaluation module can calculate the evaluation value of the intrusion detection sample set under each preset indicator, and when the evaluation value is lower than the preset threshold, adjust the intrusion detection sample set until the evaluation value corresponding to the intrusion detection sample set is adjusted to not lower than the preset threshold.
  • the data viewing module can display some or all intrusion detection samples to the development and testing personnel or operation and maintenance personnel according to their commands.
  • the calling interface layer is responsible for providing the application programming interface (API). Specifically, it can call the corresponding data from the data generation layer through the API and provide it to the upper-layer application according to the usage requirements of the upper-layer application.
  • the calling interface layer can provide part of the offline detection samples in the offline detection sample set in the data generation layer to the upper-layer application through the API to train the intrusion detection model, or provide another part of the offline detection samples in the offline detection sample set in the data generation layer to the upper-layer application through the API to realize the offline evaluation of the intrusion detection model, and can also provide the online detection samples in the data generation layer to the upper-layer application through the API to realize the online evaluation of the device to be detected that is deployed with the intrusion detection model.
  • API application programming interface
  • the application scenario layer is responsible for interacting with external devices to apply the intrusion detection sample set to various possible application scenarios.
  • the application scenario layer can use the offline detection samples provided by the calling interface layer to train the intrusion detection model in IDS development, and can also use the online detection samples provided by the calling interface layer to evaluate the detection effect of the equipment to be detected with the intrusion detection model deployed in the vehicle-cloud operation and maintenance, and can also use the offline detection samples provided by the calling interface layer to evaluate the detection effect of the intrusion detection model in the IDS test, and can also provide the intrusion detection sample set provided by the calling interface layer to a third-party testing and certification agency, so that the equipment to be detected can be successfully shipped after obtaining the certification of the third-party certification agency, etc.
  • the sample types in the intrusion detection sample set can be made richer and more comprehensive.
  • this construction method can also construct adaptive and rich intrusion detection samples for each type of vehicle model, effectively improving the accuracy of using intrusion detection samples to evaluate vehicle network security.
  • the data processing method provided by this application can also be extended to any information system that has a demand for network security.
  • it can also be applied in the field of smart home, by attacking smart home products and scrambling to obtain rich attack samples, and studying the attacker portrait in the smart home scenario.
  • it can also be applied in the field of industrial control, By introducing attack samples into the digital twin model to study intrusion defense, the robustness of the industrial control system can be enhanced.
  • a vehicle or a component to be tested on a vehicle can be virtually constructed through a cloud server, and then the above data processing operation can be performed on the virtual vehicle or vehicle component to obtain an intrusion detection sample set.
  • the anti-attack capability of the vehicle can be known in advance before the actual assembly of the vehicle, so that the vehicle can be actually built only when it is determined that the vehicle can better defend against attacks, effectively saving manpower and material costs.
  • the above mainly introduces the solution provided by the present application from the perspective of the interaction between various network elements.
  • the above-mentioned network elements include hardware structures and/or software modules corresponding to the execution of various functions.
  • the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
  • FIG8 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application, and the data processing device can be any device with processing capabilities, such as a server, or can also be a chip or circuit, such as a chip or circuit that can be set in a server, or can also be a server cluster composed of multiple servers.
  • the data processing device 800 may include a processor 801, a memory 802, and a transceiver 803, and may further include a bus system, and the processor 801, the memory 802, and the transceiver 803 may be connected through the bus system.
  • each step of the above method can be completed by an integrated logic circuit of hardware in the processor 801 or an instruction in the form of software.
  • the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor 801.
  • the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc.
  • the storage medium is located in the memory 802, and the processor 801 reads the information in the memory 802 and completes the steps of the above method in combination with its hardware.
  • the processor 801 can be a chip.
  • the processor 801 can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD) or other integrated chips.
  • FPGA field programmable gate array
  • ASIC application specific integrated circuit
  • SoC system on chip
  • CPU central processor unit
  • NP network processor
  • DSP digital signal processor
  • MCU microcontroller unit
  • PLD programmable logic device
  • the memory 802 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
  • the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
  • the volatile memory can be a random access memory (RAM), which is used as an external cache.
  • RAM random access memory
  • many forms of RAM are available, such as static RAM.
  • SRAM Static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDR SDRAM double data rate SDRAM
  • ESDRAM enhanced SDRAM
  • SLDRAM synchlink DRAM
  • DR RAM direct rambus RAM
  • the memory 802 is used to store instructions
  • the processor 801 is used to execute the instructions stored in the memory 802 to implement the method corresponding to the data processing device in any one or more of the above Figures 3, 4, or 6.
  • the processor 801 attacks the device to be detected by calling the transceiver 803 to obtain a first type of attack sample, obtains a second type of attack sample by applying noise to the first type of attack sample, and then constructs an intrusion detection sample set based on the first type of attack sample and the second type of attack sample.
  • the first type of attack samples may include real attack samples and simulated attack samples.
  • the real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.
  • the real attack sample may correspond to one or more of the following attack types: identity identification number ID non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit ECU attack, identical ID attack, controller area network CAN scan attack, unified diagnostic service UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, unknown stack connection attack.
  • attack types identity identification number ID non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit ECU attack, identical ID attack, controller area network CAN scan attack, unified diagnostic service UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error
  • the simulated attack samples may correspond to one or more of the following attack types: ID Fuzz attack, data Fuzz attack, Dos attack on CAN, Dos attack on ETH, malformed packet injection attack, and port scanning attack.
  • the processor 801 is specifically used to: traverse each attack type of a preset plurality of attack types, and when traversing each attack type: call the transceiver 803 to execute the attack behavior corresponding to the attack type on the device to be detected, and obtain the traffic data generated by the device to be detected for the attack behavior, if the traffic data is attack traffic, mark the traffic data as a first type attack sample.
  • the processor 801 determines that the traffic data is normal traffic, it marks the traffic data as a non-attack sample, and then constructs an intrusion detection sample set based on the first type of attack samples, the second type of attack samples and the non-attack samples.
  • the processor 801 is specifically configured to: apply noise to the first type of attack sample to obtain a disturbance sample, input the disturbance sample into the attack recognition model, obtain a recognition result output by the attack recognition model, and then adjust the disturbance sample according to the recognition result, until the recognition result corresponding to the adjusted disturbance sample indicates that it is impossible to determine whether the adjusted disturbance sample is an attack sample, and then determine the adjusted disturbance sample as a second type of attack sample.
  • the recognition result is used to indicate whether the disturbance sample is an attack sample.
  • the processor 801 may also perform feature extraction on the intrusion detection sample set to obtain an offline detection sample set, which is used to evaluate the detection performance of the intrusion detection model.
  • the processor 801 is specifically configured to: determine the message type of each intrusion detection sample in the intrusion detection sample set, and for each intrusion detection sample whose message type is a TCP message, Feature extraction is performed on all intrusion detection samples of TCP connections to obtain an offline detection sample. For each intrusion detection sample whose message type is CAN (FD) message or UDP message, feature extraction is performed on the intrusion detection sample of each CAN (FD) message or UDP message to obtain an offline detection sample.
  • FD CAN
  • UDP UDP
  • the extracted features include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.
  • the processor 801 after the processor 801 constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and then input the online detection sample set into the device to be detected through the test tool.
  • the online detection sample is used to evaluate the detection performance of the device to be detected that is deployed with an intrusion detection model.
  • the processor 801 can also determine an evaluation value corresponding to the intrusion detection sample set based on the values of the intrusion detection sample set under various preset indicators, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.
  • the preset indicators may include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, service coverage indicator, data labeling indicator, balance indicator, feature independence indicator, and ease of use indicator.
  • the first type of attack sample may be obtained by attacking any of the following areas: the entire device to be detected; one or more physical areas of the device to be detected; or one or more functional areas of the device to be detected.
  • FIG9 is a schematic diagram of another data processing device provided in an embodiment of the present application.
  • the data processing device 900 may be exemplarily a data processing device as described in any of the above embodiments, or may be a chip or circuit, such as a chip or circuit that can be arranged in a data processing device.
  • the data processing device 900 may implement the steps performed by the data processing device in any one or more of the corresponding methods shown in FIG3, FIG4 or FIG6 above.
  • the data processing device 900 may include an attack unit 901, a perturbation unit 902, and a construction unit 903, and may also include one or more of a feature extraction unit 904, a format conversion unit 905, and an adjustment unit 906.
  • the attack unit 901 is used to obtain a first type of attack sample by attacking the device to be detected;
  • the perturbation unit 902 is used to obtain a second type of attack sample by applying noise to the first type of attack sample;
  • the construction unit 903 is used to construct an intrusion detection sample set according to the first type of attack sample and the second type of attack sample.
  • the feature extraction unit 904 is used to: extract features from the intrusion detection sample set to obtain an offline detection sample set, and the offline detection sample set is used to evaluate the detection performance of the intrusion detection model.
  • the format conversion unit 905 is used to: convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and input the online detection sample set into the device to be detected through the test tool, and the online detection sample set is used to evaluate the detection performance of the device to be detected with the intrusion detection model deployed.
  • the adjustment unit 906 is used to determine an evaluation value corresponding to the intrusion detection sample set according to the value of the intrusion detection sample set under various preset indicators, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.
  • the attack unit 901 injects traffic into the device to be detected to attack the device to be detected.
  • the attack unit may be a sending unit, a transmitter, an output interface, a pin or a circuit when it is such as traffic.
  • the storage unit is used to store computer instructions
  • the attack unit 901, the perturbation unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905 and the adjustment unit 906 are respectively connected to the storage unit for communication, and respectively execute the computer instructions stored in the storage unit, so that the data processing device 900 can be used to execute the method executed by the data processing device in any of the above embodiments.
  • the attack unit 901, the perturbation unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905 and the adjustment unit 906 can be a general-purpose central processing unit (CPU), a microprocessor, and an application specific integrated circuit (ASIC).
  • the storage unit is a storage unit within the chip, such as a register, a cache, etc.
  • the storage unit can also be a storage unit within the data processing device 900 that is located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
  • ROM read-only memory
  • RAM random access memory
  • the division of the units of the above data processing device 900 is only a division of logical functions, and in actual implementation, all or part of them can be integrated into one physical entity, or they can be physically separated.
  • the attack unit 901, the disturbance unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905 and the adjustment unit 906 can be implemented by the processor 801 of Figure 8 above.
  • the present application also provides a data processing device, which includes a processor, the processor is connected to a memory, the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the data processing device implements the method described in any one of the embodiments in Figures 3, 4 or 6.
  • the present application also provides a data processing device, which includes a processor and a memory, the memory is used to store computer program instructions, and the processor is used to run the computer program instructions to implement the method described in any one of the embodiments in Figures 3, 4 or 6.
  • the present application also provides a chip, which may include a processor and an interface, and the processor is used to read instructions through the interface to execute the method described in any of the embodiments in Figures 3, 4 or 6.
  • the present application also provides a data processing system, which may include the aforementioned device to be detected and a data processing apparatus.
  • the present application also provides a computer-readable storage medium, which stores a computer program.
  • a computer program When the computer program is executed, the method described in any one of the embodiments in Figures 3, 4 or 6 is implemented.
  • the present application also provides a computer program product, which, when executed on a processor, implements the method described in any one of the embodiments shown in FIG. 3 , FIG. 4 or FIG. 6 .
  • a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program and/or a computer.
  • applications running on a computing device and a computing device can be components.
  • One or more components may reside in a process and/or an execution thread, and a component may be located on one computer and/or distributed between two or more computers.
  • these components may be executed from various computer-readable media having various data structures stored thereon.
  • a component may, for example, be based on a program or a process having one or more data packets (e.g., from interacting with another component in a local system, a distributed system and/or a network). Data between two components is communicated through local and/or remote processes, such as the Internet (for example, via signals to interact with other systems).
  • the disclosed systems, devices and methods can be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
  • Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
  • the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
  • the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application.
  • the aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Security & Cryptography (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

一种数据处理方法、装置、存储介质及程序产品,适用于网络安全技术领域,用以构建更加全面的入侵检测样本集。其中,方法包括:通过攻击待检测设备,获得第一类型攻击样本,通过对第一类型攻击样本施加噪声,获得第二类型攻击样本,根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。通过该方法,入侵检测样本集中既能覆盖已知入侵场景下的攻击样本又能覆盖未知入侵场景下的攻击样本,该入侵检测样本集中的入侵检测样本更加全面,进而,使用更加全面的入侵检测样本集测评待检测设备,也能提高对待检测设备进行安全评估的准确性。

Description

一种数据处理方法、装置、存储介质及程序产品 技术领域
本申请涉及网络安全技术领域,提供了一种数据处理方法、装置、存储介质及程序产品。
背景技术
近年来,汽车网络安全攻击事件频发。据了解,仅在2021年,全球范围内公开报道的汽车网络安全攻击事件就高达256件,相较2018年增加了225%。面对如今智能汽车快速发展的大趋势,汽车网络安全问题也应同样被重视起来。
为了提高汽车的网络安全性,业界针对于某些典型的入侵场景构建了入侵检测样本,并在汽车出厂之前,使用这些入侵检测样本对汽车进行攻击测试,以评估汽车的防攻击性能。虽然这种方式能确保仅出厂防攻击性能较好的汽车,但由于入侵检测样本仅是针对于某些典型的入侵场景构建得到的,因此,该入侵检测样本显然并不全面,不利于提高对汽车进行网络安全评估的准确性。
因此,目前对于汽车入侵检测样本的构建方面,还有待进一步研究。
发明内容
本申请提供一种数据处理方法、装置、存储介质及程序产品,用以构建更加全面的入侵检测样本集,以提高对待检测设备(如车辆)进行网络安全评估的准确性。
第一方面,本申请提供一种数据处理方法,适用于数据处理装置,该数据处理装置可以是具有处理能力的任意设备,如服务器或由多个服务器构成的服务器集群。该方法包括:数据处理装置通过攻击待检测设备,获得第一类型攻击样本,通过对第一类型攻击样本施加噪声,获得第二类型攻击样本,进而根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。
在上述设计中,第一类型攻击样本是通过攻击待检测设备而得到的,属于已知入侵场景下的攻击样本,而第二类型攻击样本是通过对第一类型攻击样本施加噪声而得到的,可以认为是通过对已知入侵场景下的攻击样本进行一些变形而得到的未知入侵场景下的攻击样本,如此,结合第一类型攻击样本和第二类型攻击样本构建入侵检测样本集,使得入侵检测样本集中既能覆盖已知入侵场景下的攻击样本,又能覆盖未知入侵场景下的攻击样本,从而该入侵检测样本集中的入侵检测样本更加全面,进而,使用更加全面的入侵检测样本集评估待检测设备,也能提高对待检测设备进行网络安全评估的准确性。
一种可能的设计中,第一类型攻击样本可以包括真实攻击样本和模拟攻击样本,其中,真实攻击样本是通过人为攻击待检测设备而得到的,模拟攻击样本是通过攻击工具攻击待检测设备而得到的。
在上述设计中,人为攻击方式能用于构造隐蔽性较强的或依赖业务逻辑的攻击行为,以获取到易于在真实攻击环境中标记的攻击样本,而攻击工具攻击方式则能通过模拟攻击流量的产生,获取到不易在真实攻击环境中标记的攻击样本。如此,通过结合人为攻击方式和攻击工具攻击方式综合构建第一类型攻击样本,能使第一类型攻击样本充分覆盖真实攻击环境中可能存在的各种攻击类型的攻击样本,提高第一类型攻击样本的全面性。
进一步地设计中,真实攻击样本可以对应如下攻击类型中的一项或多项:身份标识号(identity document,ID)不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻击、ID来源非指定电子控制单元(electronic control unit,ECU)攻击、出现相同ID攻击、控制器局域网络(controller area network,CAN)扫描攻击、统一诊断服务(unified diagnostic services,UDS)执行敏感操作攻击、消息认证错误攻击、ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击、未知的入栈连接攻击。
在上述设计中,通过给出易于在真实攻击环境中进行标记的多种攻击类型,能便于用户根据实际需求选择其中一种或多种攻击类型来构建真实攻击样本,以适用于不同的真实攻击场景,提高构建真实攻击样本的灵活性和通用性。
进一步地设计中,模拟攻击样本可以对应如下攻击类型中的一项或多项:ID模糊(Fuzz)攻击、数据Fuzz攻击、CAN的拒绝服务(denial-of-service,Dos)攻击、以太网(ethnic,ETH)的Dos攻击、畸形包注入攻击、端口扫描攻击。
在上述设计中,通过给出不易在真实攻击环境中进行标记的多种攻击类型,能便于用户根据实际需求选择其中一种或多种攻击类型来构建模拟攻击样本,以适用于不同的模拟攻击场景,提高构建模拟攻击样本的灵活性和通用性。
一种可能的设计中,数据处理装置通过攻击待检测设备,获得第一类型攻击样本,包括:数据处理装置遍历预设的多种攻击类型中的每种攻击类型,在遍历每种攻击类型时:对待检测设备执行该种攻击类型对应的攻击行为,并获取待检测设备对于该攻击行为所产生的流量数据,进而,当确定该流量数据是攻击流量时,将该流量数据标记为一个第一类型攻击样本。
在上述设计中,预设的多种攻击类型例如可以是前述设计中所给出的易于在真实攻击环境中进行标记的多种攻击类型和不易在真实攻击环境中进行标记的多种攻击类型,如此,通过对已知的每种攻击类型都进行流量采集和标记,能使第一类型攻击样本充分覆盖到已知的各种攻击类型,提高第一类型攻击样本的丰富性和全面性。
进一步地设计中,数据处理装置在获取待检测设备对于任一攻击行为所产生的流量数据后,若确定该流量数据不是攻击流量,而是正常流量,则可以将该流量数据标记为一个非攻击样本,进而,在遍历完全部的攻击类型后,根据第一类型攻击样本、第二类型攻击样本和非攻击样本,一起构建得到入侵检测样本集。
在上述设计中,入侵检测样本集中既包含攻击样本(包括前述的第一类型攻击样本和第二类型攻击样本)又包含非攻击样本,如此,不仅能使入侵检测样本集中的样本信息更加完善,还能在利用入侵检测样本集对待检测设备进行攻击测试时,根据待检测设备是否能拦截攻击样本以及是否能不拦截非攻击样本,更为准确地定义待检测设备的防攻击效果。
进一步地设计中,数据处理装置在获取待检测设备对于任一攻击行为所产生的流量数据后,若确定该流量数据既不满足攻击流量的特征又不满足正常流量的特征,则可以将该流量数据确定为一个无标记样本。其中,无标记样本属于未知样本,即在当前技术手段下无法识别出是攻击样本还是非攻击样本的样本。一些场景中,无标记样本属于异常样本,比如,当数据处理装置的攻击测试造成待检测设备的软硬件系统故障时,待检测设备本身可能会产生一些异常数据,这些异常数据既不满足正常流量的特征,也不满足攻击流量的特征,但仍会被数据处理装置采集下来。进而,针对于无标记样本,数据处理装置可以直 接进行丢弃,以节省入侵检测样本集的数据量,也可以在遍历完全部的攻击类型后,根据攻击样本、无标记样本和非攻击样本,一起构建得到入侵检测样本集,以将攻击待检测设备时所真实存在的全部样本都添加在入侵检测样本集中,提高入侵检测样本集中的样本丰富性,也能便于后续通过其它分析对无标记样本进行标记或执行其它一些操作。
进一步地设计中,入侵检测样本集中的攻击样本和非攻击样本占据同一比例,比如,攻击样本和非攻击样本各占全部样本的50%。如此,通过均分攻击样本和非攻击样本,能使入侵检测样本集中的攻击样本和非攻击样本维持平衡,便于后续取出相同比例的数据进行待检测设备的攻击测试,提高攻击测试结果的可信度。
进一步地设计中,数据处理装置在构建得到入侵检测样本集后,若确定入侵检测样本集中的攻击样本占据的比例和非攻击样本占据的比例不同,则可以通过对占据比例较多的样本进行裁剪,使得裁剪后的攻击样本占据的比例和非攻击样本占据的比例相同。其中,占据比例较多的样本通常为非攻击样本,也称为上下文数据,如此,通过裁剪非攻击样本,能维持入侵检测样本集中攻击样本和非攻击样本的平衡性。
一种可能的设计中,数据处理装置通过对第一类型攻击样本施加噪声,得到第二类型攻击样本,包括:数据处理装置先对第一类型攻击样本施加噪声以得到扰动样本,再将扰动样本输入攻击识别模型,并获得攻击识别模型输出的识别结果,当该识别结果指示无法确定扰动样本是否为攻击样本时,将扰动样本确定为第二类型攻击样本,否则,根据该识别结果调节扰动样本,并将调节后的扰动样本再次输入攻击识别模型,不断重复上述过程,直至调节后的扰动样本对应的识别结果指示无法确定调节后的扰动样本是否为攻击样本时,将调节后的扰动样本确定为第二类型攻击样本。
在上述设计中,第二类型攻击样本是在已知攻击类型的第一类型攻击样本的基础上,通过施加噪声而得到的无法被识别的样本,可以认为是在当前技术手段下无法确定攻击类型的攻击样本,也即是,未知攻击类型的攻击样本,如此,通过将该攻击样本添加在入侵检测样本集中,能在真实攻击测试场景中避免将这些未知攻击类型的攻击样本误报为非攻击样本的现象发生。
一种可能的设计中,数据处理装置根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以通过对入侵检测样本集进行特征提取,获得离线检测样本集,该离线检测样本集用于输入至入侵检测模型,以根据入侵检测模型输出的检测结果测评入侵检测模型的检测性能。比如,当离线检测样本集中属于攻击样本的离线检测样本越多地被入侵检测模型检测为攻击样本,同时属于非攻击样本的离线检测样本越多地被入侵检测模型检测为非攻击样本,则意味着入侵检测模型的检测效果较好。反之,当离线检测样本集中属于攻击样本的离线检测样本越多地被入侵检测模型检测为非攻击样本,而属于非攻击样本的离线检测样本越多地被入侵检测模型检测为攻击样本,则意味着入侵检测模型的检测效果较差。进而,当检测效果较差时,数据处理装置还可以调节入侵检测模型的参数,并使用调节后的入侵检测模型重新检测离线检测样本集,重复上述过程,直至得到检测效果较好的入侵检测模型为止。
在上述设计中,入侵检测样本集能支持在离线状态下测评入侵检测模型的检测效果(简称为离线测评),以便于根据该检测效果不断优化入侵检测模型,得到检测效果较好的入侵检测模型,提高入侵检测模型在待检测设备上的落地效果。
进一步地设计中,数据处理装置对入侵检测样本集进行特征提取,获得离线检测样本 集,包括:数据处理装置先确定入侵检测样本集中的每个入侵检测样本的报文类型,进而,针对于报文类型为传输控制协议(transmission control protocol,TCP)报文的各个入侵检测样本,通过对属于同一TCP连接的全部入侵检测样本进行特征提取,获得一个离线检测样本,而针对于报文类型为可变数据速率的CAN(CAN with flexible data rate,CAN(FD))报文或UDP报文的各个入侵检测样本,则通过对每个CAN(FD)报文或UDP报文的入侵检测样本进行特征提取,获得一个离线检测样本。
在上述设计中,通过对属于同一TCP连接的各个入侵检测样本进行联合的特征提取,能将一个TCP连接线上所采集到的全部入侵检测样本聚合为一个离线检测样本,有效减少离线检测样本的样本数量的同时,使每个离线检测样本所包含的特征更有关联性。此外,通过对不属于TCP连接的入侵检测样本(诸如属于CAN(FD)报文和UDP报文的入侵检测样本)进行单独的特征分析,则还能不遗漏没有关联关系的每个入侵检测样本,确保离线检测样本集覆盖全部的入侵检测样本。
进一步地设计中,离线测评所提取的特征可以包括如下特征中的一项或多项:时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。比如,对于属于TCP报文的入侵检测样本,所提取的特征可以包括此处所示出的全部特征,而对于属于CAN(FD)报文或UDP报文的入侵检测样本,则可以包括时间戳、频率特征、协议类型、内容特征、丢包率和错误包数量。
在上述设计中,通过给出离线测评对应的多种可提取特征,能便于用户根据实际离线场景中的应用层需求选择其中的一项或多项特征进行入侵检测样本集到离线检测样本集的转化,以适用于不同的离线测评场景,提高入侵检测样本集在离线测评领域中的通用性。
进一步地设计中,为保持各个离线检测样本的格式一致性,数据处理装置可以对每个入侵检测样本都提取前述的所有特征,进而,在某一入侵检测样本的某一特征不存在时,将该入侵检测样本集的该特征配置为预设字符,预设字符比如可以为数字、字母、符号或者其中一项或多项的组合形式等。
一种可能的设计中,数据处理装置根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以对入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集,进而通过测试工具将在线检测样本集输入待检测设备,该在线检测样本集用于测评部署有入侵检测模型的待检测设备的检测性能。
在上述设计中,入侵检测样本集能支持在在线状态下测评部署有入侵检测模型的待检测设备的检测效果(简称为在线测评),以便于根据该检测效果确定待检测设备的防攻击性能,确保仅出厂防攻击效果较好的待检测设备。
进一步地设计中,测试工具可以为CANoe、PCAN、Technica或其它能实现在线测评的工具,而格式转换后的入侵检测样本集可以为.PCAP、.ASC、.BLF或其它在线测试工具所对应的格式。
在上述设计中,通过使入侵检测样本支持转换为多种测试工具所对应的格式,能便于用户根据实际在线场景中的测试工具选择所对应的格式进行转换,以适用于不同的在线测评场景,提高入侵检测样本集在在线测评领域中的通用性。
一种可能的设计中,数据处理装置根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以根据入侵检测样本集在各个预设指标下的值,确定入侵检测样本集对应的评估值,当评估值低于预设阈值时,调节入侵检测样本集。
在上述设计中,预设指标例如可以是根据待检测设备所属的体系架构的特性而设置的,通过参照预设指标对已构建的入侵检测样本集进行评估,能根据评估结果有效优化入侵检测样本集,使优化后的入侵检测样本集更加适用于待检测设备所属的体系架构。
进一步地设计中,预设指标可以包括如下指标中的一项或多项:数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标、平衡性指标、特征独立性指标、易用性指标。其中,数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标和衡性指标属于定量指标,而特征独立性指标和易用性指标属于定性指标,通过结合定量指标和定性指标综合评判入侵检测样本集的构建质量,能使评判结果更加全面,且更具有说服性。
进一步地设计中,为使各个预设指标的度量具有可比性,各个预设指标的取值范围可以配置成同一范围,比如[0,1]。
进一步地设计中,入侵检测样本集对应的评估值可以是入侵检测样本集在各个预设指标下的值的加权平均值,其中,各个预设指标对应的权重可以相同也可以不同,比如,一个具体的示例中,可以配置各个定量的预设指标均对应第一权重,各个定性的预设指标均对应第二权重,且第一权重大于第二权重。
在上述设计中,通过减小定性的预设指标的权重,能降低需要通过人为经验判断的预设指标对评估值的影响,使得评估过程更加关注理性的评判标准,同时也不放弃需要人为经验的评判标准。
一种可能的设计中,第一类型攻击样本可以是通过攻击如下任一区域得到的:整个待检测设备;待检测设备的一个或多个物理区域;或者,待检测设备的一个或多个功能区域。示例性地,当待检测设备为车辆时,一个或多个物理区域可以包括车载主干区域、左前车身区域、右前车身区域、左后车身区域或右后车身区域中的一个或多个,一个或多个功能区域可以包括车载中控网关区域、车身控制区域、座舱控制区域、动力控制区域、底盘控制区域或信息娱乐区域等中的一个或多个。
在上述设计中,数据处理装置可以针对于整个待检测设备构建一个入侵检测样本集,也可以针对于待检测设备中的一个或多个物理区域构建一个入侵检测样本集,也可以针对于待检测设备中的一个或多个功能区域构建一个入侵检测样本集,或者,还可以针对于一个或多个物理区域和一个或多个功能区域组合构建一个入侵检测样本集,可见,该入侵检测样本集的构建方法可以适用于各种不同的构建场景,有助于提高入侵检测样本集的灵活性、通用性和易用性。
第二方面,本申请提供一种数据处理装置,该数据处理装置可以为具有处理能力的任意设备,如服务器或由服务器构成的服务器集群。该数据处理装置包括:攻击单元,用于通过攻击待检测设备,获得第一类型攻击样本;扰动单元,用于通过对第一类型攻击样本施加噪声,获得第二类型攻击样本;构建单元,用于根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。
一种可能的设计中,第一类型攻击样本可以包括真实攻击样本和模拟攻击样本,真实攻击样本是通过人为攻击待检测设备得到的,模拟攻击样本是通过攻击工具攻击待检测设备得到的。
一种可能的设计中,真实攻击样本可以对应如下攻击类型中的一项或多项:ID不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻 击、ID来源非指定ECU攻击、出现相同ID攻击、CAN扫描攻击、UDS执行敏感操作攻击、消息认证错误攻击、ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击、未知的入栈连接攻击。
一种可能的设计中,模拟攻击样本可以对应如下攻击类型中的一项或多项:ID Fuzz攻击、数据Fuzz攻击、CAN的Dos攻击、ETH的Dos攻击、畸形包注入攻击、端口扫描攻击。
一种可能的设计中,攻击单元具体用于:遍历预设的多种攻击类型中的每种攻击类型,在遍历每种攻击类型时:对待检测设备执行该种攻击类型对应的攻击行为,并获取待检测设备对于该攻击行为所产生的流量数据,若该流量数据是攻击流量,则将该流量数据标记为一个第一类型攻击样本。
一种可能的设计中,攻击单元在获取待检测设备对于攻击行为所产生的流量数据之后,还用于:若流量数据是正常流量,则将流量数据标记为一个非攻击样本;对应的,构建单元具体用于:根据第一类型攻击样本、第二类型攻击样本和非攻击样本,构建得到入侵检测样本集。
一种可能的设计中,扰动单元具体用于:对第一类型攻击样本施加噪声,得到扰动样本,将扰动样本输入攻击识别模型,获得攻击识别模型输出的识别结果,根据识别结果调节扰动样本,直至调节后的扰动样本对应的识别结果指示无法确定扰动样本是否为攻击样本时,将调节后的扰动样本确定为第二类型攻击样本。其中,攻击识别模型输出的识别结果用于指示扰动样本是否属于攻击样本。
一种可能的设计中,数据处理装置还可以包括特征提取单元,特征提取单元用于:对入侵检测样本集进行特征提取,获得离线检测样本集,该离线检测样本集用于测评入侵检测模型的检测效果。
一种可能的设计中,特征提取单元具体用于:确定入侵检测样本集中的每个入侵检测样本的报文类型,对于报文类型为TCP报文的各个入侵检测样本,通过对属于同一TCP连接的全部入侵检测样本进行特征提取,获得一个离线检测样本,而对于报文类型为CAN(FD)报文或UDP报文的各个入侵检测样本,通过对每个CAN(FD)报文或UDP报文的入侵检测样本进行特征提取,获得一个离线检测样本。
一种可能的设计中,前述特征提取的特征包括如下特征中的一项或多项:时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。
一种可能的设计中,数据处理装置还可以包括格式转换单元,格式转换单元用于:对入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集,并通过测试工具将在线检测样本集输入待检测设备,其中,在线检测样本集用于测评部署有入侵检测模型的待检测设备的检测性能。
一种可能的设计中,数据处理装置还可以包括调节单元,调节单元用于:根据入侵检测样本集在各个预设指标下的值,确定入侵检测样本集对应的评估值,当评估值低于预设阈值时,调节入侵检测样本集。
一种可能的设计中,预设指标包括如下指标中的一项或多项:数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标、平衡性指标、特征独立性指标、易用性指标。
一种可能的设计中,第一类型攻击样本可以是通过攻击如下任一区域得到的:整个待检测设备;待检测设备的一个或多个物理区域;或者,待检测设备的一个或多个功能区域。
第三方面,本申请提供一种数据处理装置,包括处理器,处理器与存储器相连,存储器用于存储计算机程序,处理器用于执行存储器中存储的计算机程序,以使得数据处理装置执行如上述第一方面中任一项设计所述的方法。
第四方面,本申请提供一种数据处理装置,包括处理器和存储器,存储器存储计算机程序指令,处理器运行计算机程序指令以实现如上述第一方面中任一项设计所述的方法。
第五方面,本申请提供一种数据处理装置,包括处理器、存储器和收发器,存储器存储计算机程序指令,处理器运行计算机程序指令,以调用收发器实现如上述第一方面中任一项设计所述的方法。
第六方面,本申请提供一种芯片,该芯片可以包括处理器和接口,处理器用于通过接口读取指令,以执行如上述第一方面中任一项设计所述的方法。
第七方面,本申请提供一种数据处理系统,该数据处理系统可以包括待检测设备和数据处理装置,数据处理装置用于执行如上述第一方面中任一项设计所述的方法。
第八方面,本申请提供一种计算机可读存储介质,该计算机可读存储介质存储有计算机程序,当计算机程序被运行时,实现如上述第一方面中任一项所述的方法。
第九方面,本申请提供一种计算机程序产品,当该计算机程序产品在处理器上运行时,实现如上述第一方面中任一项所述的方法。
上述第二方面至第九方面的有益效果,具体请参照上述第一方面中相应设计可以达到的技术效果,这里不再重复赘述。
附图说明
图1示例性示出本申请实施例提供的一种可能的系统架构示意图;
图2示例性示出本申请实施例提供的一种车辆的分区架构示意图;
图3示例性示出本申请实施例提供的一种数据处理方法的流程示意图;
图4示例性示出本申请实施例提供的一种第一类型攻击样本的获取流程示意图;
图5示例性示出本申请实施例提供的一种入侵检测样本集的应用场景示意图;
图6示例性示出本申请实施例提供的一种评估入侵检测样本集的流程示意图;
图7示例性示意出本申请实施例提供的一种开发数据处理方案的设计架构图;
图8示例性示出本申请实施例提供的一种数据处理装置的结构示意图;
图9示例性示出本申请实施例提供的另一种数据处理装置的结构示意图。
具体实施方式
本申请所公开的数据处理方法可以用于构建入侵检测样本集,该入侵检测样本集可以用于训练入侵检测模型,也可以用于检测入侵检测模型或部署有入侵检测模型的待检测设备的检测效果。其中,待检测设备可以是具有通信能力的任意终端设备,尤其可以是对网络安全具有一定需求的终端设备。一些示例中,该终端设备可以包括但不限于:智能运输设备,诸如汽车、轮船、无人机、火车、货车、卡车、飞行汽车等;智能家居设备,诸如电视、扫地机器人、智能台灯、音响系统、智能照明系统、电器控制系统、家庭背景音乐、家庭影院系统、对讲系统、视频监控等;智能制造设备, 诸如机器人、工业设备、工控机、智能物流、智能工厂等。或者,该终端设备也可以是计算机设备,例如台式机、个人计算机、服务器等。还应当理解的是,该终端设备也可以是便携式电子设备,诸如手机、平板电脑、掌上电脑、耳机、音响、穿戴设备(如智能手表)、车载设备、虚拟现实设备、增强现实设备等。便携式电子设备的示例包括但不限于搭载或者其它操作系统的便携式电子设备。上述便携式电子设备也可以是诸如具有触敏表面(例如触控面板)的膝上型计算机(Laptop)等。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行描述,显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。
需要说明的是,本申请实施例中的术语“系统”和“网络”可被互换使用。“至少一个”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B的情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指的这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b,或c中的至少一项(个),可以表示:a,b,c,a-b,a-c,b-c,或a-b-c,其中a,b,c可以是单个,也可以是多个。
以及,除非有特别说明,本申请实施例提及“第一”、“第二”等序数词是用于对多个对象进行区分,不用于限定多个对象的优先级或者重要程度。例如,第一类型攻击样本和第二类型攻击样本,只是为了区分不同类型的攻击样本,而并不是表示这两种攻击样本的优先级或者重要程度等的不同。
另外,本申请实施例中“连接”可以理解为电连接,两个电学元件连接可以是两个电学元件之间的直接或间接连接。例如,A与B连接,既可以是A与B直接连接,也可以是A与B之间通过一个或多个其它电学元件间接连接,例如A与B连接,也可以是A与C直接连接,C与B直接连接,A与B之间通过C实现了连接。在一些场景下,“连接”也可以理解为耦合,如两个电感之间的电磁耦合。总之,A与B之间连接,可以使A与B之间能够传输电能。
图1示例性示出本申请实施例提供的一种可能的系统架构示意图,如图1所示,该系统架构中包括待检测设备100和数据处理装置200。其中,待检测设备100可以是对网络安全具有一定要求的设备,比如车辆。数据处理装置200可以是具有数据处理能力的任意装置,比如服务器或者由多个服务器所构成的服务器集群,或者还可以是芯片或电路,比如设置在服务器或服务器集群中的芯片或电路。一个具体的示例中,该数据处理装置200可以是云服务器,云服务器可以通过无线方式连接待检测设备100。应理解,本申请实施例对该系统架构中的待检测设备100的数量和数据处理装置200的数量不作限定,例如,一个数据处理装置200可以如图1所示意的只连接一个待检测设备100,也可以同时连接多个待检测设备100。以及,本申请实施例中的数据处理装置200可以将所有的功能集成在一个独立的物理设备上,也可以将功能分布在多个独立的物理设备上,对此本申请实施例也不作具体限定。
进一步示例性地,继续参照图1所示,该系统架构中还可以包括数据库300、模型训练装置400和测试工具500中的一项或多项,或者还可以包括其它设备,如路由设备、无线中继设备、无线回传设备及操作管理维护设备等。其中,数据库300可以 用于存储数据处理装置200所构建的入侵检测样本集,该数据库300可以是独立于数据处理装置200以外的存储单元,比如数据库服务器,也可以是数据处理装置200的内部存储单元,如高速缓存存储器、随机存储器、寄存器、主存储器或只读存储器等。模型训练装置400可以是具有算法开发和验证能力的任意设备,如模型训练服务器。该模型训练装置400中部署有入侵检测系统(intrusion detection system,IDS),该IDS是汽车开放系统架构(automotive open system architecture,AUTOSAR)针对于汽车网络安全所发布的一种测试系统,属于目前最主流的车载测试系统之一,能在有限的车载硬件资源和车载存储资源的基础上结合规则和机器学习的方式训练出高准确性、高时效性和高鲁棒性的入侵检测模型。测试工具500是指能向待检测设备100注入测试流量的工具,比如可以是CAN测试工具、TCP测试工具或UDP测试工具等。且,该测试工具500对外可呈现用户界面,用户通过在用户界面上点击相应的测试命令(比如流量注入类型、流量注入数量及流量注入频率等),可驱动测试工具500按照对应的测试命令自动向待检测设备100注入测试流量。
进一步示例性地,当包括待检测设备100、数据处理装置200、数据库300、模型训练装置400和测试工具500时,数据处理装置200可以分别与待检测设备100、数据库300、模型训练装置400和测试工具500连接,待检测设备100还可以与模型训练装置400和测试工具500连接。在实施中,数据处理装置200可以针对于待检测设备100构建入侵检测样本集,并将该入侵检测样本集存储在数据库300中。之后,在使用时,数据处理装置200可以将数据库300中的入侵检测样本集转化为离线检测样本集,并可以选取其中一部分离线检测样本发送给模型训练装置400,由模型训练装置400使用该部分离线检测样本训练得到入侵检测模型后,将入侵检测模型部署在待检测设备100中,以便于待检测设备100利用入侵检测模型识别攻击流量。进而,在需要进行离线测评时,数据处理装置200还可以将转化得到的另一部分离线检测样本发送给模型训练装置400,由模型训练装置400使用之前训练好的入侵检测模型检测该部分离线检测样本以得到离线测评信息,数据处理装置200根据该离线测评信息测评模型训练装置400所训练的入侵检测模型的检测效果。相对应的,在需要进行在线测评时,数据处理装置200可以将数据库300中的入侵检测样本集转化为在线检测样本集,进而通过在线工具将部分或全部的在线检测样本输入待检测设备100,并获取待检测设备100针对于该在线检测样本产生的在线测评信息,根据该在线测评信息测评部署有入侵检测模型的待检测设备100的检测效果。
根据上述内容可知,入侵检测样本集不仅会用来训练入侵检测模型,还会作为测试样本用来测试入侵检测模型的质量和部署有入侵检测模型的待检测设备的质量。当该入侵检测样本集中的样本类型足够充分时,依据足够充分的入侵检测样本训练得到的入侵检测模型的检测效果也会更好,进而部署有入侵检测模型的待检测设备的防攻击性能也会更好,相应地,依据足够充分的入侵检测样本测评入侵检测模型或待检测设备所得到的测评结果的可靠性也会更好。因此,如何构建样本类型足够充分的入侵检测样本集,对于提高入侵检测模型的检测效果、提高待检测设备的防攻击效果、提高测评入侵检测模型的可靠性以及提高测评待检测设备的可靠性至关重要。
然而,业界在构建入侵检测样本集时,仅仅只是会利用攻击工具模拟一些典型的入侵场景来对待检测设备进行攻击,导致入侵检测样本集中也只会包含典型的入侵场 景所对应的入侵检测样本,该入侵检测样本集中的样本类型非常有限,不利于提高入侵检测模型的检测效果、提高待检测设备的防攻击效果、提高测评入侵检测模型的可靠性以及提高测评待检测设备的可靠性,进而也就不利于入侵检测技术在待检测设备上的落地。
有鉴于此,本申请实施例提供一种数据处理方法,用以构建样本类型更为丰富的入侵检测样本集,使其既能覆盖到已知入侵场景下的入侵检测样本,又能覆盖到未知入侵场景下的入侵检测样本,以提高使用该入侵检测样本集训练得到的入侵检测模型的检测效果,提高部署有入侵检测模型的待检测设备的防攻击效果,提高使用该入侵检测样本集测评入侵检测模型的可靠性,以及提高使用该入侵检测样本集测评待检测设备的可靠性。
需要说明的是,本申请实施例中的入侵检测样本集可以是针对于整个待检测设备构建得到的,也可以是针对于待检测设备中的一个或多个区域构建得到的。示例性地,当待检测设备为车辆时,图2示出了本申请实施例提供的一种车辆的分区架构示意图,其中:
图2中(A)示出的是车辆的物理分区架构图,该架构根据物理区域的不同将整个车辆划分为车载主干区域、右前车身区域、左前车身区域、右后车身区域和左后车身区域,右前车身区域、左前车身区域、右后车身区域和右后车身区域作为车载主干区域的分支区域而存在。其中,车载主干区域中部署有整车控制器(vehicle control unit,VCU),每个分支区域中部署有各自的控制节点(即Z1、Z2、Z3、Z4)和与控制节点相连的若干ECU,每个分支区域中的全部ECU通过CAN(FD)与控制节点相连,且每个分支区域中的控制节点通过ETH与车载主干区域中的VCU相连;
图2中(B)示出的是车辆的功能分区架构图,该架构根据实现功能的不同将整个车辆划分为车载中控网关功能区域和K个其它功能区域,K个其它功能区域作为车载中控网关功能区域的分支区域而存在,各自实现不同的功能,比如,一些场景中,K个其它功能区域可以包括车身控制域、座舱控制域、动力控制域、底盘控制域或信息娱乐域等中的一项或多项,K为正整数。其中,车载中控网关功能区域中部署有GateWay网关,每个其它功能区域中部署有各自的域控制器(即D1、D2、……、D4K)和与域控制器相连的若干ECU,每个其它功能区域中的全部ECU通过CAN(FD)与域控制器相连,且每个其它功能区域中的域控制器通过ETH与车载中控网关功能区域中的GateWay相连。
基于图2中(A)或图2中(B)所示意的车载分区架构,本申请实施例中可以针对于整个车辆构建一个入侵检测样本集,也可以针对于车辆中的一个或多个物理区域构建一个入侵检测样本集,还可以针对于车辆中的一个或多个功能区域构建一个入侵检测样本集,或者,还可以针对于一个或多个物理区域以及一个或多个功能区域的组合构建一个入侵检测样本集等。其中,物理区域可以是车载主干区域,也可以是右前车身区域、左前车身区域、右后车身区域或左后车身区域,功能区域可以是车载中控网关功能区域,也可以是其它功能区域,具体不作限定。
需要说明的是,具体针对于哪些区域构建入侵检测样本集,可以根据数据处理装置的处理能力和用户的实际需求进行设置。比如,当数据处理装置的处理能力较强时,可以选择针对于整个车辆构建入侵检测样本集,以利用高效的处理能力构建得到比较完整的全局 入侵检测样本集。其中,该全局入侵检测样本集可以适用于全局攻击场景,也可以适用于局部攻击场景,通用性较好。反之,当数据处理装置的处理能力不强时,可以根据实际业务所涉及到的区域,针对性地选取这些区域来构建局部的入侵检测样本集。其中,该局部的入侵检测样本集属于全局的入侵检测样本集的一个子集,虽然仅可以适用于局部攻击场景,但其需要处理的区域规模变小,因此能较好地降低构建入侵检测样本集的难度,降低入侵检测样本集的复杂程度。
下面以构建整个待检测设备对应的入侵检测模型为例,通过具体的实施例来介绍上述技术问题是如何被解决的。需要说明的是,在下文的描述中,数据处理装置可以是图1所示意的数据处理装置200,也可以是能够支持数据处理装置实现所需功能的其它通信节点、通信装置或通信系统,例如芯片、芯片系统、电路或电路系统,具体不作限定。
基于图1所示意的系统架构,图3示例性示出本申请实施例提供的一种数据处理方法的流程示意图,该数据处理方法可以由数据处理装置执行,比如图1所示意的数据处理装置200。如图3所示,该方法包括:
步骤301,数据处理装置通过攻击待检测设备,获得第一类型攻击样本。
其中,当针对于整个待检测设备构建全局入侵检测样本集,数据处理装置可以针对于整个待检测设备发起攻击行为,以获得整个待检测设备对应的第一类型攻击样本。反之,当针对于一个或多个区域构建局部入侵检测样本集时,数据处理装置可以仅针对于该一个或多个区域发起攻击行为,以获得该一个或多个区域对应的第一类型攻击样本。
一个示例中,待检测设备为车辆,该车辆可以是分布式电子电气架构、域集中式电子电气架构、车辆集中式电子电气架构或未来可能出现的任一车载架构下的车辆。如此,数据处理装置通过对属于每种车载架构的车辆进行攻击,能构建出每种车载架构所对应的入侵检测样本集,进而能便于用户根据实际需求选择其中一种或多种车载架构所对应的入侵检测样本集进行使用,提高用户的使用体验。
进一步地示例中,为了提高任一车载架构所对应的入侵检测样本集的准确性,数据处理装置还可以针对于属于该车载架构的多台车辆一起进行攻击,以使第一类型攻击样本能覆盖到该车载架构下的多台车辆,避免只攻击一台车辆时由于该车辆出现问题而导致的样本采集不准确的问题。
进一步地示例中,当针对于属于一个车载架构的多台车辆进行攻击时,数据处理装置在获取到多台车辆对应的大量的第一类型攻击样本后,还可以通过聚类或模型识别等手段对这些第一类型攻击样本进行筛选,比如筛除其中与其它的第一类型攻击样本差异较大的第一类型攻击样本,而仅保留比较相似的第一类型攻击样本。如此,通过提前清理明显存在问题的第一类型攻击样本,能避免后续对明显存在问题的第一类型攻击样本进行无意义地处理,有效节省数据处理装置的计算资源。
本申请实施例中,第一类型攻击样本可以包括真实攻击样本和模拟攻击样本。其中,真实攻击样本是通过人为攻击待检测设备而得到的,例如可以是在渗透测试的过程中采集得到的。其中,渗透测试过程是指攻击人员站在黑客的角度以攻破待检测设备为目的真实地对车辆进行攻击,这种攻击方式能够构造出隐蔽性较强的攻击行为和依赖业务逻辑的攻击行为,能获取到易于在真实攻击环境中标记的攻击样本。相应地,模拟攻击样本是通过攻击工具攻击待检测设备而得到的,例如可以是利用攻击工具模拟产生攻击流量并将该攻击流量自动地注入到待检测设备后从待检测设备中采集的,该种攻击方式能够获取到不易 在真实攻击环境中标记的攻击样本。如此,通过结合人为攻击方式和攻击工具攻击方式综合构建第一类型攻击样本,能使第一类型攻击样本充分覆盖到真实攻击环境中可能存在的各种攻击类型的攻击样本,提高第一类型攻击样本的全面性。
一种可选地实施方式中,第一类型攻击样本可以是在现阶段已知的入侵场景中根据已知的各种攻击类型攻击待检测设备而得到的。其中,已知的各种攻击类型包括易于在真实攻击环境中进行采集和标注的攻击类型(简称为第一种攻击类型)和不易在真实攻击环境中进行采集和标注的攻击类型(简称为第二种攻击类型),因此,通过采用人为方式按照第一种攻击类型真实地攻击待检测设备,可以采集得到上述真实攻击样本,通过采用攻击工具按照第二种攻击类型模拟地攻击待检测设备,可以集中采集得到上述模拟攻击样本。
举例来说,假设待检测设备为车辆,且车辆中的各个ECU通过CAN(FD)总线和ETH总线进行通信,则,请参照表1,其示出了待检测设备中可能存在的攻击类型以及每种攻击样本所对应的攻击类型的对应关系表。
表1

参照表1所示,第一种攻击类型可以包括如下攻击类型中的一项或多项:
ID不存在攻击,是指通过将CAN报文中的ID修改为某个不存在的ID来攻击车辆;
重放攻击,是指通过发送车辆之前已接收过的CAN报文来欺骗车辆;
篡改攻击,是指通过篡改CAN报文中携带的数据来攻击车辆;
数据长度错误攻击,是指通过修改CAN报文的数据长度来攻击车辆;
信号超出定义范围攻击,是指通过将CAN报文中的信号数值修改为大于最大规定值的值或小于最小规定值的值来攻击车辆;
上下文错误攻击,是指通过在CAN网络上发布不适合某种状态的特定消息或信号来攻击车辆,比如在刹车状态下发出加速信号;
ID来源非指定ECU攻击,是指通过非指定的ECU发布理应由指定ECU发布的CAN报文来攻击车辆;
出现相同ID攻击,是指通过在两个ECU发布的CAN报文中携带相同的ID来攻击车辆;
CAN扫描攻击,是指通过发送CAN扫描消息来侵入车辆;
UDS执行敏感操作攻击,是指通过在UDS报文中携带敏感信息来攻击车辆;
消息认证错误攻击,是指通过修改CAN总线的消息认证流程使得车辆无法成功认证消息;
ECU身份欺骗攻击,是指通过更改ETH消息中的ECU身份来欺骗车辆;
中间人攻击,是指通过将一台受入侵者控制的设备虚拟放置在ETH连接的两个ECU之间来攻击车辆;
ECU认证错误攻击,是指通过修改ETH的ECU认证流程使得车辆无法成功ECU;
暴力破解攻击,是指通过穷举的方式破译ETH密钥来攻击车辆;
应用层协议错误攻击,是指通过修改应用层协议使得车辆无法获取正确的应用层协议进行消息交互来攻击车辆;
未知的出栈连接攻击,是指通过给定一个虚假的出栈连接来攻击车辆;
未知的入栈连接攻击,是指通过给定一个虚假的入栈连接来攻击车辆。
在上述第一种攻击类型中,ID不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻击、ID来源非指定ECU攻击、出现相同ID攻击、CAN扫描攻击、UDS执行敏感操作攻击和消息认证错误攻击属于CAN(FD)通信方式下存在的攻击类型,而ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击和未知的入栈连接攻击属于ETH通信方式下存在的攻击类型。
继续参照表1所示,第二种攻击类型可以包括如下攻击类型中的一项或多项:
ID模糊(Fuzz)攻击,是指通过模糊处理CAN报文中的ID来攻击车辆;
数据Fuzz攻击,是指通过模糊处理CAN报文中的数据来攻击车辆;
CAN的Dos攻击,是指通过停止某部分CAN总线的收发服务来攻击车辆;
ETH的Dos攻击,是指通过停止某部分ETH总线的收发服务来攻击车辆;
畸形包注入攻击,是指通过注入畸形包来攻击车辆;
端口扫描攻击,是指通过发送端口扫描消息来侵入车辆。
在上述第二种攻击类型中,ID模糊(Fuzz)攻击、数据Fuzz攻击和CAN的Dos攻击属于CAN(FD)通信方式下存在的攻击类型,而ETH的Dos攻击、畸形包注入攻击和端口扫描攻击属于ETH通信方式下存在的攻击类型。
在上述实施方式中,通过给出已知入侵场景中可能存在的各种攻击类型,能便于用户根据实际需求选择其中一种或多种攻击类型来构建第一类型攻击样本,且还可以采用真实攻击或模拟攻击的方式获取所需要的真实攻击样本或模拟攻击样本,以适用于不同的攻击场景,提高构建第一类型攻击样本的灵活性和通用性。
进一步地示例中,数据处理装置可以通过多种方式攻击待检测设备来获取第一类型攻击样本,比如,请参照图4,其示出了本申请实施例提供的一种获取第一类型攻击样本的具体流程示意图,该流程包括:
步骤401,数据处理装置获取预设的多种攻击类型。
其中,预设的多种攻击类型示例性地可以包括上述表1所示意的全部攻击类型,以便使第一类型攻击样本能充分覆盖到已知的各种攻击类型,提高第一类型攻击样本的丰富性和全面性。
步骤402,数据处理装置确定预设的多种攻击类型中是否存在未被遍历过的攻击类型,若是,则执行步骤403,若否,则执行步骤409。
步骤403,数据处理装置对待检测设备执行未被遍历过的攻击类型所对应的攻击行为,并获取待检测设备对于该攻击行为产生的流量数据。
示例性地,针对于第一种攻击类型,可以通过人为方式提前编写好所需要的攻击代码,并将该攻击代码与所对应的第一种攻击类型映射存储在数据处理装置中。如此,数据处理装置在获取到未被遍历过的攻击类型后,如果确定该攻击类型属于第一种攻击类型,则可以从本地直接获取该第一种攻击类型对应的攻击代码,并自动按照该攻击代码生成对应的攻击行为后攻击待检测设备。反之,如果确定该攻击类型属于第二种攻击类型,则数据处理装置可以调用攻击工具,通过攻击工具生成该第二种攻击类型所对应的攻击行为后攻击对待检测设备。如此,整个攻击流程可以由数据处理装置自动实现,有助于提高对整个攻击流程的统一管理,且还不需要等待人为现场编程,有助于节省样本采集的时延。
进一步地示例中,当待检测设备为车辆时,数据处理装置可以通过数据线接入待检测设备的车载诊断(on-board diagnostics,OBD)接口,该OBD接口示例性地可以是II型车载诊断系统的接口,即OBD-II接口。最初时,由于全部的攻击类型都还未被遍历过,因此数据处理装置可以从全部的攻击类型中按照随机方式、或者按照顺序方式、或者按照其它方式选取一种攻击类型,然后按照该种攻击类型攻击待检测设备,并通过OBD接口获取待检测设备针对于该种攻击类型所产生的流量数据。进而,在分析完该流量数据后,数据处理装置可以再从未被遍历过的攻击类型中选取一种攻击类型继续攻击待检测设备,重复上述过程,直至全部的攻击类型都被遍历过为止。
进一步示例性地,上述整个攻击过程的流量采集操作可以通过数据处理装置自动实现。具体的,可以是在数据处理装置中配置采集时长,数据处理装置在开始执行攻击行为后,即可启动计时,在计时时间内,数不断从待检测设备的OBD接口处采集待检测设备输出的流量数据,直至计时到所配置的采集时长时,结束采集。其中,采集时长可以为几个小时、几天、几个星期甚至几个月,具体的,可以由本领域技术人员根据实际需求进行配置。 比如,当预设的攻击类型的数量较多时,攻击待检测设备所需的时间较长,采集时长也可以配置的越长,如数周。反之,当预设的攻击类型的数量较少时,攻击待检测设备所需的时间较短,采集时长也可以配置的越短,如数天。
步骤404,数据处理装置判断流量数据是否为攻击流量,若是,则执行步骤405,若否,则执行步骤406。
步骤405,数据处理装置将流量数据标记为一个第一类型攻击样本,之后执行步骤402。
示例性地,在确定某一流量数据满足攻击流量的特征后,若该流量数据是在第一种攻击类型下采集得到的,则数据处理装置可以将该攻击流量标记为真实攻击样本,反之,若该流量数据是在第二种攻击类型下采集得到的,则数据处理装置可以将该攻击流量标记为模拟攻击样本。
步骤406,数据处理装置判断流量数据是否为正常流量,若是,则执行步骤407,若否,则执行步骤408。
步骤407,数据处理装置将流量数据标记为一个非攻击样本,之后执行步骤402。
一些场景中,非攻击样本也称为上下文数据或上下文样本。
步骤408,数据处理装置确定流量数据为一个无标记样本,之后执行步骤402。
示例性地,当流量数据既不满足攻击流量的特征,也不满足正常流量的特征时,意味着该流量数据在当前的技术手段下无法识别其样本类型,该情况下,数据处理装置可以不对该流量数据进行标记,而是将其作为一个无标记样本。一些场景中,该无标记样本属于异常样本,比如,当数据处理装置的攻击测试造成待检测设备的软硬件系统故障时,待检测设备本身可能会产生一些异常数据,这些异常数据既不满足正常流量的特征,也不满足攻击流量的特征,但仍会被数据处理装置采集下来。
步骤409,数据处理装置结束攻击流程。
在上述示例中,数据处理装置通过攻击待检测设备,既能获取到攻击样本,如真实攻击样本和模拟攻击样本,又能获取到非攻击样本,甚至还能获取到无标记样本,这种获取方式能够获取到多种类型的样本,便于提高后续构建入侵检测样本集的样本丰富性。
需要说明的是,图4只是示例性介绍一种获取第一类型攻击样本的可能方式,本申请实施例并不限定只能采用这种方式获取第一类型攻击样本。比如,另一种可能的获取方式中,数据处理装置还可以组合预设的多种攻击类型中的至少两种攻击类型,一次性地对待检测设备执行至少两种攻击类型所对应的攻击行为,进而获取待检测设备生成的流量数据后,从该流量数据中分离出至少两种攻击类型各自对应的流量数据,根据至少两种流量数据获得第一类型攻击样本。应理解,可能的获取方式还有很多,此处不再进行一一列举。
步骤302,数据处理装置通过对第一类型攻击样本施加噪声,获得第二类型攻击样本。
示例性地,数据处理装置在获取到第一类型攻击样本后,可以对第一类型攻击样本进行格式转换,获得文本格式或二进制格式的第一类型攻击样本,并将文本格式或二进制格式的第一类型攻击样本存储在原始数据库中。
进一步示例性地,数据处理装置可以遍历原始数据库中的全部第一类型攻击样本,在遍历每个第一类型攻击样本时:对该第一类型攻击样本施加噪声以得到扰动样本,再将扰动样本输入攻击识别模型,并获得攻击识别模型输出的识别结果,当该识别结果指示无法确定扰动样本是否为攻击样本时,将扰动样本确定为一个第二类型攻击样本,否则,根据该识别结果调节扰动样本,并将调节后的扰动样本再次输入攻击识别模型,不断重复上述 过程,直至调节后的扰动样本对应的识别结果指示无法确定调节后的扰动样本是否为攻击样本时,将调节后的扰动样本确定为一个第二类型攻击样本。
其中,攻击识别模型可以是具有识别能力的任意模型,具体的,可以是采用人工智能(artificial intelligence,AI)算法训练得到的神经网络模型,该AI算法比如可以是生成对抗网络(generative adversarial networks,GAN)算法。GAN算法可以学习已知的一些攻击样本和非攻击样本的特征,并根据学习到的特征构建攻击识别模型,使攻击识别模型能够识别出输入样本属于攻击样本的概率。具体的,攻击识别模型中可以包含生成器和鉴别器,输入至攻击识别模型的任一第一类型攻击样本首先由生成器所接收。当该第一类型攻击样本具有N维特征空间(N为正整数)时,生成器可以通过对该第一类型攻击样本施加与N维特征空间一一对应的N维噪声向量,生成扰动样本,进而将扰动样本输入鉴别器。鉴别器鉴别出扰动样本属于攻击样本的概率后,如果该概率大于50%,则通知生成器按照第一方向调节N维噪声向量,如果该概率小于50%,则通知生成器按照第二方向调节N维噪声向量。进而,生成器使用调节后的N维噪声向量重新加扰第一类型攻击样本生成新的扰动样本,并发送给鉴别器,由鉴别器重新鉴别新的扰动样本属于攻击样本的概率,若该概率为50%,则可以将该新的扰动样本作为一个第二类型攻击样本,否则,重复执行上述过程,直至获得属于攻击样本的概率为50%的扰动样本为止。
如此,通过生成器和鉴别器的博弈,攻击识别模型可以生成无法识别样本类型的样本,但由于该样本是对真实攻击待检测设备所得到的攻击样本(即真实攻击样本和模拟攻击样本)进行加扰得到的,大概率属于攻击样本,因此,该样本可以认为是在当前技术手段下无法确定攻击类型的攻击样本,该攻击样本在真实识别过程中容易被误报。基于此,数据处理装置可以将该样本作为第二类型攻击样本,后续存入入侵检测样本集,以便于能在真实攻击测试场景中避免将这些未知攻击类型的攻击样本误报为非攻击样本的现象发生。
此外,由于第二类型攻击样本是通过AI算法中的生成器和鉴别器对抗而生成的,因此,第二类型样本也可以称为AI对抗样本,或者也可以具有其它名称,本申请实施例对此不作具体限定。
步骤303,数据处理装置根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。
示例性地,数据处理装置根据上述步骤301获得第一类型攻击样本(包括真实攻击样本和模拟攻击样本)和非攻击样本,以及根据上述步骤302中获得第二类型攻击样本后,可以将这些样本统一保存为.pcap格式,进而根据.pcap格式的全部样本构建得到入侵检测样本集。其中,.pcap格式是与现有的IDS系统相对接的格式,如果是对接其它系统,则数据处理装置也可以将这些样本保存为其它系统所对接的格式,本申请实施例对此不作具体限定。
进一步示例性地,针对于上述步骤301中获取的无标记样本,数据处理装置可以直接进行丢弃,以节省入侵检测样本集的数据量,也可以在遍历完全部的攻击类型后,根据攻击样本、无标记样本和非攻击样本,一起构建得到入侵检测样本集,以将攻击待检测设备时所真实存在的全部样本都添加在入侵检测样本集中,提高入侵检测样本集中的样本丰富性,也能便于后续通过其它分析对无标记样本进行标记或执行其它一些操作,或者一些特例中,还可以将其标记为异常样本后添加在入侵检测样本集中,具体不作限定。
进一步示例性地,数据处理装置在构建得到入侵检测样本集后,若确定入侵检测样本 集中的攻击样本(包括第一类型攻击样本和第二类型攻击样本)占据的比例和非攻击样本占据的比例不同,则可以通过对占据比例较多的样本进行裁剪,使得裁剪后的攻击样本占据的比例和非攻击样本占据同一比例,如攻击样本和非攻击样本各占全部样本的50%。其中,占据比例较多的样本通常为非攻击样本,在某些特殊情况下也可能是攻击样本,通过裁剪占据比例较多的样本,能使入侵检测样本集中的攻击样本和非攻击样本维持平衡,便于后续取出相同比例的数据进行攻击测试,提高攻击测试结果的可信度。
在上述构建方式中,入侵检测样本集中既能覆盖已知入侵场景下的第一类型攻击样本,又能覆盖未知入侵场景下的第二类型攻击样本,还能覆盖非攻击样本,如此,不仅能使入侵检测样本集中的样本信息更加完善,还能在利用入侵检测样本集对待检测设备进行攻击测试时,根据待检测设备是否能拦截攻击样本以及是否能不拦截非攻击样本,更为准确地定义待检测设备的防攻击效果。
上述内容介绍了入侵检测样本集的具体构建过程,下面再对已构建的入侵检测样本集的应用进行详细介绍。
图5示例性示出本申请实施例提供的一种入侵检测样本集的应用场景示意图,如图5所示,该入侵检测样本集可以应用于模型训练场景、离线测评场景或在线测评场景中的一个或多个场景。其中,图示实线示出的是模型训练场景的应用流程,图示虚线示出的是离线测评场景的应用流程,图示双节点线示出的是在线测评场景的应用流程。下面参照图5分别对这三个应用场景进行详细介绍。
模型训练场景
参照图5所示意的实线部分,模型训练是指在早期算法开发和验证中,使用入侵检测样本集训练得到入侵检测模型。由于入侵检测模型所需的训练样本具有自己的特征格式,而入侵检测样本集中的入侵检测样本未必与该特征格式相匹配,因此在训练入侵检测模型之前,数据处理装置还可以先对入侵检测样本集进行特征提取,并根据提取得到的特征构建得到离线检测样本集。如此,在需要训练入侵检测模型时,数据处理装置即可直接从离线检测样本集中选取部分离线检测样本作为训练集数据,输入至模型训练装置,由模型训练装置利用训练集数据训练入侵检测模型。其中,训练集数据中可以包括攻击类型的离线检测样本和正常类型的离线检测样本,攻击类型的离线检测样本是通过对入侵检测样本集中的第一类型攻击样本和/或第二类型攻击样本进行特征提取而得到的,正常类型的离线检测样本是对入侵检测样本集中的非攻击样本进行特征提取而得到的,且攻击类型的离线检测样本和正常类型的离线检测样本还可以保持数量的一致性,以便于能基于均衡的数据训练得到效果较好的入侵检测模型。
需要说明的是,离线检测样本集中的各个离线检测样本可以保存为文本格式,如.csv。其中,该.csv格式是与现有的IDS系统中的模型训练装置相对接的格式,如果对接的是其它类型的模型训练装置,则可以保存为其它系统所适配的格式,具体不作限定。
本申请实施例中,在进行特征提取时,数据处理装置可以孤立地分析每个入侵检测样本,也即是对每个入侵检测样本进行特征提取后得到一个离线检测样本,也可以联合多个入侵检测样本进行集中分析,比如对具有关联关系的多个入侵检测样本进行特征提取后得到一个离线检测样本。其中,具有关联关系的多个入侵检测样本比如可以是属于同一个连接的多个入侵检测样本、流量数据来自于同一个ECU的多个入侵检测样本或者流量数据 发送至同一个ECU的多个入侵检测样本等。
进一步示例性地,以具有关联关系的多个入侵检测样本为属于同一个连接的多个入侵检测样本为例,当待检测设备为车辆时,车辆中的报文类型通常可以包括TCP报文、CAN(FD)报文和UDP报文,其中,TCP报文是在至少两个ECU之间建立连接后在该连接上传输的报文,而CAN(FD)报文和UDP报文则是通过广播的方式发送在对应的总线上后由所需的ECU节点从总线中获取的报文,这三种报文的内容中既包含源互联网协议(internet protocol,IP)地址又包含目的IP地址。可见,TCP报文具有连接的概念,而CAN(FD)报文和UDP报文不具有连接的概念,因此,数据处理装置在构建得到入侵检测样本集后,可以先确定入侵检测样本集中的每个入侵检测样本的报文类型,进而,针对于报文类型为TCP报文的各个入侵检测样本,根据报文内容中包含的源IP地址和目的IP地址,获取属于同一TCP连接的全部入侵检测样本,并对这些入侵检测样本进行特征提取后得到一个离线检测样本。反之,针对于报文类型为CAN(FD)报文和UDP报文的每个入侵检测样本,单独地对每个入侵检测样本进行特征提取后得到一个离线检测样本。如此,通过联合一个TCP连接线上的全部入侵检测样本聚合得到一个离线检测样本,能在有效减少离线检测样本集中的样本数量的同时,使每个离线检测样本所包含的特征更有关联性,而通过对不具有连接概念的每个入侵检测样本进行单独的特征分析,则还能不遗漏没有关联关系的每个入侵检测样本,确保离线检测样本集能覆盖到全部的入侵检测样本。
进一步示例性地,上述特征提取所提取的特征可以包括如下特征中的一项或多项:时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。其中,时间戳是指采集到报文的时间,包括日期和时刻等。频率特征是指在源IP和目的IP之间收发报文的通信频率。协议类型是指报文所适配的协议,比如上表3中的任一种协议类型。内容特征是指报文中携带的实质内容,比如数据或指令等。丢包率是指在源IP和目的IP之间收发报文时的报文丢失比例。错误包数量是指在源IP和目的IP之间收发报文时出现的错误报文的数量。连接持续时间是指一个连接自开始建立至断开之间的时间间隔。连接发起方是指请求建立连接的ECU。连接接收方是指接收连接发起方所发送请求的ECU。
进一步示例性地,为保持各个离线检测样本的格式一致性,数据处理装置还可以针对于每类入侵检测样本都提取上述的全部特征。具体的,对于属于TCP报文的入侵检测样本,由于其存在连接的概念,因此可以提取到上述的全部特征。而对于属于CAN(FND)报文或UDP报文的入侵检测样本,由于其不存在连接的概念,因此仅可以提取到上述特征中的时间戳、频率特征、协议类型、内容特征、丢包率和错误包数量,而无法提取到连接持续时间、连接发起方和连接接收方。该情况下,为了保持离线检测样本的格式一致性,数据处理装置还可以将无法提取到的特征配置为预设字符,该预设字符比如可以为数字、字母、符号或者其中一项或多项的组合形式等。
需要说明的是,上述内容只是给出一种特征提取的示例,至于实际应用时需要提取哪些特征,则可以由本领域技术人员根据经验进行设置,本申请实施例对此不作限定。
在模型训练场景中,通过给出多种可提取的特征,能便于用户根据实际模型训练场景中的应用层需求选择其中的一项或多项特征进行入侵检测样本集到离线检测样本集的转化,以适用于不同的模型训练场景,提高入侵检测样本集在模型训练领域中的通用性。
离线测评场景
参照图5所示意的虚线部分,离线测评是指在早期算法开发和验证中,使用入侵检测样本集测试在模型训练场景中训练得到的入侵检测模型的检测效果。由于已经在模型训练场景中提取得到了离线检测样本集,因此,在使用该离线检测样本集中的部分离线检测样本训练得到入侵检测模型后,数据处理装置还可以从离线检测样本集中选取另一部分离线检测样本作为测试集数据,输入至模型训练装置,由模型训练装置利用该测试集数据测试入侵检测模型得到离线测评信息,进而将离线测评信息发送给数据处理装置,由数据处理装置根据离线测评信息评估入侵检测模型的检测效果。
其中,测试集数据中也可以包括攻击类型的离线检测样本和正常类型的离线检测样本,且攻击类型的离线检测样本和正常类型的离线检测样本还可以保持数量的一致性。如此,模型训练装置获取到入侵检测模型针对于每个离线检测样本的检测结果后,结合其中将攻击类型的离线检测样本识别为攻击样本的数量以及将正常类型的离线检测样本识别为非攻击样本的数量得到识别正确的离线检测样本的总数量,以及结合其中将攻击类型的离线检测样本识别为非攻击样本的数量以及将正常类型的离线检测样本识别为攻击样本的数量得到识别错误的离线检测样本的总数量,进而将识别正确的离线检测样本的总数量和识别错误的离线检测样本的总数量携带在离线测评信息中发送给数据处理装置。进而,若识别正确的离线检测样本的总数量越多,识别错误的离线检测样本的总数量越少,则数据处理装置可以确定入侵检测模型的检测效果越好,反之,则检测效果越差。
需要说明的是,上述离线测评的具体实现过程仅是一个示例,实际操作中还可能存在其它的实现方式,比如模型训练装置也可以直接将每个离线检测样本的检测结果作为离线测评信息发送给数据处理装置,由数据处理装置自行统计识别正确的离线检测样本的总数量和识别错误的离线检测样本的总数量后完成离线测评,可能的实现方式有很多,此处不再一一进行列举。
在离线测评场景中,入侵检测样本集能支持在离线状态下测评入侵检测模型的检测效果,如此,能便于模型训练装置根据该检测效果不断优化入侵检测模型,得到检测效果更好的入侵检测模型,为入侵检测模型在待检测设备上的落地提供依据。
在线测评场景
参照图5所示意的双节点线部分,在线测评是指在后期算法部署中,使用入侵检测样本集测试部署有入侵检测模型的待检测设备的检测效果。经过上述模型训练和离线测评后,入侵检测模型能具有较好的检测效果,但该检测效果只是在脱离待检测设备的基础上测得的,无法表征真实应用在待检测设备后的实际效果。因此,还需要将入侵检测模型部署在待检测设备上,并对部署有入侵检测模型的待检测设备进行真实地攻击后,根据待检测设备的反应确定在待检测设备中应用入侵检测模型的实际检测效果。
具体实施中,数据处理装置可以对入侵检测样本集中的入侵检测样本进行格式转换,获得与测试工具格式相匹配的在线检测样本,进而通过测试工具将在线检测样本输入至待检测设备,并获得待检测设备中部署的入侵检测模型对于在线检测样本所产生的在线测评信息,根据该在线测评信息测评部署有入侵检测模型的待检测设备的检测性能。其中,格式转换也可以以连接为单元进行,比如针对于属于同一TCP连接的各个入侵检测样本,首先聚合这些入侵检测样本得到一个初步的在线检测样本后,再将该初步的在线检测样本的格式转换为当前场景下的测试工具所适配的格式,得到一个在线检测样本。而针对于属于CAN(FD)报文或UDP报文的各个入侵检测样本,则可以直接将每个入侵检测样本的格式 转换为当前场景下的测试工具所适配的格式,得到对应的一个在线检测样本。
举例来说,当待检测设备为车辆时,测试工具可以通过车辆的OBD接口实时地将不同的在线检测样本依次注入至车辆,针对于每个在线检测样本,若车辆识别其为攻击样本,则可以通过VCU向云端的安全运营中心(security operations center,SOC)进行报警。进而,在全部的在线检测样本都注入完成后,SOC结合全部的在线检测样本的数量和在此期间车辆的报警记录信息,确定其中识别正确的在线检测样本的总数量和识别错误的在线检测样本的总数量,然后根据这两个总数量生成在线测评信息后发送给数据处理装置。进而,若识别正确的在线检测样本的总数量越多,识别错误的在线检测样本的总数量越少,则数据处理装置可以确定部署有入侵检测模型的待检测设备的检测效果越好,反之,则检测效果越差。
进一步地,考虑到CAN(FD)报文类型的测试注入需要CANoe或PCAN等测试工具,而一台数据的测试注入需要CANoe或Technica等测试工具,因此,本申请实施例中的测试工具具体可以为CANoe、PCAN、Technica或其它能实现在线测评的工具,与此对应的,格式转换后的在线检测样本可以为.PCAP、.ASC、.BLF或其它能实现在线测试的工具所对应的格式,数据处理装置将.pcap格式的入侵检测样本转换为测试工具所对应的格式的在线检测样本后,通过测试工具将该在线检测样本注入待检测设备。如此,通过使入侵检测样本集支持转换为多种测试工具所对应的格式,能便于用户根据实际在线场景中的测试工具选择所对应的格式进行转换,以适用于不同的在线测评场景,提高入侵检测样本集在在线测评领域中的通用性。
在在线测评场景中,入侵检测样本集能支持在在线状态下测评部署有入侵检测模型的待检测设备的检测效果,如此,能便于用户根据该检测效果确定待检测设备的防攻击性能,确保仅出厂防攻击效果较好的待检测设备。
上述内容详细介绍了数据处理装置如何构建和应用入侵检测样本集,除此之外,一些场景中,数据处理装置还可以评估所构建的入侵检测样本集的质量,下面示例性地介绍一种具体的评估过程。
请参照图6所示,其为本申请实施例提供的一种评估入侵检测样本集的流程示意图,该流程包括:
步骤601,数据处理装置获取入侵检测样本集。
步骤602,数据处理装置计算入侵检测样本集在每个预设指标下的值,并根据入侵检测样本集在各个预设指标下的值,计算得到入侵检测样本集对应的评估值。
其中,各个预设指标可以是根据待检测设备所属的体系架构的特性而设置的,示例性地可以包括定量指标和定性指标,定量指标是指可以通过准确的数量进行定义的评估指标,定性指标是指不能直接量化而需要通过其它途径实现量化的评估指标。
举例来说,表2示出了本申请实施例提供的一种可能的预设指标的示意表:
表2
参照表2所示,该示例中,各个预设指标可以包括数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、平衡性指标、特征独立性指标和易用性指标中的一项或多项。其中,数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标和平衡性指标属于定量指标,特征独立性指标和易用性指标属于定性指标。且,为了保持各个预设指标衡量的一致性,各个预设指标的取值范围还可以保持一致,比如均设置为[0,1],也即是,每个预设指标的值可以为0至1中的任一实数,且包括0和1。
下面先对表2所列出的各个预设指标进行详细介绍。
数据冗余指标
数据冗余指标用于指示入侵检测样本集的非冗余程度,示例性地可以表示为入侵检测样本集中的非冗余的入侵检测样本的数量和全部的入侵检测样本的数量的比值。比如,当入侵检测样本集中共存在100个入侵检测样本时,如果其中的5个入侵检测样本均对应同一ID,剩余的95个入侵检测样本对应不同的ID,则该入侵检测样本集在数据冗余指标下的值为95/100。
需要说明的是,数据冗余指标也可以表示为其它形式,只要能保证与冗余的入侵检测样本的数量和全部的入侵检测样本的数量的比值呈负相关即可。如此,当全部的入侵检测样本都不同时,数据冗余指标的值最大,入侵检测样本集中的有效样本的数量最多,样本的充足性最好。随着相同的入侵检测样本的数量增大,数据冗余指标的值逐渐变小,入侵检测样本集中的有效样本的数量逐渐变少,样本的充足性逐渐变差。直至全部的入侵检测样本都相同时,数据冗余指标的值最小,入侵检测样本集中的有效样本的数量最少,样本的充足性最差。
攻击覆盖指标
攻击覆盖指标用于指示入侵检测样本集中已包含的攻击类型的覆盖程度,示例性地可以表示为入侵检测样本集所覆盖的攻击类型的数量与待检测设备中可能存在的全部攻击类型的数量的比值。比如,假设待检测设备中可能存在的全部攻击类型为表1中所示意的25种攻击类型,而入侵检测样本集中的第一类型攻击样本是按照其中的10种攻击类型攻击待检测设备得到的,则该入侵检测样本集在攻击覆盖指标下的值为10/25。
需要说明的是,攻击覆盖指标也可以表示为其它形式,只要能保证与所覆盖的攻击类型的数量与全部可能的攻击类型的数量的比值呈正相关即可。如此,当入侵检测样本集覆盖全部可能的攻击类型时,攻击覆盖指标的值最大,入侵检测样本集中的样本类型最多,样本的多样性最好。随着所覆盖的攻击类型的数量变少,攻击覆盖指标的值逐渐变小,入侵检测样本集中的样本类型逐渐变少,样本的多样性逐渐变差。直至未覆盖任一攻击类型时,攻击覆盖指标的值最小,入侵检测样本集中的样本类型最少,样本的多样性最差。
协议覆盖指标
协议覆盖指标用于指示入侵检测样本集中已包含的通信协议的覆盖程度,示例性地可以表示为入侵检测样本集所覆盖的通信协议的数量与待检测设备中可能存在的全部通信协议的数量的比值。示例性地,请参照表3所示,其为本申请实施例提供的一种车联网领域中可能存在的通信协议的示意表,应理解,随着车联网技术的发展,未来还可能会出现新的通信协议,因此表3中的通信协议还可以随之更新,本申请实施例对此不作具体限定。
表3
假设待检测设备为车辆,车辆中可能存在的全部通信协议为表3中所示意的8种协议,而入侵检测样本集所覆盖的通信协议包括CAN(FD)、DoCAN、DDS、MQTT和HTTP(S),则该入侵检测样本集在协议覆盖指标下的值为5/8。
需要说明的是,协议覆盖指标也可以表示为其它形式,只要能保证与所覆盖的通信协议的数量与全部可能的通信协议的数量的比值呈正相关即可。如此,当入侵检测样本集覆盖全部可能的通信协议时,协议覆盖指标的值最大,该入侵检测样本集中的样本是通过对全部的通信协议消息都进行攻击而得到的,该入侵检测样本集的样本来源最多,样本的丰富性最好。随着所覆盖的协议的数量变少,协议覆盖指标的值逐渐变小,入侵检测样本集中的样本来源逐渐变少,样本的丰富性逐渐变差。直至未覆盖任一协议时,协议覆盖指标的值最小,入侵检测样本集中的样本来源最少,样本的丰富性最差。
业务覆盖指标
业务覆盖指标用于指示入侵检测样本集中已包含的业务的覆盖程度,示例性地可以表示为入侵检测样本集所包含的业务的数量与待检测设备中可能存在的全部业务的数量的比值。比如,假设待检测设备中可能存在的全部业务包括远程控制、日志传输、空中(over-the-air,OTA)更新、诊断服务、视频传输和网络管理这6种业务,而入侵检测样本集所包含的业务为远程控制、OTA更新和诊断服务,则该入侵检测样本集在业务覆盖指标下的值为3/6。
需要说明的是,业务覆盖指标也可以表示为其它形式,只要能保证与所覆盖的业务的数量与全部可能的业务的数量的比值呈正相关即可。如此,当入侵检测样本集覆盖待检测设备中的全量业务时,业务覆盖指标的值最大,样本的业务适用性最好。随着所覆盖的业务的数量变少,业务覆盖指标的值逐渐变小,入侵检测样本集的业务适用性逐渐变差,直至未覆盖任一业务时,协议覆盖指标的值最小,入侵检测样本集的业务适用性最差。
数据标记指标
数据标记指标用于指示入侵检测样本集中的入侵检测样本的可标记程度,示例性地可以表示为入侵检测样本集中已标记的入侵检测样本的数量与全部的入侵检测样本的数量的比值。其中,已标记的入侵检测样本可以包括上述步骤301中获得的真实攻击样本、模拟攻击样本和非攻击样本以及上述步骤302中获得的第二类型攻击样本,未标记的入侵检测样本包括上述步骤301中获得的无标记样本。举例来说,当入侵检测样本集中共存在100个入侵检测样本时,如果其中的20个入侵检测样本为第一类型攻击样本,20个入侵检测样本为第二类型攻击样本,55个入侵检测样本为非攻击样本,5个入侵检测样本为无标记 样本,则该入侵检测样本集中已标记的入侵检测样本的数量为95,因此,该入侵检测样本集在数据标记指标下的值为95/100。
需要说明的是,数据标记指标也可以表示为其它形式,只要能保证与已标记的入侵检测样本的数量与全部的入侵检测样本的总数量的比值呈正相关即可。如此,当入侵检测样本集中的全部的入侵检测样本都已标记时,意味着全部的入侵检测样本都已明确划分为非攻击样本和攻击样本,而不包含不确定的样本,该入侵检测样本集中的样本清晰度最高。随着已标记的入侵检测样本的数量变少,入侵检测样本集中包含的不确定的样本逐渐变多,样本的清晰度逐渐变差,直至全部的入侵检测样本都未标记时,入侵检测样本集中包含最多的不确定样本,样本的清晰度最差。
平衡性指标
平衡性指标用于指示入侵检测样本集中的攻击样本和非攻击样本的均衡程度,示例性地可以表示为入侵检测样本集中的全部入侵检测样本的总数量与攻击样本和非攻击样本的数量差的差值与全部的入侵检测样本的总数量的比值。比如,当入侵检测样本集中共存在100个入侵检测样本时,如果其中的45个入侵检测样本为攻击样本,55个入侵检测样本为非攻击样本,则该入侵检测样本集中的攻击样本和非攻击样本的数量差为10,因此,该入侵检测样本集在平衡性指标下的值可以为(100-10)/100。
需要说明的是,平衡性指标也可以表示为其它形式,只要能保证与攻击样本和非攻击样本的数量差呈负相关即可。如此,当入侵检测样本集中的攻击样本和非攻击样本各占50%时,平衡性指标的值最大,入侵检测样本集中的样本均衡性最好,后续更容易从入侵检测样本集中取出等量的攻击样本和非攻击样本进行测试。随着攻击样本和非攻击样本的数量差变大,入侵检测样本集中的样本均衡性逐渐变差,越不容易从入侵检测样本集中取出等量的攻击样本和非攻击样本进行测试。直至入侵检测样本集中全部为攻击样本或全部为非攻击样本时,入侵检测样本集中的样本均衡性最差,无法从入侵检测样本集中取出等量的攻击样本和非攻击样本进行测试。
特征独立性指标
特征独立性指标用于指示使用入侵检测样本集进行离线测评时所提取的特征的独立性,相互独立的特征越多,则特征独立性指标的取值越大,相互独立的特征越少,则特征独立性指标的取值越小。
需要说明的是,特征是否独立可以由本领域技术人员根据经验进行判断,比如可以由工程师、专家或者第三方机构中的至少两者进行判断,以综合各方经验得到较为准确的评估结果。
易用性指标
易用性指标用于指示入侵检测样本集所能兼容的应用场景的通用性,该应用场景示例性地可以是使用入侵检测样本集进行在线测评时所能支持的测试工具的格式,所能支持的测试工具的格式越多,则易用性指标的取值越大,所能支持的测试工具的格式越少,则易用性指标的取值越小。
需要说明的是,入侵检测样本集所能兼容的应用场景也可以由本领域技术人员根据经验进行判断,比如可以由工程师、专家或者第三方机构中的至少两者进行判断,以综合各方经验得到较为准确的评估结果。
进一步地,按照上述表2所示意的各个预设指标,在计算得到入侵检测样本集在每个 预设指标下的值之后,数据处理装置还可以根据各个预设指标对应的权重对入侵检测样本集在各个预设指标下的值进行加权平均,并将计算得到的加权平均值作为入侵检测样本集的评估值。其中,各个预设指标对应的权重可以相同,也可以不同,比如,一个示例中,可以配置各个定量的预设指标均对应第一权重,各个定性的预设指标均对应第二权重,且第一权重大于第二权重。如此,通过减小定性的预设指标的权重,能降低需要通过人为经验判断的预设指标对入侵检测样本对应的评估值的影响,使得评估过程更加关注理性的评判标准。
步骤603,数据处理装置判断入侵检测样本集对应的评估值是否低于预设阈值,若是,则执行步骤604,若否,则执行步骤605。
步骤604,数据处理装置调节入侵检测样本集,之后执行步骤602。
示例性地,当入侵检测样本集对应的评估值低于预设阈值时,意味着入侵检测样本集的质量不能满足要求,此时,数据处理装置可以根据入侵检测样本集在各个预设指标下的值调节入侵检测样本集,使得调节后的入侵检测样本集在一个或多个预设指标下的值变大,进而增大入侵检测样本集对应的评估值。比如,当入侵检测样本集在平衡性指标下的值较低时,数据处理装置可以通过裁剪数量较多的样本,使得攻击样本和非攻击样本趋近于一致,以提高入侵检测样本集在平衡性指标下的值。或者,当入侵检测样本集在数据冗余指标下的值较低时,数据处理装置可以通过删除冗余的入侵检测样本,使得冗余的入侵检测样本的数量和全部的入侵检测样本的数量的比值减小,以提高入侵检测样本集在数据冗余指标下的值。
需要说明的是,上述内容只是示例性地给出一种可能的调节方式,在实际操作中,还可以由本领域技术人员根据经验选择其它方式调节入侵检测样本集,本申请实施例对此不作具体限定。
步骤605,数据处理装置存储入侵检测样本集。
示例性地,当入侵检测样本集对应的评估值不低于预设阈值时,意味着入侵检测样本集的质量能够满足需求,该情况下,数据处理装置可以以文本格式将入侵检测样本集存储在数据库中。其中,文本格式示例性地可以是.cpap或者其它IDS所支持的格式。
在上述实施方式中,通过参照预设指标对已构建的入侵检测样本集进行评估,能根据评估结果不断优化入侵检测样本集,使优化后的入侵检测样本集更加适用于待检测设备所属的体系架构。且,通过结合定量指标和定性指标综合评判入侵检测样本集的构建质量,还能使评判结果更加全面,且更具有说服性。
应理解,通过预设指标评估入侵检测样本集的质量,这只是一种可选地评估方式,实际操作中还可能存在其它的评估方式,比如也可以通过人为经验直接评估,或者通过第三方机构间接评估,或者通过对比历史入侵检测样本集进行评估等,本申请实施例对此不作具体限定。
基于上述内容,图7示例性示意出本申请实施例提供的一种开发上述数据处理方案的设计架构图,该设计架构图示例性地可以是呈现给开发测试人员或运维人员的界面,由开发测试人员或运维人员根据该界面中需要实现的各个功能编写相应的程序代码。具体的,如图7所示,该设计架构中包括硬件工具层、数据生成层、调用接口层和应用场景层,下面对每个层的内容进行详细地介绍。
硬件工具层负责提供硬件接口工具,该硬件接口工具用于接入待检测设备,为采集待检测设备的流量数据提供支撑。示例性地,当待检测设备为车辆时,硬件工具层主要可以包括面向信号的CAN(FD)工具和面向服务的ETH工具,这两个工具可以接入在车辆的OBD接口上,以支持数据采集模块采集车辆针对于攻击行为所产生的流量数据。
数据生成层负责构建和查看入侵检测样本集,主要包括数据采集模块、AI生成样本模块、特征提取模块、格式转换模块、样本集评估模块和数据查看模块,还可以包括数据库。实施中,数据采集模块可以通过硬件工具层采集待检测设备对于攻击行为所产生的流量数据,通过分析该流量数据,标记得到其中的第一类型攻击样本和非攻击样本,并将第一类型攻击样本和非攻击样本存入数据库。AI生成样本模块可以对数据采集模块标记的第一类型攻击样本添加噪声,通过AI对抗算法处理得到容易被误报的第二类型攻击样本后,将第二类型攻击样本存入数据库。如此,数据库中的第一类型攻击样本、第二类型攻击样本和非攻击样本构成了入侵检测样本集。进而,特征提取模块可以对数据库中的入侵检测样本集进行特征提取后构成离线检测样本集,格式转换模块可以对数据库中的入侵检测样本集进行格式转换后构成在线检测样本集。样本集评估模块可以计算入侵检测样本集在各个预设指标下的评估值,并在评估值低于预设阈值时,调节入侵检测样本集,直至调节到入侵检测样本集对应的评估值不低于预设阈值为止。数据查看模块可以根据开发测试人员或运维人员的命令向其显示部分或全部入侵检测样本。
调用接口层负责提供应用编程接口(application programming interface,API),具体的,可以根据上层应用的使用需求,通过API从数据生成层中调用相应的数据提供给上层应用。比如,调用接口层可以通过API将数据生成层中的离线检测样本集中的部分离线检测样本提供给上层应用以训练入侵检测模型,也可以通过API将数据生成层中的离线检测样本集中的另一部分离线检测样本提供给上层应用以实现对入侵检测模型的离线测评,还可以通过API将数据生成层中的在线检测样本提供给上层应用以实现对部署有入侵检测模型的待检测设备的在线测评。
应用场景层负责与外界设备交互,以将入侵检测样本集应用到各种可能的应用场景中。一些典型的应用场景中,应用场景层可以利用调用接口层提供的离线检测样本在IDS开发中训练入侵检测模型,也可以利用调用接口层提供的在线检测样本在车云运维中测评部署有入侵检测模型的待检测设备的检测效果,也可以利用调用接口层提供的离线检测样本在IDS测试中测评入侵检测模型的检测效果,还可以将调用接口层提供的入侵检测样本集提供给第三方检测认证机构,以便于使待检测设备获得第三方认证机构的认证后顺利出厂,等等。
本申请的上述实施例中,通过结合已知攻击类型的第一类型攻击样本、未知攻击类型的第二类型攻击样本和非攻击样本构建入侵检测样本集,能使入侵检测样本集中的样本类型更加丰富全面,如此,当应用在车联网领域时,即使不同车型的网络拓扑和流量特征存在显著差异,这种构建方式也能针对于每种类型的车型构建出适配且丰富的入侵检测样本,有效提高使用入侵检测样本评估车辆网络安全的准确性。
应理解,本申请提供的数据处理方法还可以推广至任意对网络安全有需求的信息系统中。比如,还可以应用在智能家居领域中,通过攻击智能家居产品并进行加扰以获得丰富的攻击样本,研究智能家居场景下的攻击者画像。或者,还可以应用在工业控制领域中, 通过在数字孪生模型中引入攻击样本研究入侵防御,增强工控系统的鲁棒性。具体的,可以通过云服务器虚拟地构建出一台车辆或者一台车辆上所需测试的部件,进而对这台虚拟的车辆或车辆部件进行上述数据处理操作以获得入侵检测样本集,如此,能在真实组装车辆之前提前获知该车辆的防攻击能力,以便于在确定车辆能较好地防攻击的情况下才实际去造车,有效地节省人力和物力成本。
此外,随着系统架构的演变和新场景的出现,本申请提供的数据处理方法对类似的技术问题,同样适用,本申请对此也不作具体限定。
需要说明的是,上述各个信息的名称仅仅是作为示例,随着通信技术的演变,上述任意信息均可能改变其名称,但不管其名称如何发生变化,只要其含义与本申请上述信息的含义相同,则均落入本申请的保护范围之内。
上述主要从各个网元之间交互的角度对本申请提供的方案进行了介绍。可以理解的是,上述实现各网元为了实现上述功能,其包含了执行各个功能相应的硬件结构和/或软件模块。本领域技术人员应该很容易意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,本发明能够以硬件或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本发明的范围。
根据前述方法,图8为本申请实施例提供的一种数据处理装置的结构示意图,该数据处理装置可以为具有处理能力的任意设备,如服务器,或者也可以为芯片或电路,比如可设置于服务器中的芯片或电路,或者还可以为多个服务器构成的服务器集群。如图8所示,该数据处理装置800可以包括处理器801、存储器802和收发器803,还可以进一步包括总线系统,处理器801、存储器802和收发器803可以通过总线系统相连。
在实现过程中,上述方法的各步骤可以通过处理器801中的硬件的集成逻辑电路或者软件形式的指令完成。结合本申请实施例所公开的方法的步骤可以直接体现为硬件处理器执行完成,或者用处理器801中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器802,处理器801读取存储器802中的信息,结合其硬件完成上述方法的步骤。
应理解,上述处理器801可以是一个芯片。例如,该处理器801可以是现场可编程门阵列(field programmable gate array,FPGA),可以是专用集成芯片(application specific integrated circuit,ASIC),还可以是系统芯片(system on chip,SoC),还可以是中央处理器(central processor unit,CPU),还可以是网络处理器(network processor,NP),还可以是数字信号处理电路(digital signal processor,DSP),还可以是微控制器(micro controller unit,MCU),还可以是可编程控制器(programmable logic device,PLD)或其他集成芯片。
可以理解,本申请实施例中的存储器802可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(read-only memory,ROM)、可编程只读存储器(programmable ROM,PROM)、可擦除可编程只读存储器(erasable PROM,EPROM)、电可擦除可编程只读存储器(electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(random access memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态 随机存取存储器(static RAM,SRAM)、动态随机存取存储器(dynamic RAM,DRAM)、同步动态随机存取存储器(synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(double data rate SDRAM,DDR SDRAM)、增强型同步动态随机存取存储器(enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(synchlink DRAM,SLDRAM)和直接内存总线随机存取存储器(direct rambus RAM,DR RAM)。应注意,本文描述的系统和方法的存储器旨在包括但不限于这些和任意其它适合类型的存储器。
本申请实施例中,存储器802用于存储指令,处理器801用于执行该存储器802存储的指令,实现如上图3、图4或图6中所示的任意一项或任意多项中数据处理装置所对应的方法。具体来说,处理器801通过调用收发器803攻击待检测设备,获得第一类型攻击样本,通过对第一类型攻击样本施加噪声,获得第二类型攻击样本,进而根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。
在一种可选地实施方式中,第一类型攻击样本可以包括真实攻击样本和模拟攻击样本,真实攻击样本是通过人为攻击待检测设备得到的,模拟攻击样本是通过攻击工具攻击待检测设备得到的。
在一种可选地实施方式中,真实攻击样本可以对应如下攻击类型中的一项或多项:身份标识号ID不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻击、ID来源非指定电子控制单元ECU攻击、出现相同ID攻击、控制器局域网络CAN扫描攻击、统一诊断服务UDS执行敏感操作攻击、消息认证错误攻击、ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击、未知的入栈连接攻击。
在一种可选地实施方式中,模拟攻击样本可以对应如下攻击类型中的一项或多项:ID Fuzz攻击、数据Fuzz攻击、CAN的Dos攻击、ETH的Dos攻击、畸形包注入攻击、端口扫描攻击。
在一种可选地实施方式中,处理器801具体用于:遍历预设的多种攻击类型中的每种攻击类型,在遍历每种攻击类型时:通过调用收发器803对待检测设备执行攻击类型对应的攻击行为,并获取待检测设备对于攻击行为所产生的流量数据,若流量数据是攻击流量,则将流量数据标记为一个第一类型攻击样本。
在一种可选地实施方式中,处理器801在获取待检测设备对于攻击行为所产生的流量数据之后,若确定该流量数据是正常流量,则将该流量数据标记为一个非攻击样本,进而根据第一类型攻击样本、第二类型攻击样本和非攻击样本,构建得到入侵检测样本集。
在一种可选地实施方式中,处理器801具体用于:对第一类型攻击样本施加噪声,得到扰动样本,并将扰动样本输入攻击识别模型,获得攻击识别模型输出的识别结果,进而根据识别结果调节扰动样本,直至调节后的扰动样本对应的识别结果指示无法确定调节后的扰动样本是否为攻击样本时,将调节后的扰动样本确定为第二类型攻击样本。其中,识别结果用于指示扰动样本是否属于攻击样本。
在一种可选地实施方式中,处理器801在根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以对入侵检测样本集进行特征提取,获得离线检测样本集,该离线检测样本集用于测评入侵检测模型的检测性能。
在一种可选地实施方式中,处理器801具体用于:确定入侵检测样本集中的每个入侵检测样本的报文类型,对于报文类型为TCP报文的各个入侵检测样本,通过对属于同一 TCP连接的全部入侵检测样本进行特征提取,获得一个离线检测样本,对于报文类型为CAN(FD)报文或UDP报文的各个入侵检测样本,通过对每个CAN(FD)报文或UDP报文的入侵检测样本进行特征提取,获得一个离线检测样本。
在一种可选地实施方式中,特征提取出的特征包括如下特征中的一项或多项:时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。
在一种可选地实施方式中,处理器801在根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以对入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集,进而通过测试工具将在线检测样本集输入待检测设备,该在线检测样本用于测评部署有入侵检测模型的待检测设备的检测性能。
在一种可选地实施方式中,处理器801在根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集之后,还可以根据入侵检测样本集在各个预设指标下的值,确定入侵检测样本集对应的评估值,当评估值低于预设阈值时,调节入侵检测样本集。
在一种可选地实施方式中,预设指标可以包括如下指标中的一项或多项:数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标、平衡性指标、特征独立性指标、易用性指标。
在一种可选地实施方式中,第一类型攻击样本可以是通过攻击如下任一区域得到的:整个待检测设备;待检测设备的一个或多个物理区域;或者,待检测设备的一个或多个功能区域。
该数据处理装置800所涉及的与本申请实施例提供的数据处理装置的技术方案相关的概念,解释和详细说明及其他步骤请参见前述方法或其他实施例中关于这些内容的描述,此处不做赘述。
基于以上实施例以及相同构思,图9为本申请实施例提供的另一种数据处理装置的示意图,该数据处理装置900示例性地可以为上述任一实施例中所述的数据处理装置,也可以为芯片或电路,比如可设置于数据处理装置中的芯片或电路。该数据处理装置900可以实现如上图3、图4或图6中所示的任一项或任多项对应的方法中数据处理装置所执行的步骤。
如图9所示,该数据处理装置900可以包括攻击单元901、扰动单元902和构建单元903,示例性地还可以包括特征提取单元904、格式转换单元905和调节单元906中的一项或多项。攻击单元901用于通过攻击待检测设备,获得第一类型攻击样本;扰动单元902用于通过对第一类型攻击样本施加噪声,获得第二类型攻击样本;构建单元903用于根据第一类型攻击样本和第二类型攻击样本,构建得到入侵检测样本集。一些场景中,特征提取单元904用于:对入侵检测样本集进行特征提取,获得离线检测样本集,该离线检测样本集用于测评入侵检测模型的检测性能。另一些场景中,格式转换单元905用于:对入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集,并通过测试工具将在线检测样本集输入待检测设备,该在线检测样本集用于测评部署有入侵检测模型的待检测设备的检测性能。另一些场景中,调节单元906用于根据入侵检测样本集在各个预设指标下的值,确定入侵检测样本集对应的评估值,当该评估值低于预设阈值时,调节入侵检测样本集。
实施中,攻击单元901通过向待检测设备注入流量,以实现对待检测设备的攻击。该 攻击单元在诸如流量时可以为发送单元、发送器、输出接口、管脚或电路。当数据处理装置900包含存储单元时,该存储单元用于存储计算机指令,攻击单元901、扰动单元902、构建单元903、特征提取单元904、格式转换单元905和调节单元906分别与存储单元通信连接,分别执行存储单元存储的计算机指令,使数据处理装置900可以用于执行上述任一实施例中数据处理装置所执行的方法。其中,攻击单元901、扰动单元902、构建单元903、特征提取单元904、格式转换单元905和调节单元906可以是一个通用中央处理器(CPU),微处理器,特定应用集成电路(application specific intergrated circuit,ASIC)。存储单元为芯片内的存储单元,如寄存器、缓存等,存储单元还可以是数据处理装置900内的位于该芯片外部的存储单元,如只读存储器(read only memory,ROM)或可存储静态信息和指令的其他类型的静态存储设备,随机存取存储器(random access memory,RAM)等。
该数据处理装置900所涉及的与本申请实施例提供的技术方案相关的概念,解释和详细说明及其他步骤请参见前述方法或其他实施例中关于这些内容的描述,此处不做赘述。
应理解,以上数据处理装置900的单元的划分仅仅是一种逻辑功能的划分,实际实现时可以全部或部分集成到一个物理实体上,也可以物理上分开。本申请实施例中,攻击单元901、扰动单元902、构建单元903、特征提取单元904、格式转换单元905和调节单元906可以由上述图8的处理器801实现。
根据本申请实施例提供的数据处理方法,本申请还提供一种数据处理装置,该数据处理装置包括处理器,处理器与存储器相连,存储器用于存储计算机程序,处理器用于执行存储器中存储的计算机程序,以使得数据处理装置实现如图3、图4或图6中任一实施例所述的方法。
根据本申请实施例提供的数据处理方法,本申请还提供一种数据处理装置,该数据处理装置包括处理器和存储器,存储器用于存储计算机程序指令,处理器用于运行计算机程序指令以实现如图3、图4或图6中任一实施例所述的方法。
根据本申请实施例提供的数据处理方法,本申请还提供一种芯片,该芯片可以包括处理器和接口,处理器用于通过接口读取指令,以执行如图3、图4或图6中任一实施例所述的方法。
根据本申请实施例提供的数据处理方法,本申请还提供一种数据处理系统,该数据处理系统可以包括前述的待检测设备和数据处理装置。
根据本申请实施例提供的数据处理方法,本申请还提供一种计算机可读存储介质,该计算机可读存储介质存储有计算机程序,当计算机程序被运行时,实现如图3、图4或图6中任一实施例所述的方法。
根据本申请实施例提供的数据处理方法,本申请还提供一种计算机程序产品,当该计算机程序产品在处理器上运行时,实现如图3、图4或图6中任一实施例所述的方法。
在本说明书中使用的术语“部件”、“单元”、“系统”等用于表示计算机相关的实体、硬件、固件、硬件和软件的组合、软件、或执行中的软件。例如,部件可以是但不限于,在处理器上运行的进程、处理器、对象、可执行文件、执行线程、程序和/或计算机。通过图示,在计算设备上运行的应用和计算设备都可以是部件。一个或多个部件可驻留在进程和/或执行线程中,部件可位于一个计算机上和/或分布在两个或更多个计算机之间。此外,这些部件可从在上面存储有各种数据结构的各种计算机可读介质执行。部件可例如根据具有一个或多个数据分组(例如来自与本地系统、分布式系统和/或网络间的另一部件交互的 二个部件的数据,例如通过信号与其它系统交互的互联网)的信号通过本地和/或远程进程来通信。
本领域普通技术人员可以意识到,结合本文中所公开的实施例描述的各种说明性逻辑块(illustrative logical block)和步骤(step),能够以电子硬件、或者计算机软件和电子硬件的结合来实现。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统、装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(read-only memory,ROM)、随机存取存储器(random access memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。

Claims (31)

  1. 一种数据处理方法,其特征在于,包括:
    通过攻击待检测设备,获得第一类型攻击样本;
    通过对所述第一类型攻击样本施加噪声,获得第二类型攻击样本;
    根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵检测样本集。
  2. 如权利要求1所述的方法,其特征在于,所述第一类型攻击样本包括真实攻击样本和模拟攻击样本,所述真实攻击样本是通过人为攻击所述待检测设备得到的,所述模拟攻击样本是通过攻击工具攻击所述待检测设备得到的。
  3. 如权利要求2所述的方法,其特征在于,所述真实攻击样本对应如下攻击类型中的一项或多项:
    身份标识号ID不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻击、ID来源非指定电子控制单元ECU攻击、出现相同ID攻击、控制器局域网络CAN扫描攻击、统一诊断服务UDS执行敏感操作攻击、消息认证错误攻击、ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击、未知的入栈连接攻击。
  4. 如权利要求2或3所述的方法,其特征在于,所述模拟攻击样本对应如下攻击类型中的一项或多项:
    ID模糊Fuzz攻击、数据Fuzz攻击、CAN的拒绝服务Dos攻击、以太网的Dos攻击、畸形包注入攻击、端口扫描攻击。
  5. 如权利要求1至4中任一项所述的方法,其特征在于,所述通过攻击待检测设备,获得第一类型攻击样本,包括:
    遍历预设的多种攻击类型中的每种攻击类型,在遍历所述每种攻击类型时:
    对所述待检测设备执行所述攻击类型对应的攻击行为;
    获取所述待检测设备对于所述攻击行为所产生的流量数据;
    若所述流量数据是攻击流量,则将所述流量数据标记为一个所述第一类型攻击样本。
  6. 如权利要求5所述的方法,其特征在于,
    所述获取所述待检测设备对于所述攻击行为所产生的流量数据之后,还包括:
    若所述流量数据是正常流量,则将所述流量数据标记为一个非攻击样本;
    所述根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵检测样本集,包括:
    根据所述第一类型攻击样本、所述第二类型攻击样本和所述非攻击样本,构建得到所述入侵检测样本集。
  7. 如权利要求1至6中任一项所述的方法,其特征在于,所述通过对所述第一类型攻击样本施加噪声,得到第二类型攻击样本,包括:
    对所述第一类型攻击样本施加噪声,得到扰动样本;
    将所述扰动样本输入攻击识别模型,获得所述攻击识别模型输出的识别结果,所述识别结果用于指示所述扰动样本是否属于攻击样本;
    根据所述识别结果调节所述扰动样本,直至调节后的所述扰动样本对应的识别结果指示无法确定调节后的所述扰动样本是否为所述攻击样本时,将调节后的所述扰动样本确定 为所述第二类型攻击样本。
  8. 如权利要求1至7中任一项所述的方法,其特征在于,所述根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵检测样本集之后,还包括:
    对所述入侵检测样本集进行特征提取,获得离线检测样本集;
    其中,所述离线检测样本集用于测评入侵检测模型的检测性能。
  9. 如权利要求8所述的方法,其特征在于,所述对所述入侵检测样本集进行特征提取,获得离线检测样本集,包括:
    确定所述入侵检测样本集中的每个入侵检测样本的报文类型;
    对于报文类型为传输控制协议TCP报文的各个入侵检测样本,通过对属于同一TCP连接的全部入侵检测样本进行特征提取,获得一个离线检测样本;
    对于报文类型为CAN(FD)报文或UDP报文的各个入侵检测样本,通过对每个所述CAN(FD)报文或UDP报文的入侵检测样本进行特征提取,获得一个所述离线检测样本。
  10. 如权利要求8或9所述的方法,其特征在于,所述特征提取出的特征包括如下特征中的一项或多项:
    时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。
  11. 如权利要求1至10中任一项所述的方法,其特征在于,所述根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵检测样本集之后,还包括:
    对所述入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集;
    通过所述测试工具将所述在线检测样本集输入所述待检测设备;
    其中,所述在线检测样本集用于测评部署有入侵检测模型的所述待检测设备的检测性能。
  12. 如权利要求1至11中任一项所述的方法,其特征在于,所述根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵检测样本集之后,还包括:
    根据所述入侵检测样本集在各个预设指标下的值,确定所述入侵检测样本集对应的评估值,当所述评估值低于预设阈值时,调节所述入侵检测样本集。
  13. 如权利要求12所述的方法,其特征在于,所述预设指标包括如下指标中的一项或多项:
    数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标、平衡性指标、特征独立性指标、易用性指标。
  14. 如权利要求1至13中任一项所述的方法,其特征在于,所述第一类型攻击样本是通过攻击如下任一区域得到的:
    整个所述待检测设备;
    所述待检测设备的一个或多个物理区域;
    或者,
    所述待检测设备的一个或多个功能区域。
  15. 一种数据处理装置,其特征在于,包括:
    攻击单元,用于通过攻击待检测设备,获得第一类型攻击样本;
    扰动单元,用于通过对所述第一类型攻击样本施加噪声,获得第二类型攻击样本;
    构建单元,用于根据所述第一类型攻击样本和所述第二类型攻击样本,构建得到入侵 检测样本集。
  16. 如权利要求15所述的装置,其特征在于,所述第一类型攻击样本包括真实攻击样本和模拟攻击样本,所述真实攻击样本是通过人为攻击所述待检测设备得到的,所述模拟攻击样本是通过攻击工具攻击所述待检测设备得到的。
  17. 如权利要求16所述的装置,其特征在于,所述真实攻击样本对应如下攻击类型中的一项或多项:
    身份标识ID不存在攻击、重放攻击、篡改攻击、数据长度错误攻击、信号超出定义范围攻击、上下文错误攻击、ID来源非指定电子控制单元ECU攻击、出现相同ID攻击、控制器局域网络CAN扫描攻击、统一诊断服务UDS执行敏感操作攻击、消息认证错误攻击、ECU身份欺骗攻击、中间人攻击、ECU认证错误攻击、暴力破解攻击、应用层协议错误攻击、未知的出栈连接攻击、未知的入栈连接攻击。
  18. 如权利要求16或17所述的装置,其特征在于,所述模拟攻击样本对应如下攻击类型中的一项或多项:
    ID模糊Fuzz攻击、数据Fuzz攻击、CAN的拒绝服务Dos攻击、以太网的Dos攻击、畸形包注入攻击、端口扫描攻击。
  19. 如权利要求15至18中任一项所述的装置,其特征在于,所述攻击单元具体用于:
    遍历预设的多种攻击类型中的每种攻击类型,在遍历所述每种攻击类型时:
    对所述待检测设备执行所述攻击类型对应的攻击行为;
    获取所述待检测设备对于所述攻击行为所产生的流量数据;
    若所述流量数据是攻击流量,则将所述流量数据标记为一个所述第一类型攻击样本。
  20. 如权利要求19所述的装置,其特征在于,
    所述攻击单元在获取所述待检测设备对于所述攻击行为所产生的流量数据之后,还用于:若所述流量数据是正常流量,则将所述流量数据标记为一个非攻击样本;
    所述构建单元具体用于:根据所述第一类型攻击样本、所述第二类型攻击样本和所述非攻击样本,构建得到所述入侵检测样本集。
  21. 如权利要求15至20中任一项所述的装置,其特征在于,所述扰动单元具体用于:
    对所述第一类型攻击样本施加噪声,得到扰动样本;
    将所述扰动样本输入攻击识别模型,获得所述攻击识别模型输出的识别结果,所述识别结果用于指示所述扰动样本是否属于攻击样本;
    根据所述识别结果调节所述扰动样本,直至调节后的所述扰动样本对应的识别结果指示无法确定所述扰动样本是否为所述攻击样本时,将调节后的所述扰动样本确定为所述第二类型攻击样本。
  22. 如权利要求15至21中任一项所述的装置,其特征在于,还包括特征提取单元;
    所述特征提取单元用于:
    对所述入侵检测样本集进行特征提取,获得离线检测样本集;
    其中,所述离线检测样本集输用于测评入侵检测模型的检测效果。
  23. 如权利要求22所述的装置,其特征在于,所述特征提取单元具体用于:
    确定所述入侵检测样本集中的每个入侵检测样本的报文类型;
    对于报文类型为传输控制协议TCP报文的各个入侵检测样本,通过对属于同一TCP连接的全部入侵检测样本进行特征提取,获得一个离线检测样本;
    对于报文类型为CAN(FD)报文或UDP报文的各个入侵检测样本,通过对每个所述CAN(FD)报文或UDP报文的入侵检测样本进行特征提取,获得一个所述离线检测样本。
  24. 如权利要求22或23所述的装置,其特征在于,所述特征提取出的特征包括如下特征中的一项或多项:
    时间戳、频率特征、协议类型、内容特征、丢包率、错误包数量、连接持续时间、连接发起方、连接接收方。
  25. 如权利要求15至24中任一项所述的装置,其特征在于,还包括格式转换单元;
    所述格式转换单元用于:
    对所述入侵检测样本集进行格式转换,获得与测试工具格式相匹配的在线检测样本集;
    通过所述测试工具将所述在线检测样本集输入所述待检测设备;
    其中,所述在线检测样本集用于测评部署有入侵检测模型的所述待检测设备的检测性能。
  26. 如权利要求15至25中任一项所述的装置,其特征在于,还包括调节单元;
    所述调节单元用于:
    根据所述入侵检测样本集在各个预设指标下的值,确定所述入侵检测样本集对应的评估值,当所述评估值低于预设阈值时,调节所述入侵检测样本集。
  27. 如权利要求26所述的装置,其特征在于,所述预设指标包括如下指标中的一项或多项:
    数据冗余指标、攻击覆盖指标、协议覆盖指标、业务覆盖指标、数据标记指标、平衡性指标、特征独立性指标、易用性指标。
  28. 如权利要求15至27中任一项所述的装置,其特征在于,所述第一类型攻击样本是通过攻击如下任一区域得到的:
    整个所述待检测设备;
    所述待检测设备的一个或多个物理区域;
    或者,
    所述待检测设备的一个或多个功能区域。
  29. 一种数据处理装置,其特征在于,包括处理器和存储器,所述存储器存储计算机程序指令,所述处理器运行所述计算机程序指令以实现如权利要求1至权利要求14中任一项所述的方法。
  30. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质存储有计算机程序,当所述计算机程序被运行时,实现如权利要求1至权利要求14中任一项所述的方法。
  31. 一种计算机程序产品,其特征在于,当所述计算机程序产品在处理器上运行时,实现如权利要求1至权利要求14中任一项所述的方法。
PCT/CN2023/078600 2023-02-28 2023-02-28 一种数据处理方法、装置、存储介质及程序产品 Ceased WO2024178581A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN202380070132.9A CN119968810A (zh) 2023-02-28 2023-02-28 一种数据处理方法、装置、存储介质及程序产品
PCT/CN2023/078600 WO2024178581A1 (zh) 2023-02-28 2023-02-28 一种数据处理方法、装置、存储介质及程序产品

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/078600 WO2024178581A1 (zh) 2023-02-28 2023-02-28 一种数据处理方法、装置、存储介质及程序产品

Publications (1)

Publication Number Publication Date
WO2024178581A1 true WO2024178581A1 (zh) 2024-09-06

Family

ID=92589081

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/078600 Ceased WO2024178581A1 (zh) 2023-02-28 2023-02-28 一种数据处理方法、装置、存储介质及程序产品

Country Status (2)

Country Link
CN (1) CN119968810A (zh)
WO (1) WO2024178581A1 (zh)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240370567A1 (en) * 2023-05-01 2024-11-07 Dell Products L.P. Simulating a ransomware attack in a testing environment
CN119420535A (zh) * 2024-10-31 2025-02-11 海南大学 基于cnn-lstm的水声通信入侵检测方法及系统
CN119520339A (zh) * 2024-12-12 2025-02-25 长城汽车股份有限公司 一种设备测试方法、装置及系统
CN119788394A (zh) * 2024-12-31 2025-04-08 哈尔滨工程大学 一种面向网站指纹攻击的少样本数据增强方法、系统、介质及程序产品
CN120165985A (zh) * 2025-05-19 2025-06-17 西北工业大学 一种基于嵌入式平台的can总线入侵检测方法和系统

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106656981A (zh) * 2016-10-21 2017-05-10 东软集团股份有限公司 网络入侵检测方法和装置
CN114707572A (zh) * 2022-02-24 2022-07-05 浙江工业大学 一种基于损失函数敏感度的深度学习样本测试方法与装置
CN115129607A (zh) * 2022-07-19 2022-09-30 中国电力科学研究院有限公司 电网安全分析机器学习模型测试方法、装置、设备及介质
JP2022163431A (ja) * 2021-04-14 2022-10-26 株式会社日立製作所 計算機システム及び予測プログラムの評価方法
CN115712893A (zh) * 2021-08-20 2023-02-24 华为技术有限公司 一种攻击检测方法及装置

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106656981A (zh) * 2016-10-21 2017-05-10 东软集团股份有限公司 网络入侵检测方法和装置
JP2022163431A (ja) * 2021-04-14 2022-10-26 株式会社日立製作所 計算機システム及び予測プログラムの評価方法
CN115712893A (zh) * 2021-08-20 2023-02-24 华为技术有限公司 一种攻击检测方法及装置
CN114707572A (zh) * 2022-02-24 2022-07-05 浙江工业大学 一种基于损失函数敏感度的深度学习样本测试方法与装置
CN115129607A (zh) * 2022-07-19 2022-09-30 中国电力科学研究院有限公司 电网安全分析机器学习模型测试方法、装置、设备及介质

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240370567A1 (en) * 2023-05-01 2024-11-07 Dell Products L.P. Simulating a ransomware attack in a testing environment
US12437078B2 (en) * 2023-05-01 2025-10-07 Dell Products L.P. Simulating a ransomware attack in a testing environment
CN119420535A (zh) * 2024-10-31 2025-02-11 海南大学 基于cnn-lstm的水声通信入侵检测方法及系统
CN119520339A (zh) * 2024-12-12 2025-02-25 长城汽车股份有限公司 一种设备测试方法、装置及系统
CN119788394A (zh) * 2024-12-31 2025-04-08 哈尔滨工程大学 一种面向网站指纹攻击的少样本数据增强方法、系统、介质及程序产品
CN120165985A (zh) * 2025-05-19 2025-06-17 西北工业大学 一种基于嵌入式平台的can总线入侵检测方法和系统

Also Published As

Publication number Publication date
CN119968810A (zh) 2025-05-09

Similar Documents

Publication Publication Date Title
WO2024178581A1 (zh) 一种数据处理方法、装置、存储介质及程序产品
US20200186560A1 (en) System and method for time based anomaly detection in an in-vehicle communication network
US11115433B2 (en) System and method for content based anomaly detection in an in-vehicle communication network
CN109802953B (zh) 一种工控资产的识别方法及装置
CN111770069B (zh) 一种基于入侵攻击的车载网络仿真数据集生成方法
CN115801396B (zh) 一种为每个标识符建立指纹的车辆入侵检测方法及相关装置
CN112822223B (zh) 一种dns隐蔽隧道事件自动化检测方法、装置和电子设备
CN118337487B (zh) 一种基于大数据的安全网络信息智能化控制方法及系统
CN112187812A (zh) 应用于在线办公网络的数据安全检测方法及系统
WO2024007615A1 (zh) 模型训练方法、装置及相关设备
Jo et al. Automatic whitelist generation system for ethernet based in-vehicle network
Lee et al. A Comprehensive Analysis of Datasets for Automotive Intrusion Detection Systems.
CN117633665B (zh) 一种网络数据监控方法及系统
Zou et al. Using explainable AI for neural network-based network attack detection
CN117729540A (zh) 一种基于统一边缘计算框架的感知设备云边安全管控方法
CN117391214A (zh) 模型训练方法、装置及相关设备
CN113098852A (zh) 一种日志处理方法及装置
Tavasoli et al. Trust-aware federated defense against data poisoning in ml-driven ids for cavs
Kneib A survey on sender identification methodologies for the controller area network
CN115840965B (zh) 一种信息安全保障模型训练方法和系统
Varghese et al. Novel CAN bus fuzzing framework for finding vulnerabilities in automotive systems
CN113420791B (zh) 边缘网络设备接入控制方法、装置及终端设备
CN112118256B (zh) 工控设备指纹归一化方法、装置、计算机设备及存储介质
CN110177373B (zh) 接入点身份鉴别方法及装置、电子设备、计算机可读存储介质
CN111565187B (zh) 一种dns异常检测方法、装置、设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23924559

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 202380070132.9

Country of ref document: CN

WWP Wipo information: published in national office

Ref document number: 202380070132.9

Country of ref document: CN

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23924559

Country of ref document: EP

Kind code of ref document: A1