WO2022089202A1 - 故障识别模型训练方法、故障识别方法、装置及电子设备 - Google Patents
故障识别模型训练方法、故障识别方法、装置及电子设备 Download PDFInfo
- Publication number
- WO2022089202A1 WO2022089202A1 PCT/CN2021/123363 CN2021123363W WO2022089202A1 WO 2022089202 A1 WO2022089202 A1 WO 2022089202A1 CN 2021123363 W CN2021123363 W CN 2021123363W WO 2022089202 A1 WO2022089202 A1 WO 2022089202A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- fault
- series data
- log
- time series
- period
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/30—Monitoring
- G06F11/34—Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
- G06F11/3466—Performance evaluation by tracing or monitoring
- G06F11/3476—Data logging
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/35—Clustering; Classification
- G06F16/355—Creation or modification of classes or clusters
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
- G06F18/2415—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on parametric or probabilistic models, e.g. based on likelihood ratio or false acceptance rate versus a false rejection rate
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
Definitions
- the present application relates to the field of computer technology, and relates to a fault identification model training method, a fault identification method, an apparatus and an electronic device.
- Database as a Service is a service solution that provides database resources to one or more tenants in the form of standard services based on traditional database technology.
- the embodiments of the present application expect to provide a fault identification model training method, a fault identification method, an apparatus, and an electronic device.
- the embodiment of the present application provides a method for training a fault identification model, including:
- the set two-class model is trained to obtain a fault identification model corresponding to the set fault;
- the time series data includes first time series data and second time series data
- the first time series data represents the amount of log information output at each moment of the corresponding first period
- the second time series data represents monitoring data of each performance indicator of the at least one performance indicator at each moment of the corresponding first time period.
- the embodiment of the present application also provides a fault identification method, including:
- At least one first set fault is determined based on the relevant performance index corresponding to each set fault in the at least one set fault; the relevant performance index of the first set fault includes the first performance index;
- the first set fault with the highest confidence in all the identification results is determined as the current fault;
- the end moment of the first time period is the alarm moment corresponding to the first performance indicator
- the time series data corresponding to the first time period includes the first time series data corresponding to the real-time log and the second time series data corresponding to each performance index in the at least one performance index;
- the fault identification model is obtained by training based on any of the above methods for training a fault identification model.
- the fault identification method further includes:
- each server in the at least one server Based on the historical failures of each server in the at least one server and the number of occurrences corresponding to each historical failure, determine the health scores of all the servers in the at least one server, and deploy them in the servers whose health scores are greater than or equal to the set threshold new database instance.
- the embodiment of the present application also provides a device for training a fault identification model, including:
- a construction unit configured to construct at least one positive sample and at least one negative sample corresponding to the set fault based on the time series data corresponding to each first period in the at least one first period;
- a training unit configured to train the set two-class model based on at least one positive sample and at least one negative sample corresponding to the set fault to obtain a fault identification model corresponding to the set fault;
- the time series data includes first time series data and second time series data
- the first time series data represents the amount of log information output at each moment of the corresponding first period
- the second time series data represents monitoring data of each performance indicator of the at least one performance indicator at each moment of the corresponding first time period.
- the embodiment of the present application also provides a fault identification device, including:
- a first determining unit configured to determine the time series data corresponding to the first time period in the case that the data of the first performance index is detected to trigger an alarm
- the second determination unit is configured to determine at least one first set fault based on the relevant performance index corresponding to each set fault in the at least one set fault; the relevant performance index of the first set fault includes the first performance index;
- the first identification unit is configured to input the time series data corresponding to the first time period into the fault identification model corresponding to each first set fault in the at least one first set fault, and obtain the output of each fault identification model. identification results;
- the second identification unit is configured to determine the first set fault with the highest confidence in all the identification results as the currently occurring fault; wherein,
- the end moment of the first time period is the alarm moment corresponding to the first performance indicator
- the time series data corresponding to the first time period includes the first time series data corresponding to the real-time log and the second time series data corresponding to each performance index in the at least one performance index;
- the fault identification model is trained based on any method for training a fault identification model.
- Embodiments of the present application also provide an electronic device, including: a processor and a memory configured to store a computer program that can be executed on the processor,
- the processor is configured to execute at least one of the following when running the computer program:
- Embodiments of the present application further provide a storage medium on which a computer program is stored, and when the computer program is executed by a processor, at least one of the following:
- At least one positive sample and at least one negative sample corresponding to the set fault are constructed based on the time series data corresponding to each first time period in the at least one first time period; based on the at least one positive sample corresponding to the set fault and at least one negative sample to train the set two-class model to obtain a fault identification model corresponding to the set fault. Since the training samples are constructed based on the time series data corresponding to the first time period, the first time series data included in the time series data corresponding to the first time period represents the amount of log information output at each moment of the first time period, and the log information amount refers to the amount of log information.
- the amount of information is used to measure the amount of information conveyed by the log; it is not necessary to analyze the specific content of the historical log when determining the amount of log information, that is, the server does not need to pay attention to the specific content of the historical log when constructing training samples, so it can be omitted.
- the time consumed by semantic analysis of the text content of the log improves the efficiency of obtaining training samples, thereby improving the training efficiency of the fault identification model.
- the first time series data represents the amount of log information output at each moment, and the amount of log information is not determined when determining the amount of log information.
- the text content of the real-time log needs to be semantically analyzed, so time can be saved and the efficiency of fault identification can be improved.
- FIG. 1 is a schematic flowchart of the implementation of a method for training a fault identification model provided by an embodiment of the present application
- FIG. 2 is a schematic flowchart of an implementation of determining first time series data in a method for training a fault identification model provided by an embodiment of the present application;
- FIG. 3 is a schematic diagram of an implementation flow of determining positive samples and negative samples in a method for training a fault identification model provided by an embodiment of the present application;
- FIG. 4 is a schematic flowchart of the implementation of determining a positive sample in a method for training a fault identification model provided by an embodiment of the present application
- FIG. 5 is a schematic diagram of an implementation flowchart of determining a first indicator in a method for training a fault identification model provided by an embodiment of the present application
- FIG. 6 is a schematic diagram of an implementation flowchart of a method for training a fault identification model provided by an application embodiment of the present application
- FIG. 7 is a schematic flowchart of the implementation of a fault identification method provided by an embodiment of the present application.
- FIG. 8 is a schematic flowchart of the implementation of a fault identification method provided by another embodiment of the present application.
- FIG. 9 is a schematic structural diagram of an apparatus for training a fault identification model provided by an embodiment of the present application.
- FIG. 10 is a schematic structural diagram of a fault identification device provided by an embodiment of the present application.
- FIG. 11 is a schematic structural diagram of a hardware composition of an electronic device provided by an embodiment of the present application.
- the embodiments of the present application provide a fault identification method, which performs fault identification based on the amount of log information and monitoring data of performance indicators. Since the specific content of the log does not need to be analyzed when determining the amount of log information, the time consumed by semantic analysis of the text content of the log can be saved, and the efficiency of fault identification can be improved.
- FIG. 1 shows a schematic diagram of an implementation flow of a method for training a fault identification model provided by an embodiment of the present application.
- the execution subject of the method for training a fault identification model is an electronic device, such as a computer, a server, and the like.
- the method for training a fault identification model provided by an embodiment of the present application includes:
- S101 Construct at least one positive sample and at least one negative sample corresponding to the set fault based on time series data corresponding to each first time period in at least one first time period; wherein the time series data includes first time series data and second time series data data; the first time series data represents the amount of log information output at each moment of the corresponding first time period; the second time series data represents the monitoring data of each performance index of the at least one performance index at each time of the corresponding first time period.
- the first period includes a period in which the set failure occurs and a period in which the set failure does not occur.
- the electronic device constructs a positive sample based on the first time series data and the second time series data corresponding to the time period in which the set failure occurs in the first time period; based on the first time series data and the second time series corresponding to the time period in the first time period when the set failure does not occur Data construct negative samples.
- the electronic device can determine the time period when the set failure occurs and the time period when the set failure does not occur based on the alarm time corresponding to the performance index that first triggers the alarm when the set failure occurs.
- the time period in which the set failure occurs includes the first alarm moment; the time period in which the set failure does not occur is the time period in the corresponding first time period except the time period in which the set failure occurs.
- the first performance index of the at least one performance index corresponding to the set fault triggers an alarm at the first moment in the first time period
- the electronic device may determine the second time before the first time as the time period in which the set fault occurs.
- the third time after the first time is determined as the end time of the period in which the set failure occurs.
- the second time is after the start time of the corresponding first time period
- the third time is before the end time of the corresponding first time period.
- the first duration and the second duration may be the same or different.
- the first duration is the duration between the first moment and the second moment
- the second duration is the duration between the first moment and the third moment.
- the first duration and the second duration may be 15 minutes, and of course, they may also be set according to actual conditions.
- the first time series data is the amount of log information recorded in chronological order.
- the second time series data is the monitoring data of each performance index in the at least one performance index recorded in chronological order.
- Each performance index corresponds to a set of second time series data.
- the first time series data is obtained based on the historical log of the DBaaS server.
- the second time series data is obtained based on monitoring data of performance indicators of the DBaaS server.
- At least one performance indicator is used to monitor the occurrence of set failures.
- the set failure can be connection saturation, machine disk failure, memory failure, too many slow queries, etc.
- the log information volume refers to the information volume of the log.
- the log information volume is used to measure the amount of information in the log.
- the amount of log information output at each moment is the sum of the information amount of each log in all logs printed at that moment.
- the information content of a log is determined based on the first probability and at least one second probability.
- the first probability represents the probability of occurrence of the log level corresponding to the log.
- the second probability represents the probability that each set phrase of all set phrases included in the log appears at the log level corresponding to the log.
- log levels may include FATAL, ERROR, WARN, INFO, DEBUG. in,
- the FATAL level log characterizes every serious error event that will cause the application to exit.
- the ERROR level log indicates that although an error event occurs, it still does not affect the continued operation of the system.
- WARN-level logs indicate potential error events in the system.
- INFO level logs characterize the operation of the application.
- the DEBUG level log is mainly used to understand the running status of the system in more detail when debugging the application.
- S102 Based on at least one positive sample and at least one negative sample corresponding to the set fault, train a set two-class model to obtain a fault identification model corresponding to the set fault.
- the electronic device converts each positive sample in the at least one positive sample into a corresponding first vector, and converts each negative sample in the at least one negative sample into a corresponding second vector; converts the at least one first vector and the at least one first vector into a corresponding second vector.
- the two vectors are input into the set two-class model for training, and the fault identification model corresponding to the set fault is obtained.
- the set convergence condition may be that the difference between the first model parameter and the second model parameter is less than or equal to the set threshold.
- other convergence conditions may also be set, for example, the number of training times reaches the set number of times.
- the first model parameter represents the model parameter corresponding to the k-th iterative training
- the second model parameter represents the model parameter corresponding to the k-1-th iterative training
- k is an integer greater than or equal to 1.
- the set two-class model is usually a simple convolutional neural network, for example, a layer of convolution A convolutional neural network consisting of layers and hidden layers.
- the set binary classification model can also be a decision tree model.
- At least one positive sample and at least one negative sample corresponding to the set fault are constructed; At least one positive sample and at least one negative sample corresponding to the set fault are trained on the set two-class model to obtain a fault identification model corresponding to the set fault.
- the first time series data represents the amount of log information output at each moment of the corresponding first period, and the amount of log information refers to the amount of information in the log, it is used to measure the amount of information conveyed by the log; no analysis is required to determine the amount of log information
- the specific content of the historical log can therefore save the time spent on semantic analysis of the text content of the log, improve the efficiency of obtaining training samples, and thus improve the training efficiency of the fault identification model.
- FIG. 2 shows a schematic flowchart of an implementation of determining the first time series data in a method for training a fault identification model provided by an embodiment of the present application. 2, the method for determining the first time series data includes:
- S201 Calculate the total amount of information corresponding to each log based on the first amount of information and the second amount of information corresponding to each log printed at each moment in the first period of time; wherein the first amount of information represents the corresponding amount of the log The first probability of occurrence of the log level; the second information amount represents the second probability of each set phrase appearing in the corresponding log level in all set phrases included in the log.
- the electronic device calculates the sum of the first information amount and the second information amount corresponding to each log to obtain the total information amount of the corresponding log.
- the total amount of information corresponding to each log can be calculated based on the following formula:
- H(X) represents the total information volume of a log
- H(x t ) represents the first information volume of the log
- H(x c ) represents the second information volume of the log
- P(x l ) represents the log corresponding to the log
- x l ) represents the second probability of the occurrence of the i-th set phrase in the log level x l corresponding to the log
- n represents the probability of the set phrase contained in the corresponding log quantity.
- the first probability and the second probability are determined as follows:
- a second probability corresponding to each set phrase in the corresponding log level is determined;
- the at least one first log sample is obtained by sampling historical logs.
- the electronic device samples the collected historical logs to obtain at least one first log sample.
- the collected history logs may include part or all of the history logs output in each first period in at least one first period, and the collected history logs may not include the history logs output in the first period.
- each first log sample in the at least one first log sample satisfies the following conditions:
- the first probability corresponding to the log level corresponding to the first log sample satisfies the set condition
- the alarm type corresponding to the first log sample is the alarm type monitored when the set fault occurs;
- the difference between the third probability and the fourth probability is less than or equal to the set threshold; wherein the third probability represents the probability that the alarm type corresponding to the first log sample appears in the at least one first log sample; the The fourth probability represents the probability that the alarm type corresponding to the first log sample is monitored when the set fault occurs.
- the setting condition is represented as a probability range corresponding to each log level in all log levels configured by the set failure. Since the source code of the setting component in the database determines which log level is output when a setting failure occurs, and determines the probability of each log level, the electronic device can analyze the source code of the setting component in the database , determine the probability range corresponding to each log level corresponding to the log output when the fault is set, and set the above setting conditions based on the determined probability range.
- All alarm types corresponding to the first log sample are determined by setting phrases in the logs of the ERROR level in the first log sample; a log contains at least one setting phrase; one setting phrase corresponds to one alarm type.
- the electronic device determines the first probability and the second probability based on all the first log samples in the at least one first log sample, and the details are as follows:
- the electronic device Based on the log level corresponding to each log in the at least one log included in the first log sample, the electronic device counts the number of occurrences of each log level in all the first log samples; The number of times is summed to obtain the first total number of occurrences of all log levels in all first log samples; based on the number of occurrences of each log level in all first log samples, and based on the occurrence of all log levels in all first log samples The first total number of times is calculated, and the first probability corresponding to each log level is calculated. The first probability is the quotient of the number of occurrences of the corresponding log level and the first total number of times.
- the electronic device Based on the word segmentation technology, the electronic device performs word segmentation on each log in the first log sample, obtains a word segmentation result, filters the word segmentation result, and extracts a set phrase corresponding to each log. Filtering the word segmentation result includes removing connective words, common words (eg, end, run, etc.), the name of the component, the name of the called script, and the like.
- the set phrase represents the alarm type (or alarm event) corresponding to the set fault. For example, when the fault is set as the number of connections is saturated, the set phrases can include: operation timeout, checkWillUpdateKafka error, etc.
- the electronic device counts the number of occurrences of each set phrase in the corresponding log level based on the set phrase corresponding to each log in all the first log samples; the number of occurrences of each set phrase in the corresponding log level Perform a summation operation to count the second total number of occurrences in the log levels corresponding to all the set phrases; based on the number of times each set phrase appears in the corresponding log level, and based on the corresponding second total number of times, determine The second probability that each set phrase appears in the corresponding log level.
- One log corresponds to a set log level, and the second probability is the quotient of the number of times each set phrase appears in the corresponding log level and the corresponding second total number of times.
- the first log sample satisfies the above-mentioned three conditions
- the first log sample can accurately reflect the running situation of the server when the set failure occurs, and the first probability and the second probability determined based on the first log sample are used.
- the calculated amount of log information can accurately reflect the amount of log information when a set fault occurs, thereby improving the accuracy of the trained fault identification model.
- the electronic device performs a summation operation on the total amount of information corresponding to each log printed out at the same time, and obtains the total amount of information corresponding to all the logs printed out at the corresponding time;
- the total amount of information is output, and the first time series data corresponding to the corresponding first period is output.
- the total amount of information corresponding to all logs printed at the same time is the amount of log information output at the corresponding time.
- m represents the number of all logs printed at the corresponding moment
- H(X k ) represents the total information amount of the kth log printed at the corresponding moment.
- two logs of "ERROR” level are printed at the first moment in the first period.
- the first log includes a set phrase "operation timeout”
- the second log includes a set phrase "checkWillUpdateKafka error”.
- the probability of occurrence of "ERROR” level is 1%
- the probability of occurrence of "operation timeout” in ERROR level is 10%
- the probability of occurrence of "checkWillUpdateKafka error” is 5%
- H(x s ) -logP(0.01)-logP(0.1)-logP(0.01)-logP(0.05) ⁇ 6.3
- the total amount of information corresponding to each log is calculated based on the first amount of information and the second amount of information corresponding to each log printed at each moment in the first period of time; based on the first period of time
- the total amount of information corresponding to each log printed out at each moment in the log is obtained, and the first time series data corresponding to the corresponding first period is obtained. Since the total amount of information corresponding to each log can be accurately calculated, the accuracy of the first time series data is improved, thereby improving the accuracy of the fault identification model obtained by training.
- FIG. 3 shows a schematic diagram of an implementation flow of determining positive samples and negative samples in a method for training a fault identification model provided by an embodiment of the present application.
- the construction of at least one positive sample and at least one negative sample corresponding to the set fault based on time series data corresponding to at least one first period of time includes:
- S301 Determine a first indicator related to the set fault based on time series data corresponding to a second period in the first period; wherein the second period indicates that the set fault occurs in the corresponding first period period of time.
- the first indicator includes at least one of the following:
- the amount of log information is the amount of log information.
- the electronic device can determine the corresponding alarm time in the first time period based on the alarm time corresponding to the performance index that triggers the alarm when the set fault occurs. second period. Wherein, the start time of the second time period is less than or equal to the first alarm time when the set fault occurs; the end time of the second time period is greater than or equal to the last alarm time when the set fault occurs.
- the first alarm time and the time period corresponding to 15 minutes before and after the alarm time are determined as the second time period in the corresponding first time period.
- S302 Construct at least one positive sample corresponding to the set fault based on the time series data corresponding to the first indicator in the second time period.
- a positive sample may be constructed based on time series data corresponding to the first indicator in a second time period.
- FIG. 4 shows a schematic diagram of an implementation flow of determining a positive sample in a method for training a fault identification model provided by an embodiment of the present application.
- the construction of at least one positive sample corresponding to the set fault based on the time series data corresponding to the first indicator in the second time period includes:
- S401 Determine a second number of positive samples based on the first number of the first index; wherein the first number is an integer greater than or equal to 2, and the second number is a positive integer; the second number The number is less than or equal to half of the result of the full permutation operation corresponding to the first number.
- M represents the number of positive samples
- N represents the number of first indicators.
- M is a positive integer
- N is an integer greater than or equal to 2.
- the electronic device arranges the time series data corresponding to each first indicator in the second time period in chronological order to obtain a corresponding sequence.
- the sequence corresponding to each first index is fully arranged with the first index as the minimum unit to obtain a full arrangement result.
- a positive sample corresponds to a sequence of permutations. It should be noted that each positive sample may also be a one-dimensional matrix, and a one-dimensional matrix corresponds to a sequence of one arrangement.
- the result of the full arrangement of the first indicators is 6, and a maximum of 3 positive samples can be obtained based on the three first indicators.
- S303 Construct at least one negative sample corresponding to the set fault based on the time series data corresponding to the first indicator in the third period in the first period; wherein the third period represents the corresponding first period period during which the set failure did not occur.
- the third time period is any time period in the corresponding first time period except the second time period.
- a negative sample may be constructed based on the time series data corresponding to the first indicator in a third time period.
- At least one positive sample and at least one negative sample are constructed by using the time series data corresponding to the first indicator related to the fault setting in the first period, so that the difference between the constructed positive sample and the fault setting can be improved. Correlation, thereby improving the performance of the fault identification model trained based on positive samples and negative samples, thereby improving the accuracy of the trained fault identification model.
- the accuracy of the identification result can be improved.
- FIG. 5 shows a schematic diagram of an implementation flow of determining a first indicator in a method for training a fault identification model provided by an embodiment of the present application.
- the first indicator related to the set fault is determined based on the time series data corresponding to the second time period in the first time period, including:
- S501 Determine at least two sets based on the time series data corresponding to the second time period.
- the at least two sets include a first set and a second set corresponding to each performance index in the at least one performance index; the elements in the first set represent the difference between the amount of log information output at two adjacent moments The elements in the second set represent the difference between the monitoring data of the corresponding performance index at two adjacent moments.
- the first set is determined based on the first time series data corresponding to the second time period; the corresponding second set is determined based on the second time series data corresponding to each performance index in the at least one performance index in the second time period. All elements in each of the at least two sets are arranged in chronological order.
- S502 Calculate the mean and standard deviation corresponding to each of the at least two sets.
- the electronic device may calculate the mean value corresponding to each of the at least two sets based on the mean value calculation formula; and calculate the standard deviation corresponding to each set in the at least two sets based on the standard deviation calculation formula.
- S503 Determine a first interval corresponding to each of the at least two sets based on the Three Sigma criterion and the calculated mean and standard deviation.
- the three-sigma criterion states that the probability that the data in a set of data conforming to the normal distribution is distributed in ( ⁇ -3 ⁇ , ⁇ +3 ⁇ ) is 99.73%, and the probability of being distributed outside ( ⁇ -3 ⁇ , ⁇ +3 ⁇ ) is 0.27%.
- ( ⁇ -3 ⁇ , ⁇ +3 ⁇ ) is determined as the corresponding first interval.
- ⁇ represents the mean of the corresponding set;
- ⁇ represents the standard deviation of the corresponding set.
- S504 Determine a first index related to the set fault based on the first interval corresponding to each of the at least two sets and the maximum value in the corresponding set; wherein, the set corresponding to the first index The maximum value in is not in the corresponding first interval.
- the maximum value in the first set is not in the first interval corresponding to the first set, the amount of log information is determined as the first indicator related to the set failure.
- the maximum value in the first set is in the first interval corresponding to the first set, it indicates that the amount of log information is not the first indicator related to the set fault.
- the performance index corresponding to the second set is determined as the first index related to the set fault.
- the maximum value in the second set is in the corresponding first interval, it indicates that the performance index corresponding to the second set is not the first index related to the set fault.
- each first indicator is determined based on time series data corresponding to the second period in a first period.
- the corresponding first interval is determined based on the three-sigma criterion, and the first indicator related to the set fault is selected based on the determined first interval, which can improve the accuracy of the determined first indicator .
- the first indicator related to the set fault is determined based on all the time series data corresponding to the second time period and the three-sigma criterion.
- the time series data corresponding to the second time period does not include second time series data corresponding to the first performance indicator that triggers an alarm for the first time within the second time period;
- the first interval corresponding to each set and the maximum value in the corresponding set determine the first indicator related to the set fault, including:
- a second index is determined based on the first interval corresponding to each of the at least two sets and the maximum value in the corresponding set; wherein, the maximum value in the set corresponding to the second index is not in the corresponding first index interval;
- the determined second index and the first performance index are determined as the first index related to the set failure.
- the electronic device determines, based on the second time series data corresponding to the second time period, the first performance index that triggers the alarm first in the corresponding second time period; and determines the first set based on the first time series data corresponding to the second time period; Based on the time series data corresponding to the second performance index in the second time period, a second set corresponding to the corresponding second performance index is determined.
- the second performance index represents any performance index other than the first performance index among the at least one performance index.
- the amount of log information is determined as the second indicator related to the set fault.
- the maximum value in the first set is in the first interval corresponding to the first set, it indicates that the amount of log information is not the second indicator related to the set fault.
- the performance index corresponding to the second set is determined as the second index related to the set fault.
- the maximum value in the second set is in the corresponding first interval, it indicates that the performance index corresponding to the second set is not the second index related to the set fault.
- the electronic device determines the first performance indicator and all the determined second indicators as the first indicators related to the set failure.
- the first performance index that triggers an alarm for the first time in the corresponding second time period is identified as the first index related to the set fault, and then the set fault is determined according to the above method.
- the second index is to identify the first performance index and the determined second index as the first index related to the set fault. It is not necessary to determine whether the first performance index is related to the set fault through the second time series data corresponding to the first performance index, which can save time for determining the first index and improve the training efficiency of the fault identification model.
- FIG. 6 shows a schematic diagram of an implementation flow of a method for training a fault identification model provided by the application embodiment of the present application.
- the method for training a fault identification model includes:
- S601 Sampling the historical log according to a set time interval to obtain at least one first log sample.
- the set time interval can be 1 second.
- each first log sample in the first log sample satisfies the following conditions:
- the first probability corresponding to the log level corresponding to the first log sample satisfies the set condition
- the alarm type corresponding to the first log sample is the alarm type monitored when the set fault occurs;
- the difference between the third probability and the fourth probability is less than or equal to the set threshold; wherein the third probability represents the probability that the alarm type corresponding to the first log sample appears in the at least one first log sample; the The fourth probability represents the probability that the alarm type corresponding to the first log sample is monitored when the set fault occurs.
- S602 Determine at least one first probability and at least one second probability based on the at least one first log sample; wherein, the first probability is a probability corresponding to each log level; the second probability is that each set phrase is in The corresponding second probability in the corresponding log level.
- S603 Based on the determined first probability and the second probability, determine the time series data corresponding to the historical log.
- the electronic device determines the total amount of information of each log in the historical logs printed at each moment; The amount of log information output at each time, so as to obtain the time series data corresponding to the historical log.
- the time series data corresponding to the history log please refer to the relevant descriptions in the foregoing embodiments S201 to S202, and details are not repeated here.
- S604 Based on the time series data corresponding to each first period in the at least one first period, determine a first indicator related to the set fault.
- the time series data corresponding to each first period includes first time series data and second time series data.
- the first time series data represents the amount of log information output at each moment of the corresponding first time period;
- the second time series data represents the monitoring data of each performance indicator in the at least one performance index at each moment of the corresponding first time period.
- the second time series data is obtained based on the monitoring data corresponding to each performance index of the at least one performance index monitored by the monitoring program at each moment.
- S605 Based on the time series data corresponding to the first indicator in the second time period in the first time period, construct at least one positive sample corresponding to the set fault.
- S606 Based on the time series data corresponding to the first indicator in the third time period in the first time period, construct at least one negative sample corresponding to the set fault.
- S607 Based on at least one positive sample and at least one negative sample corresponding to the set fault, train the set two-class model to obtain a fault identification model corresponding to the set fault.
- At least one first probability and at least one second probability are determined based on all the first log samples; based on the determined first probability and second probability, each of the at least one first time period is determined The first time series data corresponding to the first time period; based on the first time series data and the second time series data corresponding to the first time period, at least one positive sample and at least one negative sample corresponding to the set fault are determined; based on at least one positive sample and at least one negative sample A negative sample training is used to obtain the fault identification model corresponding to the set fault.
- the log information amount calculated based on the probability determined based on the first log sample can accurately reflect the log information when the set failure occurs and improve the accuracy of the fault identification model obtained by training.
- the first time series data represents the amount of log information output at each moment of the first period, and the amount of log information refers to the amount of information in the log, which is used to measure the amount of information conveyed by the log; it is not necessary to analyze the historical log when determining the amount of log information Therefore, the time spent on semantic analysis of the text content of the log can be saved, the efficiency of obtaining training samples is improved, and the training efficiency of the fault identification model can be improved.
- the electronic device can train fault identification models corresponding to different set faults based on the above-mentioned embodiment.
- FIG. 7 shows a schematic flowchart of an implementation of a fault identification method provided by an embodiment of the present application.
- the execution subject of the fault identification method is an electronic device, such as a computer, a server, and the like.
- the electronic device can monitor data related to performance indicators of the DBaaS server, and obtain real-time logs of the DBaaS server.
- the electronic equipment for fault identification based on the fault identification model and the electronic equipment for training the fault identification model may be the same or different.
- the fault identification method includes:
- the end time of the first time period is the alarm time corresponding to the first performance index;
- the time series data corresponding to the first time period includes the first time series data corresponding to the real-time log and each performance index in the at least one performance index corresponding second time series data.
- the electronic device determines the first time series data based on the real-time log output by the DBaaS server at each moment; and determines the first time series data corresponding to each performance indicator based on the real-time monitoring data corresponding to each performance indicator in the at least one performance indicator at each moment. 2. Time series data.
- the electronic device determines, based on the current alarm time of the first performance indicator, time series data corresponding to the first period from the determined time series data.
- the first performance index is any one of the at least one performance index.
- the time series data corresponding to the first time period includes the first time series data and the second time series data corresponding to each performance index in the at least one performance index.
- the first time period may be the first 15 minutes of the alarm time of the first performance indicator.
- the electronic device detects the data of the first performance index at 8:30 in the morning to trigger an alarm, and the first period is the period from 8:15 to 8:30.
- S702 Determine at least one first set fault based on the relevant performance index corresponding to each set fault in the at least one set fault; the relevant performance index of the first set fault includes the first performance index.
- S703 Input the time series data corresponding to the first time period into the fault identification model corresponding to each first set fault in the at least one first set fault, and obtain the identification result output by each fault identification model.
- the fault identification model is obtained by training the method for training the fault identification model corresponding to any of the foregoing embodiments.
- the identification result output by each fault identification model is a numerical value between 0 and 1.
- the identification result represents the probability that the fault that occurs is the first set fault corresponding to the fault identification model.
- S704 Determine the first set fault with the highest confidence in all the identification results as the fault that currently occurs.
- Confidence Indicates the probability that the corresponding first set fault is credible. For example, if the confidence level of a first set fault is 80%, it means that the probability that the fault that occurs is the first set fault is 80%. .
- the time series data corresponding to the first time period is determined, and the time series data corresponding to the first time period is input into at least one first set fault
- the fault identification model corresponding to each of the first set faults is obtained, the identification result output by each fault identification model is obtained, and the first set fault with the highest confidence in all the identification results is determined as the current fault.
- the first time series data represents the output of each time in the first time period determined based on the alarm time
- FIG. 8 shows a schematic flowchart of the implementation of a fault identification method provided by another embodiment of the present application.
- the fault identification method further includes:
- S705 Determine the health scores of all the servers in the at least one server based on the historical failures of each server in the at least one server and the number of occurrences corresponding to each historical failure, and determine the health scores of all the servers in the at least one server. Deploy a new database instance in .
- the electronic device inquires the health score of the corresponding server from the set score table based on the identification of the historical failure of the server and the number of times corresponding to each historical failure. in,
- a set score table is pre-stored in the electronic device, and the set score table includes scores corresponding to all set faults corresponding to each fault level in at least one fault level in different frequency ranges. That is to say, in the set score table, a corresponding score range is configured for each fault level, and a score corresponding to a different frequency range is configured for the set fault corresponding to each fault level, and the score is in the corresponding The score range corresponding to the fault level.
- all set faults can be sorted according to the priority of the fault level.
- the higher the priority of the fault level the greater the impact of the set fault corresponding to the fault level on the server. .
- the higher the priority of the fault level the larger the corresponding score range. The more times the set fault corresponding to the same fault level occurs, the lower the corresponding score.
- the score range corresponding to the first fault level is smaller than the score range of the second fault level.
- the score corresponding to the first frequency range is smaller than the score corresponding to the second frequency range.
- the electronic device may determine the health score of the corresponding server based on the historical faults that occur on the server and the number of times corresponding to each historical fault. In this way, it is possible to avoid interfering with the process executed by the server due to the health score of the server, and reduce the data processing efficiency of the server.
- the health score of the corresponding server is determined based on the historical failures of the server and the number of occurrences corresponding to each historical failure, so as to deploy a new database instance in the server whose health score is greater than or equal to the set threshold .
- the health score of the server is greater than or equal to the set threshold, it indicates that the corresponding server has less potential security risks. Deploying a new database instance on a server with a health score greater than or equal to the set threshold can reduce the need for deploying new database instances. The probability of the corresponding server failure in the process can also reduce the workload of subsequent operation and maintenance.
- the embodiment of the present application also provides a device for training a fault identification model, which is arranged on an electronic device.
- the device for training a fault identification model includes: :
- the construction unit 91 is configured to construct at least one positive sample and at least one negative sample corresponding to the set fault based on the time series data corresponding to each first period in the at least one first period;
- the training unit 92 is configured to train the set two-class model based on at least one positive sample and at least one negative sample corresponding to the set fault to obtain a fault identification model corresponding to the set fault;
- the time series data includes first time series data and second time series data
- the first time series data represents the amount of log information output at each moment of the corresponding first period
- the second time series data represents monitoring data of each performance indicator of the at least one performance indicator at each moment of the corresponding first time period.
- the apparatus for training the fault identification model further includes:
- a computing unit configured to calculate the total amount of information corresponding to each log based on the first amount of information and the second amount of information corresponding to each log printed at each moment in the first period;
- the determining unit is configured to obtain the first time series data corresponding to the corresponding first period based on the total amount of information corresponding to each log printed at each moment in the first period;
- the first amount of information represents the first probability of occurrence of the log level corresponding to the log
- the second amount of information represents the second probability that each set phrase appears in the corresponding log level among all set phrases included in the log.
- the computing unit is further configured to:
- a second probability corresponding to each set phrase in the corresponding log level is determined;
- the at least one first log sample is obtained by sampling historical logs.
- each first log sample in the at least one first log sample satisfies the following conditions:
- the first probability corresponding to the log level corresponding to the first log sample satisfies the set condition
- the alarm type corresponding to the first log sample is the alarm type monitored when the set fault occurs;
- the difference between the third probability and the fourth probability is less than or equal to the set threshold; wherein the third probability represents the probability that the alarm type corresponding to the first log sample appears in the at least one first log sample; the The fourth probability represents the probability that the alarm type corresponding to the first log sample is monitored when the set fault occurs.
- the building unit 91 is configured to:
- a first indicator related to the set fault is determined; wherein the second time period represents the time period in which the set fault occurs in the corresponding first time period ;
- At least one negative sample corresponding to the set fault is constructed based on the time series data corresponding to the first indicator in the third period in the first period;
- the third period represents a period in which the set failure does not occur in the corresponding first period
- the first indicator includes at least one of the following:
- the amount of log information is the amount of log information.
- the building unit 91 is configured to:
- At least two sets are determined based on the time series data corresponding to the second time period; wherein, the at least two sets include a first set and a second set corresponding to each performance index in the at least one performance index; the first set The elements in represent the difference between the amount of log information output at two adjacent moments; the elements in the second set represent the difference between the monitoring data of the corresponding performance index at two adjacent moments;
- a first index related to the set fault is determined; wherein, the first index in the set corresponding to the first index The maximum value is not in the corresponding first interval.
- the time series data corresponding to the second time period does not include the second time series data corresponding to the first performance index that triggers the alarm first in the second time period; the construction unit 91 is configured to:
- a second index is determined based on the first interval corresponding to each of the at least two sets and the maximum value in the corresponding set; wherein, the maximum value in the set corresponding to the second index is not in the corresponding first index interval; determining the determined second index and the first performance index as the first index related to the set fault.
- the building unit 91 is configured to:
- a second number of positive samples are selected from the full permutation results; wherein,
- the first number is an integer greater than or equal to 2, and the second number is a positive integer; the second number is less than or equal to half of the result of the full permutation operation corresponding to the first number.
- each unit included in the apparatus for training a fault identification model may be implemented by a processor in the apparatus for training a fault identification model.
- the processor needs to run the program stored in the memory to realize the functions of the above program modules.
- the above processing may be allocated to different The program module is completed, that is, the internal structure of the apparatus for training the fault identification model is divided into different program modules, so as to complete all or part of the above-described processing.
- the apparatus for training a fault identification model provided in the above embodiment and the method for training a fault identification model belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
- the embodiment of the present application further provides a fault identification device, which is arranged on the electronic device.
- the fault identification device includes:
- the first determining unit 101 is configured to determine the time series data corresponding to the first time period in the case that the data of the first performance index is detected to trigger an alarm;
- the second determination unit 102 is configured to determine at least one first set fault based on the relevant performance index corresponding to each set fault in the at least one set fault; among the relevant performance indicators of the first set fault including the first performance index;
- the first identification unit 103 is configured to input the time series data corresponding to the first time period into the fault identification model corresponding to each first set fault in the at least one first set fault, and obtain the output of each fault identification model identification result;
- the second identification unit 104 is configured to determine the first set fault with the highest confidence in all the identification results as the currently occurring fault; wherein,
- the end moment of the first time period is the alarm moment corresponding to the first performance indicator
- the time series data corresponding to the first time period includes the first time series data corresponding to the real-time log and the second time series data corresponding to each performance index in the at least one performance index;
- the fault identification model is obtained by training based on any of the above methods for training a fault identification model.
- the fault identification device further includes:
- the scoring unit is configured to determine the health score of all the servers in the at least one server based on the historical failures of each server in the at least one server and the number of occurrences corresponding to each historical failure, so that the health score is greater than or equal to the set value. Deploy a new database instance in the threshold server.
- each unit included in the fault identification device can be implemented by a processor in the fault identification device.
- the processor needs to run the program stored in the memory to realize the functions of the above program modules.
- the device for fault identification model provided by the above embodiment trains the fault identification model
- only the division of the above program modules is used as an example for illustration.
- the above processing can be allocated to different programs as required.
- the module is completed, that is, the internal structure of the fault identification device is divided into different program modules, so as to complete all or part of the processing described above.
- the fault identification device and the fault identification method embodiments provided by the above embodiments belong to the same concept, and the specific implementation process thereof is detailed in the method embodiments, which will not be repeated here.
- FIG. 11 is a schematic diagram of a hardware structure of an electronic device provided by an embodiment of the application. As shown in FIG. 11 , the electronic device includes:
- Communication interface 1 which can exchange information with other devices such as servers;
- the processor 2 is connected with the communication interface 1 to realize information exchange with other devices, and when configured to run a computer program, execute the method for training a fault identification model provided by the above one or more technical solutions; or execute the above one or more The fault identification method provided by the technical solution.
- the computer program is instead stored on the memory 3 .
- bus system 4 is used to realize the connection communication between these components.
- the bus system 4 also includes a power bus, a control bus and a status signal bus.
- the various buses are designated as bus system 4 in FIG. 11 .
- the memory 3 in the embodiment of the present application is configured to store various types of data to support the operation of the electronic device.
- Examples of such data include: any computer program configured to operate on an electronic device.
- the memory 3 may be a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memory.
- the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-only memory) Only Memory), Electrically Erasable Programmable Read-Only Memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), Magnetic Random Access Memory (FRAM, ferromagnetic random access memory), Flash Memory (Flash Memory), Magnetic Surface Memory , CD-ROM, or CD-ROM (Compact Disc Read-Only Memory); magnetic surface memory can be disk memory or tape memory.
- RAM Random Access Memory
- SRAM Static Random Access Memory
- SSRAM Synchronous Static Random Access Memory
- DRAM Dynamic Random Access Memory
- SDRAM Synchronous Dynamic Random Access Memory
- DDRSDRAM Double Data Rate Synchronous Dynamic Random Access Memory
- ESDRAM Enhanced Type Synchronous Dynamic Random Access Memory
- SLDRAM Synchronous Link Dynamic Random Access Memory
- DDRRAM Direct Memory Bus Random Access Memory
- DRRAM Direct Rambus Random Access Memory
- the methods disclosed in the above embodiments of the present application may be applied to the processor 2 or implemented by the processor 2 .
- the processor 2 may be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above-mentioned method can be completed by a hardware integrated logic circuit in the processor 2 or an instruction in the form of software.
- the above-mentioned processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
- the processor 2 may implement or execute the methods, steps, and logical block diagrams disclosed in the embodiments of this application.
- a general purpose processor may be a microprocessor or any conventional processor or the like.
- the steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor.
- the software module may be located in a storage medium, the storage medium is located in the memory 3, and the processor 2 reads the program in the memory 3, and completes the steps of the foregoing method in combination with its hardware.
- the embodiment of the present application further provides a storage medium, that is, a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 3 storing a computer program, and the above-mentioned computer program can be executed by the processor 2,
- a storage medium that is, a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 3 storing a computer program, and the above-mentioned computer program can be executed by the processor 2,
- the computer-readable storage medium may be memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disk, or CD-ROM.
- the disclosed apparatus and method may be implemented in other manners.
- the device embodiments described above are only illustrative.
- the division of the units is only a logical function division. In actual implementation, there may be other division methods.
- multiple units or components may be combined, or Can be integrated into another system, or some features can be ignored, or not implemented.
- the coupling, or direct coupling, or communication connection between the components shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical or other forms. of.
- the unit described above as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units; Some or all of the units may be selected according to actual needs to achieve the purpose of the solution in this embodiment.
- each functional unit in each embodiment of the present application may all be integrated into one processing module, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integration
- the unit can be implemented either in the form of hardware or in the form of hardware plus software functional units.
- the aforementioned program may be stored in a computer-readable storage medium, and when the program is executed, execute Including the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and other various A medium on which program code can be stored.
- ROM read-only memory
- RAM random access memory
- magnetic disk or an optical disk and other various A medium on which program code can be stored.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computer Hardware Design (AREA)
- Probability & Statistics with Applications (AREA)
- Quality & Reliability (AREA)
- Databases & Information Systems (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Debugging And Monitoring (AREA)
Abstract
本申请实施例公开了一种故障识别模型训练方法、故障识别方法、装置及电子设备。该故障识别模型训练方法包括:基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,所述时序数据包括第一时序数据和第二时序数据;第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
Description
相关申请的交叉引用
本申请基于申请号为202011164795.1,申请日为2020年10月27日的中国专利申请提出,并要求上述中国专利申请的优先权,上述中国专利申请的全部内容在此引入本申请作为参考。
本申请涉及计算机技术领域,涉及一种故障识别模型训练方法、故障识别方法、装置及电子设备。
随着计算机技术的发展,越来越多的技术应用在金融领域,传统金融业正在逐步向金融科技(Fintech)转变,然而,由于金融行业的安全性、实时性要求,金融科技也对技术提出了更高的要求。金融科技领域下,数据库即服务(DBaaS,Database as a Service)是以传统数据库技术为基础,将数据库资源以标准服务的形式提供给一个或多个租户的服务解决方案。
相关技术中,在用于提供DbaaS的服务器发生故障时,通常基于设定的性能指标的监测数据和DbaaS的相关日志对应的语义分析结果进行故障分析,得到分析结果。这种故障分析方法需要消耗大量时间分析所有日志文本的具体内容,以得到对应的语义分析结果,导致故障分析的效率较低。
发明内容
为解决相关技术问题,本申请实施例期望提供一种故障识别模型训练方法、故障识别方法、装置及电子设备。
本申请实施例提供一种训练故障识别模型的方法,包括:
基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;
基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,
所述时序数据包括第一时序数据和第二时序数据;
第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;
第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
本申请实施例还提供了一种故障识别方法,包括:
在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据;
基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标;
将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果;
将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障;其中,
所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;
所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据;
所述故障识别模型基于上述任一种训练故障识别模型的方法训练得到。
上述方案中,所述故障识别方法还包括:
基于至少一台服务器中每台服务器发生的历史故障和每种历史故障对应的发生次数,确定出至少一台服务器中所有服务器的健康评分,以在健康评分大于或等于设定阈值的服务器中部署新的数据库实例。
本申请实施例还提供了一种训练故障识别模型的装置,包括:
构建单元,用于基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;
训练单元,用于基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,
所述时序数据包括第一时序数据和第二时序数据;
第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;
第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
本申请实施例还提供了一种故障识别装置,包括:
第一确定单元,配置为在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据;
第二确定单元,配置为基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标;
第一识别单元,配置为将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果;
第二识别单元,配置为将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障;其中,
所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;
所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据;
所述故障识别模型基于任一种训练故障识别模型的方法训练得到。
本申请实施例还提供了一种电子设备,包括:处理器和配置为存储能够在处理器上运行的计算机程序的存储器,
其中,所述处理器配置为运行所述计算机程序时,执行以下至少一项:
上述任一种训练故障识别模型的方法的步骤;
上述任一种故障识别方法的步骤。
本申请实施例还提供了一种存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时以下至少一项:
上述任一种训练故障识别模型的方法的步骤;
上述任一种故障识别方法的步骤。
本申请实施例,基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型。由于训练样本是基于第一时段对应的时序数据构建得到,第一时段对应的时序数据包括的第一时序数据表征第一时段的每个时刻输出的日志信息量,而日志信息量是指日志的信息量,用于度量日志传达的信息多少;在确定日志信息量时不需要分析历史日志的具体内容,也就是说,在构建训练样本时服务器不需要关注历史日志的具体内容,因此可以省去对日志的文本内容进行语义分析消耗的时间,提高获取训练样本的效率,进而提高故障识别模型的训练效率。
在利用故障识别模型进行故障识别时,由于输入故障识别模型的时序数据中包括实时日志对应的第一时序数据,第一时序数据表征每个时刻输出的日志信息量,在确定日志信息量时不需要对实时日志的文本内容进行语义分析,因此,可以节省时间,可以提高故障识别效率。
图1为本申请实施例提供的一种训练故障识别模型的方法的实现流程示意图;
图2为本申请实施例提供的一种训练故障识别模型的方法中确定第一时序数据的实现流程示意图;
图3为本申请实施例提供的一种训练故障识别模型的方法中确定正样本和负样本的实现流程示意图;
图4为本申请实施例提供的一种训练故障识别模型的方法中确定正样本的实现流程示意图;
图5为本申请实施例提供的一种训练故障识别模型的方法中确定第一指标的实现流程示意图;
图6为本申请应用实施例提供的一种训练故障识别模型的方法的实现流程示意图;
图7为本申请实施例提供的一种故障识别方法的实现流程示意图;
图8为本申请另一实施例提供的一种故障识别方法的实现流程示意图;
图9为本申请实施例提供的训练故障识别模型的装置的结构示意图;
图10为本申请实施例提供的故障识别装置的结构示意图;
图11为本申请实施例提供的电子设备的硬件组成结构示意图。
以下结合说明书附图及具体实施例对本申请的技术方案做进一步的详细阐述。
针对上述技术问题,本申请实施例提供了一种故障识别方法,基于日志信息量以及性能指标的监测数据进行故障识别。由于在确定日志信息量时不需要分析日志的具体内容,可以节省对日志的文本内容进行语义分析消耗的时间,提高故障识别效率。
以下结合说明书附图及具体实施例对本申请的技术方案做进一步的详细阐述。
图1示出了本申请实施例提供的一种训练故障识别模型的方法的实现流程示意图。 在本申请实施例中,训练故障识别模型的方法的执行主体为电子设备,例如,电脑、服务器等。
参照图1,本申请实施例提供的训练故障识别模型的方法包括:
S101:基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;其中,所述时序数据包括第一时序数据和第二时序数据;第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
这里,第一时段包括发生设定故障的时段和未发生该设定故障的时段。
电子设备基于第一时段中发生设定故障的时段对应的第一时序数据和第二时序数据构建正样本;基于第一时段中未发生设定故障的时段对应的第一时序数据和第二时序数据构建负样本。
需要说明的是,发生设定故障时,至少一个性能指标的数据触发告警。电子设备可以基于在发生该设定故障时首个触发告警的性能指标对应的告警时刻,确定出发生该设定故障的时段以及确定出未发生该设定故障的时段。发生该设定故障的时段包括首个告警时刻;未发生该设定故障的时段为对应的第一时段中除发生该设定故障的时段之外的时段。
例如,设定故障对应的至少一个性能指标中的第一性能指标在第一时段中的第一时刻触发告警,电子设备可以将第一时刻之前的第二时刻确定为发生设定故障的时段的开始时刻,将第一时刻之后的第三时刻确定为发生设定故障的时段的结束时刻。需要说明的是,第二时刻在对应的第一时段的开始时刻之后,第三时刻在对应的第一时段的结束时刻之前。第一时长和第二时长可以相同,也可以不同。第一时长为第一时刻与第二时刻之间的时长,第二时长为第一时刻与第三时刻之间的时长。示例性地,第一时长和第二时长可以为15分钟,当然,也可以根据实际情况进行设置。
需要说明的是,第一时序数据是按时间先后顺序记录的日志信息量。第二时序数据是按时间先后顺序记录的至少一个性能指标中每个性能指标的监测数据。每个性能指标对应一组第二时序数据。其中,第一时序数据基于DBaaS服务器的历史日志得到。第二时序数据基于DBaaS服务器的性能指标的监测数据得到。至少一个性能指标用于监测是否发生设定故障。设定故障可以为连接数饱和、机器磁盘故障、内存故障、慢查询过多等。
日志信息量是指日志的信息量。日志信息量用于量度日志的信息多少。每个时刻输出的日志信息量是该时刻打印出的所有日志中每条日志的信息量的总和。一条日志的信息量基于第一概率和至少一个第二概率确定出。第一概率表征该条日志对应的日志级别出现的概率。第二概率表征该条日志中包括的所有设定词组中每个设定词组在该日志对应的日志级别下出现的概率。
示例性地,日志级别可以包括FATAL、ERROR、WARN、INFO、DEBUG。其中,
FATAL级别的日志表征每个严重的错误事件将会导致应用程序退出。
ERROR级别的日志表征虽然发生错误事件,但仍然不影响系统继续运行。
WARN级别的日志表征系统存在潜在的错误事件。
INFO级别的日志表征应用程序的运行情况。
DEBUG级别的日志主要用于在调试应用程序时更详细的了解系统的运行状态。
S102:基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型。
电子设备将至少一个正样本中的每个正样本转换成对应的第一向量,将至少一个负样本中的每个负样本转换成对应的第二向量;将至少一个第一向量和至少一个第二向量 输入至设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型。
其中,在训练设定的二分类模型的过程中,在二分类模型未达到设定的收敛条件时,更新二分类模型的相关参数,基于至少一个正样本和至少一个负样本继续训练该二分类模型。设定的收敛条件可以是第一模型参数与第二模型参数之间的差值小于或等于设定阈值,当然也可以设置其他收敛条件,例如,训练次数达到设定次数。第一模型参数表征第k次迭代训练对应的模型参数,第二模型参数表征第k-1次迭代训练对应的模型参数,k为大于或等于1的整数。
在二分类模型达到设定的收敛条件时,,停止训练,将最后一次更新得到的模型参数确定为训练完毕的二分类模型所使用的模型参数,并将训练完毕的二分类模型确定为设定故障对应的故障识别模型。
需要说明的是,由于设定故障对应的正样本和负样本的数量通常比较少,为了防止过拟合,设定的二分类模型通常为简单的卷积神经网络,例如,由一层卷积层和隐藏层构成的卷积神经网络。设定的二分类模型也可以为决策树模型。
在本实施例提供的方案中,基于至少一个第一时段中每个第一时段对应的第一时序数据和第二时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型。由于第一时序数据表征对应的第一时段的每个时刻输出的日志信息量,而日志信息量是指日志的信息量,用于度量日志传达的信息多少;在确定日志信息量时不需要分析历史日志的具体内容,因此可以省去对日志的文本内容进行语义分析消耗的时间,提高获取训练样本的效率,进而提高故障识别模型的训练效率。
在一实施例中,图2示出了本申请实施例提供的一种训练故障识别模型的方法中确定第一时序数据的实现流程示意图。参照图2,确定第一时序数据的方法包括:
S201:基于第一时段中每个时刻打印出的每条日志对应的第一信息量和第二信息量,计算出每条日志对应的总信息量;其中,所述第一信息量表征日志对应的日志级别出现的第一概率;所述第二信息量表征日志中包括的所有设定词组中每个设定词组在对应的日志级别中出现的第二概率。
电子设备计算每条日志对应的第一信息量和第二信息量之和,得到对应日志的总信息量。
这里,可以基于以下公式计算每条日志对应的总信息量:
H(X)=H(x
t)+H(x
c) (1)
H(x
t)=-logP(x
l) (2)
其中,H(X)表征一条日志的总信息量,H(x
t)表征日志的第一信息量,H(x
c)表征日志的第二信息量;P(x
l)表征日志对应的日志级别x
l出现的第一概率;P(x
i|x
l)表征在日志对应的日志级别x
l中第i个设定词组出现的第二概率,n表征对应日志中包含的设定词组的数量。
在一实施例中,按照以下方式确定出第一概率和第二概率:
基于至少一个第一日志样本中每个日志级别出现的次数,确定出每个日志级别对应的第一概率;
基于所述至少一个第一日志样本中每个设定词组在每个日志级别下出现的次数,确定出每个设定词组在对应的日志级别中对应的第二概率;其中,
所述至少一个第一日志样本通过对历史日志进行采样得到。
这里,电子设备对采集到的历史日志进行采样,得到至少一个第一日志样本。采集到的历史日志可以包括至少一个第一时段中每个第一时段输出的历史日志中的部分或全部,采集到的历史日志也可以不包括第一时段输出的历史日志。
在一实施例中,至少一个第一日志样本中每个第一日志样本均满足以下条件:
第一日志样本对应的日志级别对应的第一概率满足设定条件;
第一日志样本对应的告警类型为发生所述设定故障时监测到的告警类型;
第三概率与第四概率之间的差值小于或等于设定阈值;其中,所述第三概率表征第一日志样本对应的告警类型在所述至少一个第一日志样本中出现的概率;所述第四概率表征第一日志样本对应的告警类型在发生所述设定故障时被监测到的概率。
其中,设定条件表征为所述设定故障配置的所有日志级别中每种日志级别对应的概率范围。由于数据库中的设定组件的源码决定了在发生设定故障时输出哪种日志级别的日志,以及决定了每种日志级别出现的概率,因此,电子设备可以分析数据库中的设定组件的源码,确定出在设定故障时输出的日志对应的每种日志级别对应的概率范围,基于确定出的概率范围设置上述设定条件。
第一日志样本对应的所有告警类型由第一日志样本中ERROR级别的日志中的设定词组确定出;一条日志包含至少一个设定词组;一个设定词组对应一个告警类型。
电子设备在得到至少一个第一日志样本的情况下,基于至少一个第一日志样本中的所有第一日志样本确定出第一概率和第二概率,具体如下:
电子设备基于第一日志样本包括的至少一条日志中每条日志对应的日志级别,统计出所有第一日志样本中每种日志级别出现的次数;对所有第一日志样本中每种日志级别出现的次数进行求和运算,得到所有第一日志样本中所有日志级别出现的第一总次数;基于所有第一日志样本中每种日志级别出现的次数,以及基于所有第一日志样本中所有日志级别出现的第一总次数,计算出每种日志级别对应的第一概率。第一概率为对应的日志级别出现的次数与第一总次数的商。
电子设备基于分词技术,对第一日志样本中的每条日志进行分词,得到分词结果,对分词结果进行过滤处理,提取出每条日志对应的设定词组。对分词结果进行过滤处理包括去除连接词、常用词(例如,end、run等)、组件的名称、调用的脚本的名称等。设定词组表征设定故障对应的告警类型(或称告警事件)。例如,设定故障为连接数饱和时,设定词组可以包括:operation timeout、checkWillUpdateKafka error等。
电子设备基于所有第一日志样本中每条日志对应的设定词组,统计出每个设定词组在对应的日志级别中出现的次数;对每个设定词组在对应的日志级别中出现的次数进行求和运算,统计出所有设定词组对应的日志级别中出现的第二总次数;基于每个设定词组在对应的日志级别中出现的次数,以及基于对应的第二总次数,确定出每个设定词组在对应的日志级别中出现的第二概率。其中,一条日志对应一个设定的日志级别,第二概率为每个设定词组在对应的日志级别下出现的次数与对应的第二总次数的商。
在本实施例中,第一日志样本满足上述三个条件,第一日志样本能够准确反映出发生设定故障时服务器的运行情况,利用基于第一日志样本确定出的第一概率和第二概率计算出的日志信息量,能够准确反映出发生设定故障时的日志信息量,进而提高训练得到的故障识别模型的准确度。
S202:基于第一时段中每个时刻打印出的每条日志对应的总信息量,得到对应的第一时段对应的第一时序数据。
电子设备对同一时刻打印出的每条日志对应的总信息量进行求和运算,得到对应时刻打印出的所有日志对应的总信息量;基于第一时段中每个时刻打印出的所有日志对应 的总信息量,输出对应的第一时段对应的第一时序数据。其中,同一时刻打印出的所有日志对应的总信息量,即为对应时刻输出的日志信息量。
每个时刻打印出的所有日志对应的总信息量H(x
s)为:
其中,m表征对应时刻打印出的所有日志的数量,H(X
k)表征对应时刻打印出的第k条日志的总信息量。
这里,将通过上述公式(1)至(3)计算出的每条日志的总信息量,代入公式(4),即可得到对应时刻打印出的所有日志对应的总信息量。
例如,在第一时段中第一时刻打印出了两条“ERROR”级别的日志,第一条日志包括一个设定词组“operation timeout”,第二条日志包括一个设定词组“checkWillUpdateKafka error”,预先基于至少一个第一日志样本确定出“ERROR”级别出现的概率为1%,“operation timeout”在ERROR级别中出现的概率为10%,“checkWillUpdateKafka error”出现的概率为5%,那么基于上述公式(4)可以得到第一时刻输出的信息量的表达式为:
基于上述(1)至(3)可以得到:
H(X
1)=-logP(0.01)-logP(0.1) (6)
H(X
2)=-logP(0.01)-logP(0.05) (7)
将(6)和(7)代入(5),得到第一时刻输出的信息量为:
H(x
s)=-logP(0.01)-logP(0.1)-logP(0.01)-logP(0.05)≈6.3
在本实施例提供的方案中,基于第一时段中每个时刻打印出的每条日志对应的第一信息量和第二信息量,计算出每条日志对应的总信息量;基于第一时段中每个时刻打印出的每条日志对应的总信息量,得到对应的第一时段对应的第一时序数据。由于可以准确计算出每条日志对应的总信息量,因此提高第一时序数据的准确度,进而提高训练得到的故障识别模型的准确度。
作为本申请的另一实施例,图3示出了本申请实施例提供的一种训练故障识别模型的方法中确定正样本和负样本的实现流程示意图。参照图3,所述基于至少一个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本,包括:
S301:基于第一时段中的第二时段对应的时序数据,确定出与所述设定故障相关的第一指标;其中,所述第二时段表征对应的第一时段中发生所述设定故障的时段。
这里,所述第一指标包括以下至少之一:
在发生所述设定故障时告警的所有性能指标中的至少一个性能指标;
日志信息量。
由于在发生设定故障时,存在至少一个性能指标的监测数据触发告警,因此,电子设备可以基于在发生设定故障时触发告警的性能指标对应的告警时刻,确定出对应的第一时段中的第二时段。其中,第二时段的起始时刻小于或等于在发生设定故障时的首个告警时刻;第二时段的结束时刻大于或等于在发生设定故障时的最后一个告警时刻。
例如,将第一个告警时刻以及该告警时刻前后15分钟对应的时段,确定为对应的第一时段中的第二时段。
S302:基于所述第一指标在所述第二时段对应的时序数据,构建所述设定故障对应的至少一个正样本。
这里,可基于第一指标在一个第二时段对应的时序数据,构建一个正样本。
在一实施例中,图4示出了本申请实施例提供的一种训练故障识别模型的方法中确定正样本的实现流程示意图。参照图4,所述基于所述第一指标在所述第二时段对应的时序数据,构建所述设定故障对应的至少一个正样本,包括:
S401:基于所述第一指标的第一数量,确定出正样本的第二数量;其中,所述第一数量为大于或等于2的整数,所述第二数量为正整数;所述第二数量小于或等于所述第一数量对应的全排列运算结果的二分之一。
其中,M表征正样本的数量,N表征第一指标的数量。M为正整数,N大于或等于2的整数。
S402:将所有第一指标在第二时段对应的时序数据以第一指标为最小单位进行全排列,得到全排列结果。
这里,电子设备将每个第一指标在第二时段对应的时序数据,按照时间先后顺序进行排列,得到对应的序列。将每个第一指标对应的序列以第一指标为最小单位进行全排列,得到全排列结果。
需要说明的是,通过变换所有第一指标的时序数据的排列顺序以扩充正样本,可以节省选取正样本消耗的时间。考虑到在发生设定故障时,对应的性能指标的告警顺序并不是固定不变的,因此,变换所有第一指标的时序数据的排列顺序,还可以使得训练得到的故障识别模型对第一指标对应的时序数据的排列顺序不敏感,提升故障识别模型的泛化能力。
S403:从所述全排列结果中选出第二数量的正样本。
这里,一个正样本对应一种排列方式的序列。需要说明的是,每个正样本也可以为一维矩阵,一个一维矩阵对应一种排列方式的序列。
示例性地,第一指标的数量为3时,第一指标的全排列结果为6,基于3个第一指标最多可得到3个正样本。
S303:基于所述第一指标在所述第一时段中的第三时段对应的时序数据,构建所述设定故障对应的至少一个负样本;其中,所述第三时段表征对应的第一时段中未发生所述设定故障的时段。
第三时段为对应的第一时段中除第二时段之外的任一时段。这里,可以基于第一指标在一个第三时段对应的时序数据,构建一个负样本。
本实施例提供的方案中,通过与设定故障相关的第一指标在第一时段对应的时序数据,构建至少一个正样本以及至少一个负样本,可以提高构建出的正样本与设定故障的相关性,进而提高基于正样本和负样本训练得到的故障识别模型的性能,进而提高训练得到的故障识别模型的准确度。在利用该故障识别模型进行故障识别时,可以提高识别结果的准确度。
作为本申请的另一实施例,图5示出了本申请实施例提供的一种训练故障识别模型的方法中确定第一指标的实现流程示意图。参照图5,所述基于第一时段中的第二时段对应的时序数据,确定出与所述设定故障相关的第一指标,包括:
S501:基于第二时段对应的时序数据,确定出至少两个集合。
其中,所述至少两个集合包括一个第一集合和至少一个性能指标中每个性能指标对应的第二集合;所述第一集合中的元素表征相邻两个时刻输出的日志信息量之间的差值;所述第二集合中的元素表征对应的性能指标在相邻两个时刻的监测数据之间的差值。
这里,基于第二时段对应的第一时序数据,确定出第一集合;基于至少一个性能指标中每个性能指标在第二时段对应的第二时序数据,确定出对应的第二集合。至少两个集合中每个集合中的所有元素按照时间先后顺序排列。
S502:计算出所述至少两个集合中每个集合对应的均值和标准差。
电子设备可以基于均值计算公式,计算出至少两个集合中每个集合对应的均值;基于标准差的计算公式,计算出至少两个集合中每个集合对应标准差。
S503:基于三西格玛准则以及基于计算出的均值和标准差,确定出所述至少两个集合中每个集合对应的第一区间。
三西格玛准则指出,符合正态分布的一组数据中的数据分布在(μ-3σ,μ+3σ)中的概率为99.73%,分布在(μ-3σ,μ+3σ)之外的概率为0.27%。
因此,将(μ-3σ,μ+3σ)确定为对应的第一区间。μ表征对应集合的均值;σ表征对应集合的标准差。
S504:基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出与所述设定故障相关的第一指标;其中,所述第一指标对应的集合中的最大值未处于对应的第一区间。
这里,当第一集合中的最大值未处于第一集合对应的第一区间时,将日志信息量确定为与设定故障相关的第一指标。当第一集合中的最大值处于第一集合对应的第一区间时,表征日志信息量不是与设定故障相关的第一指标。
当第二集合中的最大值未处于对应的第一区间时,将该第二集合对应的性能指标确定为设定故障相关的第一指标。当第二集合中的最大值处于对应的第一区间时,表征该第二集合对应的性能指标不是与设定故障相关的第一指标。
需要说明的是,当第一时段的数量为至少两个时,可以按照上述方式确定出至少两份第一指标,将至少两份第一指标中的交集确定为最终的第一指标,这样可以提高第一指标的准确度。其中,每份第一指标基于一个第一时段中的第二时段对应的时序数据确定出。
本实施例提供的方案中,基于三西格玛准则确定出对应的第一区间,基于确定出的第一区间筛选出与设定故障相关的第一指标,可以提高确定出的第一指标的准确度。
需要说明的是,图5对应的实施例中是基于第二时段对应的所有时序数据以及三西格玛准则确定出与设定故障相关的第一指标。在另一实施例中,所述第二时段对应的时序数据未包括在第二时段内首个触发告警的第一性能指标对应的第二时序数据;所述基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出与所述设定故障相关的第一指标,包括:
基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出第二指标;其中,所述第二指标对应的集合中的最大值未处于对应的第一区间;
将确定出的第二指标和所述第一性能指标确定为与所述设定故障相关的第一指标。
这里,电子设备基于第二时段对应的第二时序数据,确定出在对应的第二时段内首个触发告警的第一性能指标;基于第二时段对应的第一时序数据确定出第一集合;基于第二性能指标在第二时段对应的时序数据,确定出对应的第二性能指标对应的第二集合。其中,第二性能指标表征至少一个性能指标中除第一性能指标之外的任一性能指标。
当第一集合中的最大值未处于第一集合对应的第一区间时,将日志信息量确定为与设定故障相关的第二指标。当第一集合中的最大值处于第一集合对应的第一区间时,表征日志信息量不是与设定故障相关的第二指标。
当第二集合中的最大值未处于对应的第一区间时,将该第二集合对应的性能指标确定为与设定故障相关的第二指标。当第二集合中的最大值处于对应的第一区间时,表征 该第二集合对应的性能指标不是与设定故障相关的第二指标。
在确定出所有第二指标的情况下,电子设备将第一性能指标和确定出的所有第二指标,确定为与设定故障相关的第一指标。
本实施例提供的方案中,将在对应的第二时段内首个触发告警的第一性能指标识别为与设定故障相关的第一指标,然后再按照上述方法确定出与设定故障相关的第二指标,将第一性能指标和确定出的第二指标识别为与设定故障相关的第一指标。不需要通过第一性能指标对应的第二时序数据确定第一性能指标是否与设定故障相关,可以节省确定第一指标的时间,提高故障识别模型的训练效率。
作为本申请的应用实施例,图6示出了本申请应用实施例提供的一种训练故障识别模型的方法的实现流程示意图。参照图6,训练故障识别模型的方法包括:
S601:按照设定的时间间隔对历史日志进行采样,得到至少一个第一日志样本。
设定的时间间隔可以为1秒。
其中,第一日志样本中每个第一日志样本均满足以下条件:
第一日志样本对应的日志级别对应的第一概率满足设定条件;
第一日志样本对应的告警类型为发生所述设定故障时监测到的告警类型;
第三概率与第四概率之间的差值小于或等于设定阈值;其中,所述第三概率表征第一日志样本对应的告警类型在所述至少一个第一日志样本中出现的概率;所述第四概率表征第一日志样本对应的告警类型在发生所述设定故障时被监测到的概率。
S602:基于所述至少一个第一日志样本,确定出至少一个第一概率和至少一个第二概率;其中,第一概率为每种日志级别对应的概率;第二概率为每个设定词组在对应的日志级别中对应的第二概率。
S602的实现过程请参照上述实施例中S201中的相关描述,此处不赘述。
S603:基于确定出的第一概率和第二概率,确定出历史日志对应的时序数据。
这里,电子设备基于确定出的第一概率和第二概率,确定出每个时刻打印出的历史日志中每条日志的总信息量;基于历史日志中每条日志的总信息量,确定出每个时刻输出的日志信息量,从而得到历史日志对应的时序数据。确定历史日志对应的时序数据实现方式请参照上述实施例S201~S202中的相关描述,此处不赘述。
S604:基于至少一个第一时段中每个第一时段对应的时序数据,确定出与设定故障相关的第一指标。
这里,每个第一时段对应的时序数据包括第一时序数据和第二时序数据。第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。第二时序数据基于监测程序监测到的至少一个性能指标中每个性能指标在每个时刻对应的监测数据得到。
S604的实现方式请参照上述S301中的相关描述,此处不赘述。
S605:基于所述第一指标在第一时段中的第二时段对应的时序数据,构建设定故障对应的至少一个正样本。
这里,S605的实现方式请参照上述S302中的相关描述,此处不赘述。
S606:基于所述第一指标在第一时段中的第三时段对应的时序数据,构建设定故障对应的至少一个负样本。
这里,S606的实现方式请参照上述S303中的相关描述,此处不赘述。
S607:基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型。
这里,S607的实现方式请参照上述S102中的相关描述,此处不赘述。
本实施例提供的方案中,基于所有第一日志样本确定出至少一个第一概率和至少一 个第二概率;基于确定出的第一概率和第二概率,确定出至少一个第一时段中每个第一时段对应的第一时序数据;基于第一时段对应的第一时序数据和第二时序数据,确定出设定故障对应的至少一个正样本和至少一个负样本;基于至少一个正样本和至少一个负样本训练得到设定故障对应的故障识别模型。由于第一日志样本能够准确反映出发生设定故障时服务器的运行情况,因此,利用基于第一日志样本确定出的概率计算出的日志信息量,能够准确反映出发生设定故障时的日志信息量,进而提高训练得到的故障识别模型的准确度。由于第一时序数据表征第一时段的每个时刻输出的日志信息量,而日志信息量是指日志的信息量,用于度量日志传达的信息多少;在确定日志信息量时不需要分析历史日志的具体内容,因此可以省去对日志的文本内容进行语义分析消耗的时间,提高获取训练样本的效率,进而提高故障识别模型的训练效率。
上面在介绍了设定故障对应的故障识别模型的训练方法之后,下面介绍通过上面的方式训练的到的故障识别模型进行故障识别的实现过程。需要说明的是,电子设备可以基于上述实施例训练出不同的设定故障对应的故障识别模型。
图7示出了本申请实施例提供的一种故障识别方法的实现流程示意图。在本申请实施例中,故障识别方法的执行主体为电子设备,例如,电脑、服务器等。电子设备可以监测DBaaS服务器的性能指标的相关数据,以及获取DBaaS服务器的实时日志。基于故障识别模型进行故障识别的电子设备与训练故障识别模型的电子设备,可以相同,也可以不同。
参照图7,故障识别方法包括:
S701:在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据。
其中,所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据。
电子设备基于DBaaS服务器在每个时刻输出的实时日志,确定出第一时序数据;基于至少一个性能指标中每个性能指标在每个时刻对应的实时监测数据,确定出每个性能指标对应的第二时序数据。
电子设备在检测到第一性能指标的数据触发告警的情况下,基于第一性能指标当前的告警时刻,从确定出的时序数据中,确定出第一时段对应的时序数据。其中,第一性能指标为至少一个性能指标中的任一性能指标。第一时段对应的时序数据包括第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据。
在实际应用中,第一时段可以为第一性能指标的告警时刻的前15分钟。例如,电子设备在当天上午8点30分检测到第一性能指标的数据触发告警,第一时段为8点15分到8点30分这个时段。
S702:基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标。
这里,每种设定故障对应的相关性能指标的实现方式,请参照上述确定出与设定故障相关的第一指标的相关描述,此处不赘述。
S703:将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果。
这里,故障识别模型上述任一实施例对应的训练故障识别模型的方法训练得到。每个故障识别模型输出的识别结果为0-1之间的数值。识别结果表征发生的故障是故障识别模型对应的第一设定故障的概率。
S704:将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障。
置信度(confidence):表示对应的第一设定故障可信的概率,例如某个第一设定故障的置信度为80%,则表示发生的故障为第一设定故障的概率为80%。
本实施例提供的方案中,在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据,将第一时段对应的时序数据输入至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果,将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障。基于第一时段对应的时序数据进行故障识别时,由于第一时段对应的时序数据包括实时日志对应的第一时序数据,第一时序数据表征基于告警时刻确定的第一时段中的每个时刻输出的日志信息量,在确定日志信息量时不需要分析日志的具体内容,因此可以节省对日志的文本内容进行语义分析消耗的时间,提高故障分析效率。
图8示出了本申请另一实施例提供的一种故障识别方法的实现流程示意图。参照图8,在图7对应的实施例的基础上,图8对应的实施例中,故障识别方法还包括:
S705:基于至少一台服务器中每台服务器发生的历史故障和每种历史故障对应的发生次数,确定出至少一台服务器中所有服务器的健康评分,以在健康评分大于或等于设定阈值的服务器中部署新的数据库实例。
电子设备基于服务器发生的历史故障的标识,以及基于每种历史故障对应的次数,从设定的评分表中查询对应的服务器的健康评分。其中,
电子设备中预先存储了设定的评分表,设定的评分表包括至少一种故障等级中每个故障等级对应的所有设定故障在不同的次数范围内对应的分值。也就是说,设定的评分表中为每个故障等级配置了对应的分值范围,以及为每个故障等级对应的设定故障配置了不同的次数范围对应的分值,该分值处于对应的故障等级对应的分值范围。
需要说明的是,设定的评分表中可以按照故障等级的优先级对所有设定故障进行排序,故障等级的优先级越高,表征该故障等级对应的设定故障对服务器的影响程度越大。故障等级的优先级越高,对应的分值范围越大。同一故障等级对应的设定故障发生的次数越多,对应的分值越低。
例如,当第一故障等级的优先级大于第二故障等级的优先级时,第一故障等级对应的分值范围小于第二故障等级的分值范围。
针对同一故障等级对应的设定故障,在第一次数范围大于第二次数范围时,第一次数范围对应的分值小于第二次数范围的分值。
需要说明的是,电子设备可以在服务器处于空闲状态的情况下,基于服务器发生的历史故障和每种历史故障对应的次数,确定出对应的服务器的健康评分。这样可以避免因对服务器进行健康评分而干扰服务器执行的进程,降低服务器的数据处理效率。
本实施例提供的方案中,基于服务器发生的历史故障和每种历史故障对应的发生次数,确定对应的服务器的健康评分,以在健康评分大于或等于设定阈值的服务器中部署新的数据库实例。在服务器的健康评分大于或等于设定阈值时,表征对应的服务器存在的安全隐患较小,在健康评分大于或等于设定阈值的服务器中部署新的数据库实例,可以降低在部署新的数据库实例的过程中对应的服务器发生故障的概率,还可以减少后续运维的工作量。
为实现本申请实施例的训练故障识别模型的方法,本申请实施例还提供了一种训练故障识别模型的装置,设置在电子设备上,如图9所示,该训练故障识别模型的装置包括:
构建单元91,配置为基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;
训练单元92,配置为基于所述设定故障对应的至少一个正样本和至少一个负样本, 对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,
所述时序数据包括第一时序数据和第二时序数据;
第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;
第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
在一实施例中,训练故障识别模型的装置还包括:
计算单元,配置为基于第一时段中每个时刻打印出的每条日志对应的第一信息量和第二信息量,计算出每条日志对应的总信息量;
确定单元,配置为基于第一时段中每个时刻打印出的每条日志对应的总信息量,得到对应的第一时段对应的第一时序数据;其中,
所述第一信息量表征日志对应的日志级别出现的第一概率;
所述第二信息量表征日志中包括的所有设定词组中每个设定词组在对应的日志级别中出现的第二概率。
在一实施例中,所述计算单元还配置为:
基于至少一个第一日志样本中每个日志级别出现的次数,确定出每个日志级别对应的第一概率;
基于所述至少一个第一日志样本中每个设定词组在每个日志级别下出现的次数,确定出每个设定词组在对应的日志级别中对应的第二概率;其中,
所述至少一个第一日志样本通过对历史日志进行采样得到。
在一实施例中,至少一个第一日志样本中每个第一日志样本均满足以下条件:
第一日志样本对应的日志级别对应的第一概率满足设定条件;
第一日志样本对应的告警类型为发生所述设定故障时监测到的告警类型;
第三概率与第四概率之间的差值小于或等于设定阈值;其中,所述第三概率表征第一日志样本对应的告警类型在所述至少一个第一日志样本中出现的概率;所述第四概率表征第一日志样本对应的告警类型在发生所述设定故障时被监测到的概率。
在一实施例中,构建单元91配置为:
基于第一时段中的第二时段对应的时序数据,确定出与所述设定故障相关的第一指标;其中,所述第二时段表征对应的第一时段中发生所述设定故障的时段;
基于所述第一指标在所述第二时段对应的时序数据,构建所述设定故障对应的至少一个正样本;
基于所述第一指标在所述第一时段中的第三时段对应的时序数据,构建所述设定故障对应的至少一个负样本;其中,
所述第三时段表征对应的第一时段中未发生所述设定故障的时段;
所述第一指标包括以下至少之一:
在发生所述设定故障时告警的所有性能指标中的至少一个性能指标;
日志信息量。
在一实施例中,构建单元91配置为:
基于第二时段对应的时序数据,确定出至少两个集合;其中,所述至少两个集合包括一个第一集合和至少一个性能指标中每个性能指标对应的第二集合;所述第一集合中的元素表征相邻两个时刻输出的日志信息量之间的差值;所述第二集合中的元素表征对应的性能指标在相邻两个时刻的监测数据之间的差值;
计算出所述至少两个集合中每个集合对应的均值和标准差;
基于三西格玛准则以及基于计算出的均值和标准差,确定出所述至少两个集合中每个集合对应的第一区间;
基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出与所述设定故障相关的第一指标;其中,所述第一指标对应的集合中的最大值未处于对应的第一区间。
在一实施例中,所述第二时段对应的时序数据未包括在第二时段内首个触发告警的第一性能指标对应的第二时序数据;构建单元91用于:
基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出第二指标;其中,所述第二指标对应的集合中的最大值未处于对应的第一区间;将确定出的第二指标和所述第一性能指标确定为与所述设定故障相关的第一指标。
在一实施例中,构建单元91配置为:
基于所述第一指标的第一数量,确定出正样本的第二数量;
将所有第一指标在第二时段对应的时序数据以第一指标为最小单位进行全排列,得到全排列结果;
从所述全排列结果中选出第二数量的正样本;其中,
所述第一数量为大于或等于2的整数,所述第二数量为正整数;所述第二数量小于或等于所述第一数量对应的全排列运算结果的二分之一。
实际应用时,训练故障识别模型的装置包括的各单元可由训练故障识别模型的装置中的处理器来实现。当然,处理器需要运行存储器中存储的程序来实现上述各程序模块的功能。
需要说明的是:上述实施例提供的训练故障识别模型的装置在训练故障识别模型时,仅以上述各程序模块的划分进行举例说明,实际应用中,可以根据需要而将上述处理分配由不同的程序模块完成,即将训练故障识别模型的装置的内部结构划分成不同的程序模块,以完成以上描述的全部或者部分处理。另外,上述实施例提供的训练故障识别模型的装置与训练故障识别模型的方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
为实现本申请实施例的故障识别方法,本申请实施例还提供了一种故障识别装置,设置在电子设备上,如图10所示,该故障识别装置包括:
第一确定单元101,配置为在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据;
第二确定单元102,配置为基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标;
第一识别单元103,配置为将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果;
第二识别单元104,配置为将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障;其中,
所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;
所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据;
所述故障识别模型基于上述任一种训练故障识别模型的方法训练得到。
在一实施例中,该故障识别装置还包括:
评分单元,配置为基于至少一台服务器中每台服务器发生的历史故障和每种历史故障对应的发生次数,确定出至少一台服务器中所有服务器的健康评分,以在健康评分大于或等于设定阈值的服务器中部署新的数据库实例。
实际应用时,故障识别装置包括的各单元可由故障识别装置中的处理器来实现。当然,处理器需要运行存储器中存储的程序实现上述各程序模块的功能。
需要说明的是:上述实施例提供的故障识别模型的装置在训练故障识别模型时,仅以上述各程序模块的划分进行举例说明,实际应用中,可以根据需要而将上述处理分配由不同的程序模块完成,即将故障识别装置的内部结构划分成不同的程序模块,以完成以上描述的全部或者部分处理。另外,上述实施例提供的故障识别装置与故障识别方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
基于上述程序模块的硬件实现,且为了实现本申请实施例的方法,本申请实施例还提供了一种电子设备。图11为本申请实施例提供的电子设备的硬件组成结构示意图,如图11所示,电子设备包括:
通信接口1,能够与其它设备比如服务器等进行信息交互;
处理器2,与通信接口1连接,以实现与其它设备进行信息交互,配置为运行计算机程序时,执行上述一个或多个技术方案提供的训练故障识别模型的方法;或者执行上述一个或多个技术方案提供的故障识别方法。而所述计算机程序存储在存储器3上。
当然,实际应用时,电子设备中的各个组件通过总线系统4耦合在一起。可理解,总线系统4用于实现这些组件之间的连接通信。总线系统4除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图11中将各种总线都标为总线系统4。
本申请实施例中的存储器3配置为存储各种类型的数据以支持电子设备的操作。这些数据的示例包括:配置为在电子设备上操作的任何计算机程序。
可以理解,存储器3可以是易失性存储器或非易失性存储器,也可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(ROM,Read Only Memory)、可编程只读存储器(PROM,Programmable Read-Only Memory)、可擦除可编程只读存储器(EPROM,Erasable Programmable Read-Only Memory)、电可擦除可编程只读存储器(EEPROM,Electrically Erasable Programmable Read-Only Memory)、磁性随机存取存储器(FRAM,ferromagnetic random access memory)、快闪存储器(Flash Memory)、磁表面存储器、光盘、或只读光盘(CD-ROM,Compact Disc Read-Only Memory);磁表面存储器可以是磁盘存储器或磁带存储器。易失性存储器可以是随机存取存储器(RAM,Random Access Memory),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(SRAM,Static Random Access Memory)、同步静态随机存取存储器(SSRAM,Synchronous Static Random Access Memory)、动态随机存取存储器(DRAM,Dynamic Random Access Memory)、同步动态随机存取存储器(SDRAM,Synchronous Dynamic Random Access Memory)、双倍数据速率同步动态随机存取存储器(DDRSDRAM,Double Data Rate Synchronous Dynamic Random Access Memory)、增强型同步动态随机存取存储器(ESDRAM,Enhanced Synchronous Dynamic Random Access Memory)、同步连接动态随机存取存储器(SLDRAM,Sync Link Dynamic Random Access Memory)、直接内存总线随机存取存储器(DRRAM,Direct Rambus Random Access Memory)。本申请实施例描述的存储器3旨在包括但不限于这些和任意其它适合类型的存储器。
上述本申请实施例揭示的方法可以应用于处理器2中,或者由处理器2实现。处理器2可能是一种集成电路芯片,具有信号的处理能力。在实现过程中,上述方法的各步骤可以通过处理器2中的硬件的集成逻辑电路或者软件形式的指令完成。上述的处理器2可以是通用处理器、DSP,或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。处理器2可以实现或者执行本申请实施例中的公开的各方法、步骤及 逻辑框图。通用处理器可以是微处理器或者任何常规的处理器等。结合本申请实施例所公开的方法的步骤,可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件模块组合执行完成。软件模块可以位于存储介质中,该存储介质位于存储器3,处理器2读取存储器3中的程序,结合其硬件完成前述方法的步骤。
处理器2执行所述程序时实现本申请实施例的各个方法中多核处理器对应的流程,为了简洁,在此不再赘述。
在示例性实施例中,本申请实施例还提供了一种存储介质,即计算机存储介质,具体为计算机可读存储介质,例如包括存储计算机程序的存储器3,上述计算机程序可由处理器2执行,以完成前述图1至图6对应的实施例中的所述步骤;或者完成前述图7至图8对应的实施例中的所述步骤。计算机可读存储介质可以是FRAM、ROM、PROM、EPROM、EEPROM、Flash Memory、磁表面存储器、光盘、或CD-ROM等存储器。
在本申请所提供的几个实施例中,应该理解到,所揭露的设备和方法,可以通过其它的方式实现。以上所描述的设备实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,如:多个单元或组件可以结合,或可以集成到另一个系统,或一些特征可以忽略,或不执行。另外,所显示或讨论的各组成部分相互之间的耦合、或直接耦合、或通信连接可以是通过一些接口,设备或单元的间接耦合或通信连接,可以是电性的、机械的或其它形式的。
上述作为分离部件说明的单元可以是、或也可以不是物理上分开的,作为单元显示的部件可以是、或也可以不是物理单元,即可以位于一个地方,也可以分布到多个网络单元上;可以根据实际的需要选择其中的部分或全部单元来实现本实施例方案的目的。
另外,在本申请各实施例中的各功能单元可以全部集成在一个处理模块中,也可以是各单元分别单独作为一个单元,也可以两个或两个以上单元集成在一个单元中;上述集成的单元既可以采用硬件的形式实现,也可以采用硬件加软件功能单元的形式实现。本领域普通技术人员可以理解:实现上述方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成,前述的程序可以存储于一计算机可读取存储介质中,该程序在执行时,执行包括上述方法实施例的步骤;而前述的存储介质包括:移动存储设备、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。
需要说明的是:“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。
另外,本申请实施例所记载的技术方案之间,在不冲突的情况下,可以任意组合。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。
Claims (14)
- 一种训练故障识别模型的方法,包括:基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,所述时序数据包括第一时序数据和第二时序数据;第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
- 根据权利要求1所述的方法,其中,所述方法还包括:基于第一时段中每个时刻打印出的每条日志对应的第一信息量和第二信息量,计算出每条日志对应的总信息量;基于第一时段中每个时刻打印出的每条日志对应的总信息量,得到对应的第一时段对应的第一时序数据;其中,所述第一信息量表征日志对应的日志级别出现的第一概率;所述第二信息量表征日志中包括的所有设定词组中每个设定词组在对应的日志级别中出现的第二概率。
- 根据权利要求2所述的方法,其中,所述方法还包括:基于至少一个第一日志样本中每个日志级别出现的次数,确定出每个日志级别对应的第一概率;基于所述至少一个第一日志样本中每个设定词组在每个日志级别下出现的次数,确定出每个设定词组在对应的日志级别中对应的第二概率;其中,所述至少一个第一日志样本通过对历史日志进行采样得到。
- 根据权利要求3所述的方法,其中,至少一个第一日志样本中每个第一日志样本均满足以下条件:第一日志样本对应的日志级别对应的第一概率满足设定条件;第一日志样本对应的告警类型为发生所述设定故障时监测到的告警类型;第三概率与第四概率之间的差值小于或等于设定阈值;其中,所述第三概率表征第一日志样本对应的告警类型在所述至少一个第一日志样本中出现的概率;所述第四概率表征第一日志样本对应的告警类型在发生所述设定故障时被监测到的概率。
- 根据权利要求1-4任一项所述的方法,其中,所述基于至少一个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本,包括:基于第一时段中的第二时段对应的时序数据,确定出与所述设定故障相关的第一指标;其中,所述第二时段表征对应的第一时段中发生所述设定故障的时段;基于所述第一指标在所述第二时段对应的时序数据,构建所述设定故障对应的至少一个正样本;基于所述第一指标在所述第一时段中的第三时段对应的时序数据,构建所述设定故障对应的至少一个负样本;其中,所述第三时段表征对应的第一时段中未发生所述设定故障的时段;所述第一指标包括以下至少之一:在发生所述设定故障时告警的所有性能指标中的至少一个性能指标;日志信息量。
- 根据权利要求5所述的方法,其中,所述基于第一时段中的第二时段对应的时序数据,确定出与所述设定故障相关的第一指标,包括:基于第二时段对应的时序数据,确定出至少两个集合;其中,所述至少两个集合包括一个第一集合和至少一个性能指标中每个性能指标对应的第二集合;所述第一集合中的元素表征相邻两个时刻输出的日志信息量之间的差值;所述第二集合中的元素表征对应的性能指标在相邻两个时刻的监测数据之间的差值;计算出所述至少两个集合中每个集合对应的均值和标准差;基于三西格玛准则以及基于计算出的均值和标准差,确定出所述至少两个集合中每个集合对应的第一区间;基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出与所述设定故障相关的第一指标;其中,所述第一指标对应的集合中的最大值未处于对应的第一区间。
- 根据权利要求6所述的方法,其中,所述第二时段对应的时序数据未包括在第二时段内首个触发告警的第一性能指标对应的第二时序数据;所述基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出与所述设定故障相关的第一指标,包括:基于所述至少两个集合中每个集合对应的第一区间和对应集合中的最大值,确定出第二指标;其中,所述第二指标对应的集合中的最大值未处于对应的第一区间;将确定出的第二指标和所述第一性能指标确定为与所述设定故障相关的第一指标。
- 根据权利要求5所述的方法,其中,所述基于所述第一指标在所述第二时段对应的时序数据,构建所述设定故障对应的至少一个正样本,包括:基于所述第一指标的第一数量,确定出正样本的第二数量;将所有第一指标在第二时段对应的时序数据以第一指标为最小单位进行全排列,得到全排列结果;从所述全排列结果中选出第二数量的正样本;其中,所述第一数量为大于或等于2的整数,所述第二数量为正整数;所述第二数量小于或等于所述第一数量对应的全排列运算结果的二分之一。
- 一种故障识别方法,包括:在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据;基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标;将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果;将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障;其中,所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据;所述故障识别模型基于权利要求1至8任一项所述的方法训练得到。
- 根据权利要求9所述的方法,其中,所述方法还包括:基于至少一台服务器中每台服务器发生的历史故障和每种历史故障对应的发生次数,确定出至少一台服务器中所有服务器的健康评分,以在健康评分大于或等于设定阈值的服务器中部署新的数据库实例。
- 一种训练故障识别模型的装置,包括:构建单元,配置为基于至少一个第一时段中每个第一时段对应的时序数据,构建设定故障对应的至少一个正样本和至少一个负样本;训练单元,配置为基于所述设定故障对应的至少一个正样本和至少一个负样本,对设定的二分类模型进行训练,得到所述设定故障对应的故障识别模型;其中,所述时序数据包括第一时序数据和第二时序数据;第一时序数据表征对应的第一时段的每个时刻输出的日志信息量;第二时序数据表征至少一个性能指标中每个性能指标在对应的第一时段的每个时刻的监测数据。
- 一种故障识别装置,包括:第一确定单元,配置为在检测到第一性能指标的数据触发告警的情况下,确定出第一时段对应的时序数据;第二确定单元,配置为基于至少一种设定故障中每种设定故障对应的相关性能指标,确定出至少一种第一设定故障;所述第一设定故障的相关性能指标中包含所述第一性能指标;第一识别单元,配置为将所述第一时段对应的时序数据输入所述至少一种第一设定故障中每种第一设定故障对应的故障识别模型,得到每个故障识别模型输出的识别结果;第二识别单元,配置为将所有识别结果中置信度最高的第一设定故障确定为当前发生的故障;其中,所述第一时段的结束时刻为所述第一性能指标对应的告警时刻;所述第一时段对应的时序数据包括实时日志对应的第一时序数据和至少一个性能指标中每个性能指标对应的第二时序数据;所述故障识别模型基于权利要求1至8任一项所述的方法训练得到。
- 一种电子设备,包括:处理器和配置为存储能够在处理器上运行的计算机程序的存储器,其中,所述处理器配置为运行所述计算机程序时,执行以下至少一项:权利要求1至8任一项所述的方法的步骤;权利要求9至10任一项所述的方法的步骤。
- 一种存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现以下至少一项:权利要求1至8任一项所述的方法的步骤;权利要求9至10任一项所述的方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202011164795.1 | 2020-10-27 | ||
| CN202011164795.1A CN112308126B (zh) | 2020-10-27 | 2020-10-27 | 故障识别模型训练方法、故障识别方法、装置及电子设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022089202A1 true WO2022089202A1 (zh) | 2022-05-05 |
Family
ID=74331152
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/123363 Ceased WO2022089202A1 (zh) | 2020-10-27 | 2021-10-12 | 故障识别模型训练方法、故障识别方法、装置及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112308126B (zh) |
| WO (1) | WO2022089202A1 (zh) |
Cited By (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115118464A (zh) * | 2022-06-10 | 2022-09-27 | 深信服科技股份有限公司 | 一种失陷主机检测方法、装置、电子设备及存储介质 |
| CN115113606A (zh) * | 2022-06-21 | 2022-09-27 | 哈尔滨工业大学 | 基于频域改进sdp图的航天器故障分析方法、装置及介质 |
| CN115185735A (zh) * | 2022-08-04 | 2022-10-14 | 中国平安财产保险股份有限公司 | 软件故障识别自愈方法及相关设备 |
| CN115225470A (zh) * | 2022-07-28 | 2022-10-21 | 天翼云科技有限公司 | 一种业务异常监测方法、装置、电子设备及存储介质 |
| CN115422263A (zh) * | 2022-11-01 | 2022-12-02 | 广东亿能电力股份有限公司 | 一种电力现场多功能通用型故障分析方法及系统 |
| CN115951002A (zh) * | 2023-03-10 | 2023-04-11 | 山东省计量科学研究院 | 一种气质联用仪故障检测装置 |
| CN116089231A (zh) * | 2023-02-13 | 2023-05-09 | 北京优特捷信息技术有限公司 | 一种故障告警方法、装置、电子设备及存储介质 |
| CN116781984A (zh) * | 2023-08-21 | 2023-09-19 | 深圳市华星数字有限公司 | 一种机顶盒数据优化存储方法 |
| WO2023216457A1 (zh) * | 2022-05-11 | 2023-11-16 | 中电信数智科技有限公司 | 一种核心网与基站传输网络异常预测及定位的方法 |
| CN117076131A (zh) * | 2023-10-12 | 2023-11-17 | 中信建投证券股份有限公司 | 一种任务分配方法、装置、电子设备及存储介质 |
| WO2024197741A1 (en) * | 2023-03-30 | 2024-10-03 | Dow Global Technologies Llc | Leading indicator for overall plant health using process data augmented with natural language processing from multiple sources |
| CN119902949A (zh) * | 2025-03-31 | 2025-04-29 | 天翼云科技有限公司 | 云平台监控方法、装置、计算机设备、可读存储介质和程序产品 |
| CN120408649A (zh) * | 2025-07-07 | 2025-08-01 | 齐鲁师范学院 | 一种基于大模型的漏洞挖掘方法 |
| CN121324916A (zh) * | 2025-10-17 | 2026-01-13 | 安徽思宇微电子技术有限责任公司 | 智能量测开关的端子温度监测与故障自动诊断方法 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112308126B (zh) * | 2020-10-27 | 2024-08-23 | 深圳前海微众银行股份有限公司 | 故障识别模型训练方法、故障识别方法、装置及电子设备 |
| CN114943247B (zh) * | 2022-04-15 | 2025-05-27 | 北京宝兰德软件股份有限公司 | 性能指标时序数据的波形识别方法和装置 |
| CN115168173B (zh) * | 2022-07-25 | 2025-10-03 | 阿里巴巴(中国)有限公司 | 故障预测模型训练方法、设备故障确定方法、装置及设备 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3336636A1 (en) * | 2016-12-19 | 2018-06-20 | Palantir Technologies Inc. | Machine fault modelling |
| CN109639450A (zh) * | 2018-10-23 | 2019-04-16 | 平安壹钱包电子商务有限公司 | 基于神经网络的故障告警方法、计算机设备及存储介质 |
| CN111585799A (zh) * | 2020-04-29 | 2020-08-25 | 杭州迪普科技股份有限公司 | 网络故障预测模型建立方法及装置 |
| CN112308126A (zh) * | 2020-10-27 | 2021-02-02 | 深圳前海微众银行股份有限公司 | 故障识别模型训练方法、故障识别方法、装置及电子设备 |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107526667B (zh) * | 2017-07-28 | 2020-04-28 | 阿里巴巴集团控股有限公司 | 一种指标异常检测方法、装置以及电子设备 |
| CN110545195A (zh) * | 2018-05-29 | 2019-12-06 | 华为技术有限公司 | 网络故障分析方法及装置 |
| CN110647446B (zh) * | 2018-06-26 | 2023-02-21 | 中兴通讯股份有限公司 | 一种日志故障关联与预测方法、装置、设备及存储介质 |
| CN109446049A (zh) * | 2018-11-01 | 2019-03-08 | 郑州云海信息技术有限公司 | 一种基于监督学习的服务器错误诊断方法和装置 |
| CN109828869B (zh) * | 2018-12-05 | 2020-12-04 | 南京中兴软件有限责任公司 | 预测硬盘故障发生时间的方法、装置及存储介质 |
| CN109634828A (zh) * | 2018-12-17 | 2019-04-16 | 浪潮电子信息产业股份有限公司 | 故障预测方法、装置、设备及存储介质 |
| CN110838075A (zh) * | 2019-05-20 | 2020-02-25 | 全球能源互联网研究院有限公司 | 电网系统暂态稳定的预测模型的训练及预测方法、装置 |
| CN110598802B (zh) * | 2019-09-26 | 2021-07-27 | 腾讯科技(深圳)有限公司 | 一种内存检测模型训练的方法、内存检测的方法及装置 |
| CN110647456B (zh) * | 2019-09-29 | 2022-12-27 | 苏州浪潮智能科技有限公司 | 一种存储设备的故障预测方法、系统及相关装置 |
| CN111045894B (zh) * | 2019-12-13 | 2024-02-13 | 贵州广思信息网络有限公司广州分公司 | 数据库异常检测方法、装置、计算机设备和存储介质 |
| CN111338836B (zh) * | 2020-02-24 | 2023-09-01 | 北京奇艺世纪科技有限公司 | 处理故障数据的方法、装置、计算机设备和存储介质 |
| CN111752775B (zh) * | 2020-05-28 | 2022-11-18 | 苏州浪潮智能科技有限公司 | 一种磁盘故障预测方法和系统 |
| CN111611146B (zh) * | 2020-06-18 | 2023-05-16 | 南方电网科学研究院有限责任公司 | 一种微服务故障预测方法和装置 |
-
2020
- 2020-10-27 CN CN202011164795.1A patent/CN112308126B/zh active Active
-
2021
- 2021-10-12 WO PCT/CN2021/123363 patent/WO2022089202A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3336636A1 (en) * | 2016-12-19 | 2018-06-20 | Palantir Technologies Inc. | Machine fault modelling |
| CN109639450A (zh) * | 2018-10-23 | 2019-04-16 | 平安壹钱包电子商务有限公司 | 基于神经网络的故障告警方法、计算机设备及存储介质 |
| CN111585799A (zh) * | 2020-04-29 | 2020-08-25 | 杭州迪普科技股份有限公司 | 网络故障预测模型建立方法及装置 |
| CN112308126A (zh) * | 2020-10-27 | 2021-02-02 | 深圳前海微众银行股份有限公司 | 故障识别模型训练方法、故障识别方法、装置及电子设备 |
Cited By (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023216457A1 (zh) * | 2022-05-11 | 2023-11-16 | 中电信数智科技有限公司 | 一种核心网与基站传输网络异常预测及定位的方法 |
| CN115118464A (zh) * | 2022-06-10 | 2022-09-27 | 深信服科技股份有限公司 | 一种失陷主机检测方法、装置、电子设备及存储介质 |
| CN115113606A (zh) * | 2022-06-21 | 2022-09-27 | 哈尔滨工业大学 | 基于频域改进sdp图的航天器故障分析方法、装置及介质 |
| CN115225470B (zh) * | 2022-07-28 | 2023-10-13 | 天翼云科技有限公司 | 一种业务异常监测方法、装置、电子设备及存储介质 |
| CN115225470A (zh) * | 2022-07-28 | 2022-10-21 | 天翼云科技有限公司 | 一种业务异常监测方法、装置、电子设备及存储介质 |
| CN115185735A (zh) * | 2022-08-04 | 2022-10-14 | 中国平安财产保险股份有限公司 | 软件故障识别自愈方法及相关设备 |
| CN115422263A (zh) * | 2022-11-01 | 2022-12-02 | 广东亿能电力股份有限公司 | 一种电力现场多功能通用型故障分析方法及系统 |
| CN115422263B (zh) * | 2022-11-01 | 2023-01-13 | 广东亿能电力股份有限公司 | 一种电力现场多功能通用型故障分析方法及系统 |
| CN116089231A (zh) * | 2023-02-13 | 2023-05-09 | 北京优特捷信息技术有限公司 | 一种故障告警方法、装置、电子设备及存储介质 |
| CN116089231B (zh) * | 2023-02-13 | 2023-09-15 | 北京优特捷信息技术有限公司 | 一种故障告警方法、装置、电子设备及存储介质 |
| CN115951002A (zh) * | 2023-03-10 | 2023-04-11 | 山东省计量科学研究院 | 一种气质联用仪故障检测装置 |
| WO2024197741A1 (en) * | 2023-03-30 | 2024-10-03 | Dow Global Technologies Llc | Leading indicator for overall plant health using process data augmented with natural language processing from multiple sources |
| CN116781984A (zh) * | 2023-08-21 | 2023-09-19 | 深圳市华星数字有限公司 | 一种机顶盒数据优化存储方法 |
| CN116781984B (zh) * | 2023-08-21 | 2023-11-07 | 深圳市华星数字有限公司 | 一种机顶盒数据优化存储方法 |
| CN117076131A (zh) * | 2023-10-12 | 2023-11-17 | 中信建投证券股份有限公司 | 一种任务分配方法、装置、电子设备及存储介质 |
| CN117076131B (zh) * | 2023-10-12 | 2024-01-23 | 中信建投证券股份有限公司 | 一种任务分配方法、装置、电子设备及存储介质 |
| CN119902949A (zh) * | 2025-03-31 | 2025-04-29 | 天翼云科技有限公司 | 云平台监控方法、装置、计算机设备、可读存储介质和程序产品 |
| CN120408649A (zh) * | 2025-07-07 | 2025-08-01 | 齐鲁师范学院 | 一种基于大模型的漏洞挖掘方法 |
| CN121324916A (zh) * | 2025-10-17 | 2026-01-13 | 安徽思宇微电子技术有限责任公司 | 智能量测开关的端子温度监测与故障自动诊断方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112308126A (zh) | 2021-02-02 |
| CN112308126B (zh) | 2024-08-23 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022089202A1 (zh) | 故障识别模型训练方法、故障识别方法、装置及电子设备 | |
| US10805151B2 (en) | Method, apparatus, and storage medium for diagnosing failure based on a service monitoring indicator of a server by clustering servers with similar degrees of abnormal fluctuation | |
| CN109284269B (zh) | 异常日志分析方法、装置、存储介质及服务器 | |
| US20190243743A1 (en) | Unsupervised anomaly detection | |
| WO2022001125A1 (zh) | 一种存储系统的存储故障预测方法、系统及装置 | |
| CN109710518A (zh) | 脚本审核方法及装置 | |
| WO2025129877A1 (zh) | 一种异构硬盘系统故障预警方法及装置 | |
| CN113298638A (zh) | 根因定位方法、电子设备及存储介质 | |
| CN110245077A (zh) | 一种程序异常的响应方法及设备 | |
| CN108804136A (zh) | 一种基于名称语义的配置项类型约束推断方法 | |
| Wu et al. | Invalid bug reports complicate the software aging situation | |
| CN111897696A (zh) | 服务器集群硬盘状态检测方法、装置、电子设备及存储介质 | |
| CN117216095A (zh) | 结构化查询语句检测方法、装置、设备及介质 | |
| CN113806178B (zh) | 一种集群节点故障检测方法及装置 | |
| CN118295843A (zh) | 一种故障定位方法以及计算设备 | |
| CN120994446A (zh) | 基于大语言模型与业务拓扑的告警事件根因分析方法、装置及设备 | |
| WO2023103344A1 (zh) | 一种数据处理方法、装置、设备及存储介质 | |
| CN119621386A (zh) | 一种故障判断方法、装置、电子设备及存储介质 | |
| CN118569831B (zh) | 一种数据中心健康度评分的估算方法及计算设备 | |
| US20250371271A1 (en) | Inference model training and tuning using augmented questions and answers | |
| US12613764B2 (en) | Managing data processing system failures using a predictive model as a controller and hidden knowledge from predictive models | |
| CN119561847A (zh) | 基于大模型的预警方法、装置、电子设备及存储介质 | |
| WO2024244575A1 (zh) | 电子设备及其性能检测方法、介质 | |
| CN118113508A (zh) | 网卡故障风险预测方法、装置、设备及介质 | |
| CN115114070A (zh) | 一种故障诊断方法、装置、设备及介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21884931 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 16.08.2023) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21884931 Country of ref document: EP Kind code of ref document: A1 |