WO2025236282A1 - Incident triage and root cause analysis - Google Patents

Incident triage and root cause analysis

Info

Publication number
WO2025236282A1
WO2025236282A1 PCT/CN2024/093949 CN2024093949W WO2025236282A1 WO 2025236282 A1 WO2025236282 A1 WO 2025236282A1 CN 2024093949 W CN2024093949 W CN 2024093949W WO 2025236282 A1 WO2025236282 A1 WO 2025236282A1
Authority
WO
WIPO (PCT)
Prior art keywords
correlated
metric
incident
metrics
deviation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/093949
Other languages
French (fr)
Inventor
Haoshuang Li
Liangfei Su
Wentao Li
Huai JIANG
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
eBay Inc
Original Assignee
eBay Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by eBay Inc filed Critical eBay Inc
Priority to PCT/CN2024/093949 priority Critical patent/WO2025236282A1/en
Publication of WO2025236282A1 publication Critical patent/WO2025236282A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/07Responding to the occurrence of a fault, e.g. fault tolerance
    • G06F11/0703Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
    • G06F11/079Root cause analysis, i.e. error or fault diagnosis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/07Responding to the occurrence of a fault, e.g. fault tolerance
    • G06F11/0703Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
    • G06F11/0706Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment
    • G06F11/0709Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment in a distributed system consisting of a plurality of standalone computer nodes, e.g. clusters, client-server systems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/3003Monitoring arrangements specially adapted to the computing system or computing system component being monitored
    • G06F11/3006Monitoring arrangements specially adapted to the computing system or computing system component being monitored where the computing system is distributed, e.g. networked systems, clusters, multiprocessor systems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/3003Monitoring arrangements specially adapted to the computing system or computing system component being monitored
    • G06F11/302Monitoring arrangements specially adapted to the computing system or computing system component being monitored where the computing system component is a software system
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/34Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
    • G06F11/3409Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment for performance assessment
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F11/00Error detection; Error correction; Monitoring
    • G06F11/30Monitoring
    • G06F11/34Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
    • G06F11/3452Performance evaluation by statistical analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F2201/00Indexing scheme relating to error detection, to error correction, and to monitoring
    • G06F2201/865Monitoring of software

Definitions

  • FIG. 1 is a block diagram showing an example data system that includes a data management system, according to various embodiments of the present disclosure.
  • FIG. 2 is a block diagram illustrating an example data management system that facilitates incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • FIG. 3 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • FIG. 4 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • FIG. 5 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • FIG. 6 is a diagram illustrating an example data management system that facilitates incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • FIG. 7 is a diagram illustrating an example topology graph and an example causal graph, according to various embodiments of the present disclosure.
  • FIG. 8 is a diagram illustrating an example data table that stores causal links derived from historical data and/or human experience, according to various embodiments of the present disclosure.
  • FIG. 9 is a diagram illustrating line graphs representing monitoring data associated with business metrics and correlated metrics, according to various embodiments of the present disclosure.
  • FIG. 10 is a block diagram illustrating a representative software architecture, which may be used in conjunction with various hardware architectures herein described, according to various embodiments of the present disclosure.
  • FIG. 11 is a block diagram illustrating components of a machine able to read instructions from a machine storage medium and perform any one or more of the methodologies discussed herein according to various embodiments of the present disclosure.
  • Incident triage involves assessing incidents to determine their severity, urgency, and priority for resolution.
  • An incident can refer to any event or occurrence that disrupts the normal operation of services and requires a response to mitigate its impact.
  • Incident triage helps prioritize resources and responses effectively to minimize the impact of incidents on users and the business.
  • Root cause analysis involves identifying the underlying causes of incidents. Root cause analysis not only addresses the symptoms or immediate causes of an incident but also provides insights into the underlying systemic issues and/or weaknesses in processes, systems, or infrastructure that need to be addressed to prevent recurrence.
  • Various embodiments involve a data management system that receives an alert of an incident as input.
  • the alert includes contextual information about deviating metrics associated with the incident.
  • a metric refers to a quantitative measurement that provides insight into various aspects of a service's behavior.
  • a metric can be represented by time series data.
  • the data management system Based on the contextual information, the data management system performs a time series similarity search to identify the monitoring data of the incident and filters out noise by focusing on the residuals of the time series, considering factors like offset and direction. This refined data is then mapped into the predefined categories of signals (e.g., traffic, error rate, saturation, and latency) , which are key performance indicators of services (e.g., microservices) provided by a cloud computing environment.
  • signals e.g., traffic, error rate, saturation, and latency
  • Mapped metrics are overlaid onto a topological graph that represents services and the associated infrastructure components (e.g., service mesh) , providing a comprehensive view of the cloud environment. Further, the data management system constructs a causal graph (e.g., signal causal graph) based on this topological representation and causal links derived from historical data and human experience.
  • a weighted link analysis algorithm e.g., a weighted PageRank algorithm
  • a data management system receives an incident alert associated with a deviation in a metric.
  • deviation can include a spike, a dip, a fluctuation, or a pattern, indicating a departure from an expected behavior of a metric.
  • An incident can refer to any event (e.g., anomaly event) or occurrence that disrupts the normal operation of a service (e.g., microservice) and requires a response to mitigate its impact.
  • a payment failure event can be an incident associated with the payment processing service.
  • an alert associated with an incident includes contextual data representing the deviation.
  • the contextual data can include a query (e.g., raw query) associated with the deviation in the metric.
  • the data management system identifies a plurality of correlated metrics based on the contextual data (e.g., query) included in the incident alert. Specifically, the data management system can execute the query against a database to retrieve monitoring data associated with the deviation. The data management system then performs a similarity search (e.g., a time series similarity search) based on the retrieved monitoring data to identify the plurality of correlated metrics. In response to performing the similarity search, the data management system identifies the plurality of correlated metrics and assigns a similarity value to each of the plurality of correlated metrics. Each similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation. In various embodiments, each similarity value is associated with a direction indicator corresponding to a deviation type. Examples of deviation types include spikes, surges, dips, drops, fluctuations, and patterns.
  • deviation types include spikes, surges, dips, drops, fluctuations, and patterns.
  • Example algorithms (or measures) associated with the similarity search include, without limitation, Minkowski distance measure, cosine similarity measure, Pearson correlation coefficient measure, Dynamic Time Warping (DTW) measure, kernel-based similarity measure, longest common subsequence algorithm, slope-based similarity, and k-shape algorithm.
  • Minkowski distance measure cosine similarity measure
  • Pearson correlation coefficient measure Pearson correlation coefficient measure
  • Dynamic Time Warping (DTW) measure kernel-based similarity measure
  • longest common subsequence algorithm longest common subsequence algorithm
  • slope-based similarity slope-based similarity
  • k-shape algorithm k-shape algorithm
  • the queried database stores time series monitoring data using a custom format optimized for efficient querying and storage.
  • monitoring data can be organized as time series, where each time series represents a stream of timestamped metric values associated with a specific metric and set of labels.
  • the data management system generates a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals.
  • the operation of generating the plurality of labels can include mapping the plurality of correlated metrics into the plurality of categories of signals.
  • Each signal in a category of signals e.g., traffic, error rate, saturation, and latency
  • the data management system constructs a causal graph (e.g., a signal causal graph) based on the plurality of labels and a topology graph representing the plurality of correlated metrics.
  • a causal graph e.g., a signal causal graph
  • An example causal graph also referred to as a causal inference graph, can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) .
  • a causal graph can specifically focus on the causal relationships between signals or time series data.
  • An example topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate the flow of requests between them.
  • Infrastructure components such as service mesh, sit between microservices and handle service-to-service communication, providing features like load balancing, service discovery, and security.
  • Each microservice or infrastructure component can include (or correspond to) one or more metrics.
  • the data management system uses a link analysis algorithm (e.g., weighted PageRank algorithm) to determine the root cause of an incident. The determination is at least based on the constructed causal graph and a plurality of weights assigned to the plurality of correlated metrics.
  • the root cause of the incident can correspond to a correlated metric from the plurality of correlated metrics.
  • the data management system applies the plurality of similarity values assigned to the plurality of correlated metrics as a plurality of weights to the link analysis algorithm. Applying weights to the weighted link analysis algorithm involves assigning numerical values to the links between nodes in a network. These weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes.
  • the data management system uses the weighted link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph.
  • the scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence.
  • the data management system then ranks the plurality of scores in descending order and identifies the metric with the top rank (i.e., the metric with the highest ranking) as the root cause of the incident.
  • the data management system constructs the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics.
  • Tabular data can be represented by a causal link table, which can be updated based on the outcomes of previous incident triage.
  • a causal link represents a cause-and-effect relationship between a pair of correlated metrics.
  • the plurality of causal links can be derived from historical data and/or human experience.
  • FIG. 1 is a block diagram showing an example data system 100 that includes a data management system 122 (also referred to as system 122) , according to various embodiments of the present disclosure.
  • the data system 100 can facilitate incident triage and root cause analysis in cloud computing environments.
  • the data system 100 includes one or more client devices 102, a server system 108, and a network 106 (e.g., Internet, wide-area-network (WAN) , local-area-network (LAN) , wireless network) that communicatively couples them together.
  • Each client device 102 can host a number of applications, including a client software application 104.
  • the client software application 104 can communicate data with the server system 108 via a network 106. Accordingly, the client software application 104 can communicate and exchange data with the server system 108 via network 106.
  • the server system 108 provides server-side functionality via the network 106 to the client software application 104. While certain functions of the data system 100 are described herein as being performed by the data management system 122 on the server system 108, it will be appreciated that the location of certain functionality within the server system 108 is a design choice. For example, it may be technically preferable to initially deploy certain technology and functionality within the server system 108, but to later migrate this technology and functionality to the client software application 104.
  • the server system 108 supports various services and operations that are provided to the client software application 104 by the data management system 122. Such operations include transmitting data from the data management system 122 to the client software application 104, receiving data from the client software application 104 at the data management system 122, and the data management system 122 processing data generated by the client software application 104. Data exchanges within the data system 100 may be invoked and controlled through operations of software component environments available via one or more endpoints, or functions available via one or more user interfaces of the client software application 104, which may include web-based user interfaces provided by the server system 108 for presentation at the client device 102.
  • an Application Program Interface (API) server 110 and a web server 112 is coupled to an application server 116, which hosts the data management system 122.
  • the application server 116 is communicatively coupled to a database server 118, which facilitates access to a database 120 that stores data associated with the application server 116, including data that may be generated or used by the data management system 122.
  • the API server 110 receives and transmits data (e.g., API calls, commands, requests, responses, and authentication data) between the client device 102 and the application server 116.
  • data e.g., API calls, commands, requests, responses, and authentication data
  • the API server 110 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the client software application 104 in order to invoke the functionality of the application server 116.
  • the API server 110 exposes various functions supported by the application server 116 including, without limitation, user registration; login functionality; data object operations (e.g., generating, storing, retrieving, encrypting, decrypting, transferring, access rights, licensing) ; and/or user communications.
  • the server system 108, or the data management system 122 may extract user data from one or more third-party platforms (e.g., third-party social media platforms) .
  • the extracted data may be open-source poster data associated with targeted influencers on the one or more third-party platforms 124 and may include user profile data, activity data, and media posted (either created and/or shared) by the one or more influencers.
  • the media (or media data) include text, image, video, audio, and metadata.
  • Example metadata may include hashtags and labels.
  • the web server 112 can support various functionality of the data management system 122 of the application server 116.
  • FIG. 2 is a block diagram illustrating an example data management system 200 that facilitates incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • the data management system 200 represents an example of the data management system 122 described with respect to FIG. 1.
  • the data management system 200 comprises an incident alert receiving component 210, a correlated metric identifying component 220, a metric categorizing component 230, a causal graph constructing component 240, and a root cause determining component 250.
  • one or more of the incident alert receiving component 210, the correlated metric identifying component 220, the metric categorizing component 230, the causal graph constructing component 240, and the root cause determining component 250 are implemented by one or more hardware processors 202.
  • Data generated by one or more of the incident alert receiving component 210, the correlated metric identifying component 220, the metric categorizing component 230, the causal graph constructing component 240, and the root cause determining component 250 may be stored in a database (or datastore) 260 of the data management system 200.
  • the incident alert receiving component 210 is configured to receive incident alerts generated based on the detection of incidents as input.
  • An incident alert can include contextual information (also referred to as contextual data) on the deviating metric.
  • Contextual data can include a query (e.g., raw query) associated with the deviating metric.
  • the correlated metric identifying component 220 is configured to identify a plurality of correlated metrics based on the contextual data (e.g., query) included in the incident alert.
  • the correlated metric identifying component 220 is configured to perform a similarity search (e.g., a time series similarity search) based on the monitoring data retrieved in the query results to identify the correlated metrics.
  • Each correlated metric is assigned a similarity value representing a degree of similarity between the correlated metric and the deviating metric.
  • Each similarity value can be associated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  • the metric categorizing component 230 is configured to generate a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals.
  • the metric categorizing component 230 is configured to map the plurality of correlated metrics into the plurality of categories of signals.
  • Each signal in a category of signals (e.g., traffic, error rate, saturation, and latency) corresponds to a key performance indicator for monitoring and managing the performance of a service in the cloud computing platform.
  • the causal graph constructing component 240 is configured to construct a causal graph (e.g., a signal causal graph) based on the plurality of labels and a topology graph representing the plurality of correlated metrics.
  • a causal graph can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) .
  • a topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate invocation and dependency relationships between them.
  • the root cause determining component 250 is configured to use a link analysis algorithm (e.g., weighted PageRank algorithm) to determine the root cause of an incident. The determination is at least based on the constructed causal graph and a plurality of weights assigned to the plurality of correlated metrics.
  • the root cause of the incident corresponds to a correlated metric from the plurality of correlated metrics.
  • the root cause determining component 250 is configured to apply the similarity values as weights to the link analysis algorithm. These weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes.
  • the root cause determining component 250 is configured to use the weighted link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph.
  • the scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence.
  • the plurality of scores is ranked in descending order.
  • the metric with the top rank i.e., the metric with the highest ranking
  • FIG. 3 is a flowchart illustrating an example method 300 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • example methods described herein may be performed by a machine in accordance with some embodiments.
  • method 300 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof.
  • An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture.
  • hardware processors e.g., central processing units or graphics processing units
  • Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry.
  • the operations of method 300 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 300.
  • an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel.
  • a processor receives an incident alert associated with a deviation in a metric.
  • the incident alert includes contextual data representing the deviation.
  • a processor identifies a plurality of correlated metrics based on the contextual data.
  • Contextual data can include a query (e.g., raw query) associated with the deviating metric.
  • a processor generates a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals.
  • a processor maps the plurality of correlated metrics into the plurality of categories of signals.
  • Each signal in a category of signals (e.g., traffic, error rate, saturation, and latency) corresponds to a key performance indicator for monitoring and managing the performance of a service in the cloud computing platform.
  • a processor constructs a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics.
  • a causal graph can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) .
  • a topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate the flow of requests between them.
  • a processor uses a link analysis algorithm (e.g., a weighted PageRank algorithm) to determine the root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics.
  • the determination of the root cause is further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics.
  • a causal link represents a cause-and-effect relationship between a pair of correlated metrics.
  • the root cause of the incident corresponds to one of the plurality of correlated metrics.
  • method 300 can include an operation where a graphical user interface is displayed (or caused to be displayed) by the hardware processor.
  • the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface.
  • This operation for displaying the graphical user interface can be separate from operations 302 through 310 or, alternatively, form part of one or more of operations 302 through 310.
  • FIG. 4 is a flowchart illustrating an example method 400 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 400 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture.
  • a hardware processors e.g., central processing units or graphics processing units
  • a computing device e.g., a desktop, server, laptop, mobile phone, tablet, etc.
  • Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry.
  • the operations of method 400 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 400.
  • an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown.
  • the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 400 can be performed dependently or independently from operations in method 300.
  • a processor retrieves the monitoring data associated with the deviation by executing a query against a database.
  • the query is included in the contextual data described herein.
  • the queried database stores time series monitoring data using a custom format optimized for efficient querying and storage.
  • monitoring data can be organized as time series, where each time series represents a stream of timestamped metric values associated with a specific metric and set of labels.
  • a processor performs a similarity search (e.g., a time series similarity search) based on the monitoring data associated with the deviation.
  • a similarity search e.g., a time series similarity search
  • each correlated metric is assigned a similarity value representing a degree of similarity between the correlated metric and the deviating metric.
  • Each similarity value can be associated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  • a processor identifies a plurality of correlated metrics based on the similarity search.
  • a processor assigns a similarity value to each of the plurality of correlated metrics.
  • a similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation.
  • Each similarity value can be generated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  • method 400 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor.
  • the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface.
  • This operation for displaying the graphical user interface can be separate from operations 402 through 408 or, alternatively, form part of one or more of operations 402 through 408.
  • FIG. 5 is a flowchart illustrating an example method 500 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
  • example methods described herein may be performed by a machine in accordance with some embodiments.
  • method 500 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof.
  • An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture.
  • hardware processors e.g., central processing units or graphics processing units
  • Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry.
  • the operations of method 500 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 500.
  • an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 500 can be performed dependently or independently from operations in method 300 and method 400.
  • a processor applies a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm.
  • the weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes.
  • a processor uses the link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph and the plurality of similarity values described herein.
  • Each similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation.
  • the scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence.
  • a processor identifies a correlated metric with a top rank as the root cause of the incident.
  • method 500 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor.
  • the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface.
  • This operation for displaying the graphical user interface can be separate from operations 502 through 508 or, alternatively, form part of one or more of operations 502 through 508.
  • a topology graph 610 can be generated based on data from tracing/log 622 and data from CMDB (Configuration Management Database) 624.
  • CMDB data can include details about hardware, software, network devices, applications, and their relationships and dependencies in an organization’s IT infrastructure. Tracing and log data include monitoring data that records invocation and dependency relationships between microservices.
  • FIG. 7 is a diagram 700 illustrating an example topology graph and an example causal graph, according to various embodiments of the present disclosure.
  • an example topology graph includes microservice S1, microservice S2, microservice S3, microservice S4.
  • a topology graph shows invocation and dependency relationships between microservices and the associated infrastructure components (not shown) .
  • Each microservice can be associated with (or include) one or more metrics used to assess the services' effectiveness, efficiency, and quality.
  • microservice S0 is associated with metric 702.
  • Microservice S1 is associated with metrics 704 and 706.
  • Microservice S2 is associated with metrics 708, 710, and 712.
  • Microservice S3 is associated with metric 714.
  • the numbers and types of metrics needed for a given service are determined based on various considerations, including the nature of the service (e.g., the level of availability, required response time, data processing reliability) and business needs.
  • Each metric is assigned a label and a direction representing a deviation. As shown, the label “ERROR” and the direction “+” are generated for the metric 702.
  • Microservice S1 includes metric 704 and metric 706.
  • Metric 704 is also assigned the label “ERROR” and the direction “+. ”
  • Metric 706 is assigned the label “LATENCY” and the direction “+. ”
  • Direction “+” represents a spike or a surge in deviation of a metric, whereas direction “-” represents a dip or a drop.
  • an example causal graph includes metrics 702, 704, 706, 708, 710, 712, and 714, connected by arrows marked with “CAUSE. ”
  • a saturation in metric 708 causes a latency in metric 706, which in turn causes an error in metric 704.
  • FIG. 9 is a diagram 900 illustrating line graphs representing monitoring data associated with business metrics and correlated metrics, according to various embodiments of the present disclosure.
  • a time series similarity search is performed to identify the monitoring data of the incident and filter out noise by focusing on the residuals of the time series, considering factors like offset and direction.
  • line graph 902 represents monitoring data associated with a business metric between 3: 10 am and 6: 10 am on 3/25.
  • Surge 910 is observed at 6: 10 am and is determined to be an incident (e.g., Shipping Label Failure) that triggers an incident alert.
  • Line graph 906 represents monitoring data associated with a correlated metric identified as the result of a similarity search described herein.
  • Line graph 906 illustrates that surge 912 occurred at 6: 12 am, around the same time surge 910 is observed. The deviations in both line graphs (i.e., 902 and 906) are in the same direction.
  • Line graphs 904 and 908 are illustrated to show metrics with deviation in opposite directions that can also be determined as similar (or correlated) based on a similarity search described herein.
  • line graph 904 represents monitoring data associated with a business metric between 12: 38 pm and 3: 38 pm on 4/16. Dip 914 in line graph 904 is observed at 3: 38 pm.
  • Line graph 908 represents monitoring data associated with a correlated metric identified as the result of a similarity search.
  • surge 916 is observed at the same time as dip 914. The surge 916 and the dip 914 are deviations in opposite directions.
  • Described implementations of the subject matter can include one or more features, alone or in combination as illustrated below by way of examples.
  • Example 1 A system comprising: one or more hardware processors; and at least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
  • Example 2 The system of Example 1, wherein the operation of generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises: mapping the plurality of correlated metrics into the plurality of categories of signals.
  • Example 3 The system of any one or more of Examples 1 or 2, wherein each signal in a category of signals corresponds to a key performance indicator for monitoring and managing performance of a service.
  • Example 4 The system of any one or more of Examples 1-3, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
  • Example 5 The system of any one or more of Examples 1-4, wherein the contextual data comprises a query associated with the deviation in the metric.
  • Example 6 The system of any one or more of Examples 1-5, wherein the operations comprise: executing the query against a database to retrieve monitoring data associated with the deviation; performing a similarity search based on the monitoring data associated with the deviation; and in response to performing the similarity search, identifying the plurality of correlated metrics.
  • Example 7 The system of any one or more of Examples 1-6, wherein the operations comprise: in response to performing the similarity search, assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
  • Example 8 The system of any one or more of Examples 1-7, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  • Example 9 The system of any one or more of Examples 1-8, wherein the operation of determining the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises: applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm; generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph; ranking the plurality of scores in a descending order; and identifying a correlated metric with a top rank as the root cause of the incident.
  • Example 10 The system of any one or more of Examples 1-9, wherein the operations comprise: constructing the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics, a causal link representing a cause-and-effect relationship between a pair of correlated metrics.
  • Example 11 A method comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
  • Example 12 The method of Example 11, the generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises: mapping the plurality of correlated metrics into the plurality of categories of signals.
  • Example 13 The method of any one or more of Examples 11 or12, wherein each signal in a category of signals corresponds to a key performance indicator for monitoring and managing performance of a service.
  • Example 14 The method of any one or more of Examples 11-13, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
  • Example 15 The method of any one or more of Examples 11-14, wherein the contextual data comprises a query associated with the deviation in the metric.
  • Example 16 The method of any one or more of Examples 11-15, comprising: executing the query against a database to retrieve monitoring data associated with the deviation; performing a similarity search based on the monitoring data associated with the deviation; and in response to performing the similarity search: identifying the plurality of correlated metrics, and assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
  • Example 17 The method of any one or more of Examples 11-16, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  • Example 18 The method of any one or more of Examples 11-17, the determining of the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises: applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm; generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph; ranking the plurality of scores in a descending order; and identifying a correlated metric with a top rank as the root cause of the incident.
  • Example 19 The method of any one or more of Examples 11-18, comprising: constructing the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics, a causal link representing a cause-and-effect relationship between a pair of correlated metrics.
  • Example 20 A machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
  • FIG. 10 is a block diagram illustrating an example of a software architecture 1002 that may be installed on a machine, according to some example embodiments.
  • FIG. 10 is merely a non-limiting example of a software architecture, and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein.
  • the software architecture 1002 may be executing on hardware such as a machine 1100 of FIG. 11 that includes, among other things, processors 1110, memory 1130, and input/output (I/O) components 1150.
  • a representative hardware layer 1004 is illustrated and can represent, for example, the machine 1100 of FIG. 11.
  • the representative hardware layer 1004 comprises one or more processing units 1006 having associated executable instructions 1008.
  • the executable instructions 1008 represent the executable instructions of the software architecture 1002.
  • the hardware layer 1004 also includes memory or storage modules (storage components) 1010, which also have the executable instructions 1008.
  • the hardware layer 1004 may also comprise other hardware 1012, which represents any other hardware of the hardware layer 1004, such as the other hardware illustrated as part of the machine 1200.
  • the software architecture 1002 may be conceptualized as a stack of layers, where each layer provides particular functionality.
  • the software architecture 1002 may include layers such as an operating system 1014, libraries 1016, frameworks/middleware 1018, applications 1020, and a presentation layer 1044.
  • the applications 1020 or other components within the layers may invoke API calls 1024 through the software stack and receive a response, returned values, and so forth (illustrated as messages 1026) in response to the API calls 1024.
  • the layers illustrated are representative in nature, and not all software architectures have all layers. For example, some mobile or special-purpose operating systems may not provide a frameworks/middleware 1018 layer, while others may provide such a layer. Other software architectures may include additional or different layers.
  • the operating system 1014 may manage hardware resources and provide common services.
  • the operating system 1014 may include, for example, a kernel 1028, services 1030, and drivers 1032.
  • the kernel 1028 may act as an abstraction layer between the hardware and the other software layers.
  • the kernel 1028 may be responsible for memory management, processor management (e.g., scheduling) , component management, networking, security settings, and so on.
  • the services 1030 may provide other common services for the other software layers.
  • the drivers 1032 may be responsible for controlling or interfacing with the underlying hardware.
  • the drivers 1032 may include display drivers, camera drivers, drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers) , drivers, audio drivers, power management drivers, and so forth depending on the hardware configuration.
  • USB Universal Serial Bus
  • the libraries 1016 may provide a common infrastructure that may be utilized by the applications 1020 and/or other components and/or layers.
  • the libraries 1016 typically provide functionality that allows other software components/modules to perform tasks in an easier fashion than by interfacing directly with the underlying operating system 1014 functionality (e.g., kernel 1028, services 1030, or drivers 1032) .
  • the libraries 1016 may include system libraries 1034 (e.g., C standard library) that may provide functions such as memory allocation functions, string manipulation functions, mathematic functions, and the like.
  • the libraries 1016 may include API libraries 1036 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as MPEG4, H.
  • graphics libraries e.g., an OpenGL framework that may be used to render 2D and 3D graphic content on a display
  • database libraries e.g., SQLite that may provide various relational database functions
  • web libraries e.g., WebKit that may provide web browsing functionality
  • the libraries 1016 may also include a wide variety of other libraries 1038 to provide many other APIs to the applications 1020 and other software components/modules.
  • the frameworks 1018 may provide a higher-level common infrastructure that may be utilized by the applications 1020 or other software components/modules.
  • the frameworks 1018 may provide various graphical user interface functions, high-level resource management, high-level location services, and so forth.
  • the frameworks 1018 may provide a broad spectrum of other APIs that may be utilized by the applications 1020 and/or other software components/modules, some of which may be specific to a particular operating system or platform.
  • the applications 1020 include built-in applications 1040 and/or third-party applications 1042.
  • built-in applications 1040 may include, but are not limited to, a home application, a contacts application, a browser application, a book reader application, a location application, a media application, a messaging application, or a game application.
  • the third-party applications 1042 may include any of the built-in applications 1040, as well as a broad assortment of other applications.
  • the third-party applications 1042 e.g., an application developed using the Android TM or iOS TM software development kit (SDK) by an entity other than the vendor of the particular platform
  • the third-party applications 1042 may be mobile software running on a mobile operating system such as iOS TM , Android TM , or other mobile operating systems.
  • the third-party applications 1042 may invoke the API calls 1024 provided by the mobile operating system such as the operating system 1014 to facilitate functionality described herein.
  • the applications 1020 may utilize built-in operating system functions (e.g., kernel 1028, services 1030, or drivers 1032) , libraries (e.g., system libraries 1034, API libraries 1036, and other libraries 1038) , or frameworks/middleware 1018 to create user interfaces to interact with users of the system.
  • built-in operating system functions e.g., kernel 1028, services 1030, or drivers 1032
  • libraries e.g., system libraries 1034, API libraries 1036, and other libraries 1038
  • frameworks/middleware 1018 e.g., frameworks/middleware 1018 to create user interfaces to interact with users of the system.
  • interactions with a user may occur through a presentation layer, such as the presentation layer 1044.
  • the application/module “logic” can be separated from the aspects of the application/module that interact with the user.
  • Some software architectures utilize virtual machines. In the example of FIG. 10, this is illustrated by a virtual machine 1048.
  • the virtual machine 1048 creates a software environment where applications/modules can execute as if they were executing on a hardware machine (e.g., the machine 1100 of FIG. 11) .
  • the virtual machine 1048 is hosted by a host operating system (e.g., the operating system 1014) and typically, although not always, has a virtual machine monitor 1046, which manages the operation of the virtual machine 1048 as well as the interface with the host operating system (e.g., the operating system 1014) .
  • a software architecture executes within the virtual machine 1048, such as an operating system 1050, libraries 1052, frameworks 1054, applications 1056, or a presentation layer 1058. These layers of software architecture executing within the virtual machine 1048 can be the same as corresponding layers previously described or may be different.
  • FIG. 11 illustrates a diagrammatic representation of a machine 1100 in the form of a computer system within which a set of instructions may be executed for causing the machine 1100 to perform any one or more of the methodologies discussed herein, according to an embodiment.
  • FIG. 11 shows a diagrammatic representation of the machine 1100 in the example form of a computer system, within which instructions 1116 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1100 to perform any one or more of the methodologies discussed herein may be executed.
  • the instructions 1116 may cause the machine 1100 to execute the method 300 described above with respect to FIG. 3, the method 400 described above with respect to FIG. 4, and the method 500 described above with respect to FIG. 5.
  • the instructions 1116 transform the general, non-programmed machine 1100 into a particular machine 1100 programmed to carry out the described and illustrated functions in the manner described.
  • the machine 1100 operates as a standalone device or may be coupled (e.g., networked) to other machines.
  • the machine 1100 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
  • the machine 1100 may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC) , a tablet computer, a laptop computer, a netbook, a personal digital assistant (PDA) , an entertainment media system, a cellular telephone, a smart phone, a mobile device, or any machine capable of executing the instructions 1116, sequentially or otherwise, that specify actions to be taken by the machine 1100.
  • a server computer a client computer
  • PC personal computer
  • PDA personal digital assistant
  • an entertainment media system a cellular telephone
  • smart phone a mobile device
  • mobile device or any machine capable of executing the instructions 1116, sequentially or otherwise, that specify actions to be taken by the machine 1100.
  • the term “machine” shall also be taken to include a collection of machines 1100 that individually or jointly execute the instructions 1116 to perform any one or more of the methodologies discussed herein.
  • the machine 1100 may include processors 1110, memory 1130, and I/O components 1150, which may be configured to communicate with each other such as via a bus 1102.
  • the processors 1110 e.g., a hardware processor, such as a central processing unit (CPU) , a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU) , a digital signal processor (DSP) , an application-specific integrated circuit (ASIC) , a radio-frequency integrated circuit (RFIC) , another processor, or any suitable combination thereof
  • a hardware processor such as a central processing unit (CPU) , a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU) , a digital signal processor (DSP) , an application-specific integrated circuit (ASIC) , a radio-frequency integrated circuit (RFIC) , another processor, or any suitable combination thereof
  • CPU central processing unit
  • processor is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores” ) that may execute instructions contemporaneously.
  • FIG. 11 shows multiple processors 1110, the machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor) , multiple processors with a single core, multiple processors with multiples cores, or any combination thereof.
  • the memory 1130 may include a main memory 1132, a static memory 1134, and a storage unit 1136 including machine-readable medium 1138, each accessible to the processors 1110 such as via the bus 1102.
  • the main memory 1132, the static memory 1134, and the storage unit 1136 store the instructions 1116 embodying any one or more of the methodologies or functions described herein.
  • the instructions 1116 may also reside, completely or partially, within the main memory 1132, within the static memory 1134, within the storage unit 1136, within at least one of the processors 1110 (e.g., within the processor’s cache memory) , or any suitable combination thereof, during execution thereof by the machine 1100.
  • the I/O components 1150 may include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on.
  • the specific I/O components 1150 that are included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones will likely include a touch input device or other such input mechanisms, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I/O components 1150 may include many other components that are not shown in FIG. 11.
  • the I/O components 1150 are grouped according to functionality merely for simplifying the following discussion, and the grouping is in no way limiting. In some examples, the I/O components 1150 may include output components 1152 and input components 1154.
  • the output components 1152 may include visual components (e.g., a display such as a plasma display panel (PDP) , a light-emitting diode (LED) display, a liquid crystal display (LCD) , a projector, or a cathode ray tube (CRT) ) , acoustic components (e.g., speakers) , haptic components (e.g., a vibratory motor, resistance mechanisms) , other signal generators, and so forth.
  • visual components e.g., a display such as a plasma display panel (PDP) , a light-emitting diode (LED) display, a liquid crystal display (LCD) , a projector, or a cathode ray tube (CRT)
  • acoustic components e.g., speakers
  • haptic components e.g., a vibratory motor, resistance mechanisms
  • the input components 1154 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components) , point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument) , tactile input components (e.g., a physical button, a touch screen that provides location and/or force of touches or touch gestures, or other tactile input components) , audio input components (e.g., a microphone) , and the like.
  • alphanumeric input components e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components
  • point-based input components e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument
  • tactile input components e.
  • the I/O components 1150 may include biometric components 1156, motion components 1158, environmental components 1160, or position components 1162, among a wide array of other components.
  • the motion components 1158 may include acceleration sensor components (e.g., accelerometer) , gravitation sensor components, rotation sensor components (e.g., gyroscope) , and so forth.
  • the environmental components 1160 may include, for example, illumination sensor components (e.g., photometer) , temperature sensor components (e.g., one or more thermometers that detect ambient temperature) , humidity sensor components, pressure sensor components (e.g., barometer) , acoustic sensor components (e.g., one or more microphones that detect background noise) , proximity sensor components (e.g., infrared sensors that detect nearby objects) , gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere) , or other components that may provide indications, measurements, or signals corresponding to a surrounding physical environment.
  • illumination sensor components e.g., photometer
  • temperature sensor components e.g., one or more thermometers that detect ambient temperature
  • humidity sensor components e.g., humidity sensor components
  • pressure sensor components e.g., barometer
  • acoustic sensor components e.g., one or more microphones that detect background noise
  • proximity sensor components
  • the position components 1162 may include location sensor components (e.g., a Global Positioning System (GPS) receiver component) , altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude may be derived) , orientation sensor components (e.g., magnetometers) , and the like.
  • location sensor components e.g., a Global Positioning System (GPS) receiver component
  • altitude sensor components e.g., altimeters or barometers that detect air pressure from which altitude may be derived
  • orientation sensor components e.g., magnetometers
  • the I/O components 1150 may include communication components 1164 operable to couple the machine 1100 to a network 1180 or devices 1170 via a coupling 1182 and a coupling 1172, respectively.
  • the communication components 1164 may include a network interface component or another suitable device to interface with the network 1180.
  • the communication components 1164 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, components (e.g., Low Energy) , components, and other communication components to provide communication via other modalities.
  • the devices 1170 may be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a USB) .
  • the communication components 1164 may detect identifiers or include components operable to detect identifiers.
  • the communication components 1164 may include radio frequency identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes) , or acoustic detection components (e.g., microphones to identify tagged audio signals) .
  • RFID radio frequency identification
  • NFC smart tag detection components e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes
  • Certain embodiments are described herein as including logic or a number of components, components, elements, or mechanisms. Such components can constitute either software components (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware components.
  • a “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a certain physical manner.
  • one or more computer systems e.g., a standalone computer system, a client computer system, or a server computer system
  • one or more hardware components of a computer system e.g., a processor or a group of processors
  • software e.g., an application or application portion
  • a hardware component is implemented mechanically, electronically, or any suitable combination thereof.
  • a hardware component can include dedicated circuitry or logic that is permanently configured to perform certain operations.
  • a hardware component can be a special-purpose processor, such as a field-programmable gate array (FPGA) or an ASIC.
  • a hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations.
  • a hardware component can include software encompassed within a general-purpose processor or other programmable processor. It will be appreciated that the decision to implement a hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.
  • the phrase “component” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired) , or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein.
  • hardware components are temporarily configured (e.g., programmed)
  • each of the hardware components need not be configured or instantiated at any one instance in time.
  • a hardware component comprises a general-purpose processor configured by software to become a special-purpose processor
  • the general-purpose processor may be configured as respectively different special-purpose processors (e.g., comprising different hardware components) at different times.
  • Software can accordingly configure a particular processor or processors, for example, to constitute a particular hardware component at one instance of time and to constitute a different hardware component at a different instance of time.
  • Hardware components can provide information to, and receive information from, other hardware components. Accordingly, the described hardware components can be regarded as being communicatively coupled. Where multiple hardware components exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) between or among two or more of the hardware components. In embodiments in which multiple hardware components are configured or instantiated at different times, communications between or among such hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware components have access. For example, one hardware component performs an operation and stores the output of that operation in a memory device to which it is communicatively coupled. A further hardware component can then, at a later time, access the memory device to retrieve and process the stored output. Hardware components can also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information) .
  • a resource e.g., a collection of information
  • processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors constitute processor-implemented components that operate to perform one or more operations or functions described herein.
  • processor-implemented component refers to a hardware component implemented using one or more processors.
  • the methods described herein can be at least partially processor-implemented, with a particular processor or processors being an example of hardware.
  • at least some of the operations of a method can be performed by one or more processors or processor-implemented components.
  • the one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS) .
  • SaaS software as a service
  • at least some of the operations may be performed by a group of computers (as examples of machines 1100 including processors 1110) , with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API) .
  • a client device may relay or operate in communication with cloud computing systems and may access circuit design information in a cloud environment.
  • processors 1110 or processor-implemented components are located in a single geographic location (e.g., within a home environment, an office environment, or a server farm) . In other example embodiments, the processors or processor-implemented components are distributed across a number of geographic locations.
  • the various memories i.e., 1130, 1132, 1134, and/or the memory of the processor (s) 1110) and/or the storage unit 1136 may store one or more sets of instructions 1116 and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein.
  • These instructions e.g., the instructions 1116) , when executed by the processor (s) 1110, cause various operations to implement the disclosed embodiments.
  • machine-storage medium As used herein, the terms “machine-storage medium, ” “device-storage medium, ” and “computer-storage medium” mean the same thing and may be used interchangeably.
  • the terms refer to a single or multiple storage devices and/or media (e.g., a centralized or distributed database, and/or associated caches and servers) that store executable instructions 1116 and/or data.
  • the terms shall accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media, including memory internal or external to processors.
  • machine-storage media computer-storage media and/or device-storage media
  • non-volatile memory including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
  • semiconductor memory devices e.g., erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , FPGA, and flash memory devices
  • magnetic disks such as internal hard disks and removable disks
  • magneto-optical disks magneto-optical disks
  • CD-ROM and DVD-ROM disks CD-ROM and DVD-ROM disks.
  • machine-storage media, ” “computer-storage media, ” and “device-storage media” specifically exclude
  • one or more portions of the network 1180 may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN) , a LAN, a wireless LAN (WLAN) , a WAN, a wireless WAN (WWAN) , a metropolitan-area network (MAN) , the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN) , a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks.
  • VPN virtual private network
  • WLAN wireless LAN
  • WWAN wireless WAN
  • MAN metropolitan-area network
  • PSTN public switched telephone network
  • POTS plain old telephone service
  • the network 1180 or a portion of the network 1180 may include a wireless or cellular network
  • the coupling 1182 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or another type of cellular or wireless coupling.
  • CDMA Code Division Multiple Access
  • GSM Global System for Mobile communications
  • the coupling 1182 may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT) , Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS) , High-Speed Packet Access (HSPA) , Worldwide Interoperability for Microwave Access (WiMAX) , Long-Term Evolution (LTE) standard, others defined by various standard-setting organizations, other long-range protocols, or other data transfer technology.
  • 1xRTT Single Carrier Radio Transmission Technology
  • GPRS General Packet Radio Service
  • EDGE Enhanced Data rates for GSM Evolution
  • 3GPP Third Generation Partnership Project
  • 4G fourth generation wireless (4G) networks
  • Universal Mobile Telecommunications System (UMTS) Universal Mobile Telecommunications System
  • High-Speed Packet Access HSPA
  • the instructions may be transmitted or received over the network using a transmission medium via a network interface device (e.g., a network interface component included in the communication components) and utilizing any one of a number of well-known transfer protocols (e.g., hypertext transfer protocol (HTTP) ) .
  • a network interface device e.g., a network interface component included in the communication components
  • HTTP hypertext transfer protocol
  • the instructions may be transmitted or received using a transmission medium via the coupling (e.g., a peer-to-peer coupling) to the devices 1170.
  • the terms “transmission medium” and “signal medium” mean the same thing and may be used interchangeably in this disclosure.
  • transmission medium and “signal medium” shall be taken to include any intangible medium that is capable of storing, encoding, or carrying the instructions for execution by the machine, and include digital or analog communications signals or other intangible media to facilitate communication of such software.
  • transmission medium and “signal medium” shall be taken to include any form of modulated data signal, carrier wave, and so forth.
  • modulated data signal means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
  • machine-readable medium e.g., a non-transitory computer-readable medium
  • device-readable medium mean the same thing and may be used interchangeably in this disclosure.
  • the terms are defined to include both machine-storage media and transmission media.
  • the terms include both storage devices/media and carrier waves/modulated data signals.
  • an embodiment described herein can be implemented using a non-transitory medium (e.g., a non-transitory computer-readable medium) .
  • the term “or” may be construed in either an inclusive or exclusive sense.
  • the terms “a” or “an” should be read as meaning “at least one, ” “one or more, ” or the like.
  • the presence of broadening words and phrases such as “one or more, ” “at least, ” “but not limited to, ” or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent.
  • boundaries between various resources, operations, components, engines, and data stores are somewhat arbitrary, and particular operations are illustrated in a context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within a scope of various embodiments of the present disclosure.
  • the specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Quality & Reliability (AREA)
  • General Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • Computer Hardware Design (AREA)
  • Mathematical Physics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Probability & Statistics with Applications (AREA)
  • Health & Medical Sciences (AREA)
  • Biomedical Technology (AREA)
  • Debugging And Monitoring (AREA)

Abstract

An incident alert associated with a deviation in a metric is received. A plurality of correlated metrics is identified based on contextual data. A plurality of labels for the plurality of correlated metrics is generated in accordance with a plurality of categories of signals. A causal graph is constructed based on the plurality of labels and a topology graph of the plurality of correlated metrics. A root cause of an incident associated with the incident alert is determined using a link analysis algorithm.

Description

INCIDENT TRIAGE AND ROOT CAUSE ANALYSIS TECHNICAL FIELD
The present disclosure generally relates to cloud computing technologies. More particularly, various embodiments described herein provide for systems, methods, techniques, instruction sequences, and devices that facilitate automated incident triage and root cause analysis in complex cloud computing environments.
BACKGROUND
In the realm of cloud computing, managing and maintaining the stability and efficiency of operations is a continuous challenge. This sector typically deals with complex environments composed of numerous interconnected services and infrastructure components. During operational disruptions, such as site incidents, the process of identifying and resolving these issues can be time-consuming. This is exacerbated by the vast amount of data generated by various monitoring tools, such as metrics from thousands of microservices. The ability to quickly and accurately analyze this data to determine the root cause of an incident is crucial for minimizing downtime and maintaining service reliability.
BRIEF DESCRIPTION OF THE DRAWINGS
In the drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views. To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced. Some embodiments are illustrated by way of examples, and not limitations, in the accompanying figures.
FIG. 1 is a block diagram showing an example data system that includes a data management system, according to various embodiments of the present disclosure.
FIG. 2 is a block diagram illustrating an example data management system that facilitates incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
FIG. 3 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
FIG. 4 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
FIG. 5 is a flowchart illustrating an example method for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
FIG. 6 is a diagram illustrating an example data management system that facilitates incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure.
FIG. 7 is a diagram illustrating an example topology graph and an example causal graph, according to various embodiments of the present disclosure.
FIG. 8 is a diagram illustrating an example data table that stores causal links derived from historical data and/or human experience, according to various embodiments of the present disclosure.
FIG. 9 is a diagram illustrating line graphs representing monitoring data associated with business metrics and correlated metrics, according to various embodiments of the present disclosure.
FIG. 10 is a block diagram illustrating a representative software architecture, which may be used in conjunction with various hardware architectures herein described, according to various embodiments of the present disclosure.
FIG. 11 is a block diagram illustrating components of a machine able to read instructions from a machine storage medium and perform any one or more of the  methodologies discussed herein according to various embodiments of the present disclosure.
DETAILED DESCRIPTION
The description that follows includes systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments. It will be evident, however, to one skilled in the art that the present inventive subject matter may be practiced without these specific details.
Reference in the specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present subject matter. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” appearing in various places throughout the specification are not necessarily all referring to the same embodiment.
For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the present subject matter. However, it will be apparent to one of ordinary skill in the art that embodiments of the subject matter described may be practiced without the specific details presented herein, or in various combinations, as described herein. Furthermore, well-known features may be omitted or simplified in order not to obscure the described embodiments. Various embodiments may be given throughout this description. These are merely descriptions of specific embodiments. The scope or meaning of the claims is not limited to the embodiments given.
Various embodiments include systems, methods, and non-transitory computer-readable media that facilitate automated incident triage and root cause analysis in complex cloud computing environments, according to various embodiments of the present disclosure. Incident triage involves assessing incidents  to determine their severity, urgency, and priority for resolution. An incident can refer to any event or occurrence that disrupts the normal operation of services and requires a response to mitigate its impact. Incident triage helps prioritize resources and responses effectively to minimize the impact of incidents on users and the business. Root cause analysis involves identifying the underlying causes of incidents. Root cause analysis not only addresses the symptoms or immediate causes of an incident but also provides insights into the underlying systemic issues and/or weaknesses in processes, systems, or infrastructure that need to be addressed to prevent recurrence.
Various embodiments involve a data management system that receives an alert of an incident as input. The alert includes contextual information about deviating metrics associated with the incident. A metric refers to a quantitative measurement that provides insight into various aspects of a service's behavior. A metric can be represented by time series data. Based on the contextual information, the data management system performs a time series similarity search to identify the monitoring data of the incident and filters out noise by focusing on the residuals of the time series, considering factors like offset and direction. This refined data is then mapped into the predefined categories of signals (e.g., traffic, error rate, saturation, and latency) , which are key performance indicators of services (e.g., microservices) provided by a cloud computing environment. Mapped metrics are overlaid onto a topological graph that represents services and the associated infrastructure components (e.g., service mesh) , providing a comprehensive view of the cloud environment. Further, the data management system constructs a causal graph (e.g., signal causal graph) based on this topological representation and causal links derived from historical data and human experience. The application of a weighted link analysis algorithm (e.g., a weighted PageRank algorithm) on this graph, where weights are derived from the similarity scores of the time series analysis, allows for the precise determination of the root cause.
This approach addresses the complexity of diagnosing incidents in cloud environments with a high degree of automation and accuracy. Unlike prior  solutions, such as basic alert systems or manual triage processes that do not integrate deep analytical techniques into their workflows, various embodiments provide a systematic, automated, and data-driven way to quickly pinpoint root causes. This reduces the Mean Time to Recovery (MTTR) and improves the reliability of cloud services, which is critical in maintaining operational efficiency and customer trust in cloud platforms.
In various embodiments, a data management system receives an incident alert associated with a deviation in a metric. Examples of deviation can include a spike, a dip, a fluctuation, or a pattern, indicating a departure from an expected behavior of a metric. An incident can refer to any event (e.g., anomaly event) or occurrence that disrupts the normal operation of a service (e.g., microservice) and requires a response to mitigate its impact. For example, a payment failure event can be an incident associated with the payment processing service. In various embodiments, an alert associated with an incident (also referred to as an incident alert) includes contextual data representing the deviation. The contextual data can include a query (e.g., raw query) associated with the deviation in the metric.
In various embodiments, the data management system identifies a plurality of correlated metrics based on the contextual data (e.g., query) included in the incident alert. Specifically, the data management system can execute the query against a database to retrieve monitoring data associated with the deviation. The data management system then performs a similarity search (e.g., a time series similarity search) based on the retrieved monitoring data to identify the plurality of correlated metrics. In response to performing the similarity search, the data management system identifies the plurality of correlated metrics and assigns a similarity value to each of the plurality of correlated metrics. Each similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation. In various embodiments, each similarity value is associated with a direction indicator corresponding to a deviation type. Examples of deviation types include spikes, surges, dips, drops, fluctuations, and patterns.
Example algorithms (or measures) associated with the similarity search include, without limitation, Minkowski distance measure, cosine similarity measure, Pearson correlation coefficient measure, Dynamic Time Warping (DTW) measure, kernel-based similarity measure, longest common subsequence algorithm, slope-based similarity, and k-shape algorithm.
In various embodiments, the queried database stores time series monitoring data using a custom format optimized for efficient querying and storage. In the database, monitoring data can be organized as time series, where each time series represents a stream of timestamped metric values associated with a specific metric and set of labels.
In various embodiments, the data management system generates a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals. In particular, the operation of generating the plurality of labels can include mapping the plurality of correlated metrics into the plurality of categories of signals. Each signal in a category of signals (e.g., traffic, error rate, saturation, and latency) can correspond to a key performance indicator for monitoring and managing the performance of a service in the cloud computing platform.
In various embodiments, the data management system constructs a causal graph (e.g., a signal causal graph) based on the plurality of labels and a topology graph representing the plurality of correlated metrics.
An example causal graph, also referred to as a causal inference graph, can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) . A causal graph can specifically focus on the causal relationships between signals or time series data.
An example topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate the flow of requests between them. Infrastructure components, such as service mesh, sit between microservices and handle service-to-service communication, providing features like  load balancing, service discovery, and security. Each microservice or infrastructure component can include (or correspond to) one or more metrics.
In various embodiments, the data management system uses a link analysis algorithm (e.g., weighted PageRank algorithm) to determine the root cause of an incident. The determination is at least based on the constructed causal graph and a plurality of weights assigned to the plurality of correlated metrics. The root cause of the incident can correspond to a correlated metric from the plurality of correlated metrics. In particular, the data management system applies the plurality of similarity values assigned to the plurality of correlated metrics as a plurality of weights to the link analysis algorithm. Applying weights to the weighted link analysis algorithm involves assigning numerical values to the links between nodes in a network. These weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes. The data management system uses the weighted link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph. The scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence. The data management system then ranks the plurality of scores in descending order and identifies the metric with the top rank (i.e., the metric with the highest ranking) as the root cause of the incident.
In various embodiments, the data management system constructs the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics. Tabular data can be represented by a causal link table, which can be updated based on the outcomes of previous incident triage. A causal link represents a cause-and-effect relationship between a pair of correlated metrics. The plurality of causal links can be derived from historical data and/or human experience.
Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the appended drawings. The present disclosure may, however, be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein.
FIG. 1 is a block diagram showing an example data system 100 that includes a data management system 122 (also referred to as system 122) , according to various embodiments of the present disclosure. By including the data management system 122, the data system 100 can facilitate incident triage and root cause analysis in cloud computing environments. As shown, the data system 100 includes one or more client devices 102, a server system 108, and a network 106 (e.g., Internet, wide-area-network (WAN) , local-area-network (LAN) , wireless network) that communicatively couples them together. Each client device 102 can host a number of applications, including a client software application 104. The client software application 104 can communicate data with the server system 108 via a network 106. Accordingly, the client software application 104 can communicate and exchange data with the server system 108 via network 106.
The server system 108 provides server-side functionality via the network 106 to the client software application 104. While certain functions of the data system 100 are described herein as being performed by the data management system 122 on the server system 108, it will be appreciated that the location of certain functionality within the server system 108 is a design choice. For example, it may be technically preferable to initially deploy certain technology and functionality within the server system 108, but to later migrate this technology and functionality to the client software application 104.
The server system 108 supports various services and operations that are provided to the client software application 104 by the data management system 122. Such operations include transmitting data from the data management system 122 to the client software application 104, receiving data from the client software application 104 at the data management system 122, and the data management system 122 processing data generated by the client software application 104. Data exchanges within the data system 100 may be invoked and controlled through operations of software component environments available via one or more endpoints, or functions available via one or more user interfaces of the client  software application 104, which may include web-based user interfaces provided by the server system 108 for presentation at the client device 102.
With respect to the server system 108, an Application Program Interface (API) server 110 and a web server 112 is coupled to an application server 116, which hosts the data management system 122. The application server 116 is communicatively coupled to a database server 118, which facilitates access to a database 120 that stores data associated with the application server 116, including data that may be generated or used by the data management system 122.
The API server 110 receives and transmits data (e.g., API calls, commands, requests, responses, and authentication data) between the client device 102 and the application server 116. Specifically, the API server 110 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the client software application 104 in order to invoke the functionality of the application server 116. The API server 110 exposes various functions supported by the application server 116 including, without limitation, user registration; login functionality; data object operations (e.g., generating, storing, retrieving, encrypting, decrypting, transferring, access rights, licensing) ; and/or user communications.
The server system 108, or the data management system 122 may extract user data from one or more third-party platforms (e.g., third-party social media platforms) . The extracted data may be open-source poster data associated with targeted influencers on the one or more third-party platforms 124 and may include user profile data, activity data, and media posted (either created and/or shared) by the one or more influencers. The media (or media data) include text, image, video, audio, and metadata. Example metadata may include hashtags and labels.
Through one or more web-based interfaces (e.g., web-based user interfaces) , the web server 112 can support various functionality of the data management system 122 of the application server 116.
FIG. 2 is a block diagram illustrating an example data management system 200 that facilitates incident triage and root cause analysis in cloud computing  environments, according to various embodiments of the present disclosure. For some embodiments, the data management system 200 represents an example of the data management system 122 described with respect to FIG. 1. As shown, the data management system 200 comprises an incident alert receiving component 210, a correlated metric identifying component 220, a metric categorizing component 230, a causal graph constructing component 240, and a root cause determining component 250. According to various embodiments, one or more of the incident alert receiving component 210, the correlated metric identifying component 220, the metric categorizing component 230, the causal graph constructing component 240, and the root cause determining component 250 are implemented by one or more hardware processors 202. Data generated by one or more of the incident alert receiving component 210, the correlated metric identifying component 220, the metric categorizing component 230, the causal graph constructing component 240, and the root cause determining component 250 may be stored in a database (or datastore) 260 of the data management system 200.
The incident alert receiving component 210 is configured to receive incident alerts generated based on the detection of incidents as input. An incident alert can include contextual information (also referred to as contextual data) on the deviating metric. Contextual data can include a query (e.g., raw query) associated with the deviating metric.
The correlated metric identifying component 220 is configured to identify a plurality of correlated metrics based on the contextual data (e.g., query) included in the incident alert. In particular, the correlated metric identifying component 220 is configured to perform a similarity search (e.g., a time series similarity search) based on the monitoring data retrieved in the query results to identify the correlated metrics. Each correlated metric is assigned a similarity value representing a degree of similarity between the correlated metric and the deviating metric. Each similarity value can be associated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
The metric categorizing component 230 is configured to generate a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals. In particular, the metric categorizing component 230 is configured to map the plurality of correlated metrics into the plurality of categories of signals. Each signal in a category of signals (e.g., traffic, error rate, saturation, and latency) corresponds to a key performance indicator for monitoring and managing the performance of a service in the cloud computing platform.
The causal graph constructing component 240 is configured to construct a causal graph (e.g., a signal causal graph) based on the plurality of labels and a topology graph representing the plurality of correlated metrics. A causal graph can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) . A topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate invocation and dependency relationships between them.
The root cause determining component 250 is configured to use a link analysis algorithm (e.g., weighted PageRank algorithm) to determine the root cause of an incident. The determination is at least based on the constructed causal graph and a plurality of weights assigned to the plurality of correlated metrics. The root cause of the incident corresponds to a correlated metric from the plurality of correlated metrics. In particular, the root cause determining component 250 is configured to apply the similarity values as weights to the link analysis algorithm. These weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes. The root cause determining component 250 is configured to use the weighted link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph. The scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence. The plurality of scores is ranked in descending order. The metric with the top rank (i.e., the metric with the highest ranking) is determined to be the root cause of the incident.
FIG. 3 is a flowchart illustrating an example method 300 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 300 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 300 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 300. Depending on the embodiment, an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel.
At operation 302, a processor receives an incident alert associated with a deviation in a metric. The incident alert includes contextual data representing the deviation.
At operation 304, a processor identifies a plurality of correlated metrics based on the contextual data. Contextual data can include a query (e.g., raw query) associated with the deviating metric.
At operation 306, a processor generates a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals. In particular, a processor maps the plurality of correlated metrics into the plurality of  categories of signals. Each signal in a category of signals (e.g., traffic, error rate, saturation, and latency) corresponds to a key performance indicator for monitoring and managing the performance of a service in the cloud computing platform.
At operation 308, a processor constructs a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics. A causal graph can be a graphical representation used in causal inference to illustrate the causal relationships between variables (e.g., metrics) . A topology graph can include a plurality of microservices and infrastructure components connected by arrows that indicate the flow of requests between them.
At operation 310, a processor uses a link analysis algorithm (e.g., a weighted PageRank algorithm) to determine the root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics. The determination of the root cause is further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics. A causal link represents a cause-and-effect relationship between a pair of correlated metrics. The root cause of the incident corresponds to one of the plurality of correlated metrics.
Though not illustrated, method 300 can include an operation where a graphical user interface is displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 302 through 310 or, alternatively, form part of one or more of operations 302 through 310.
FIG. 4 is a flowchart illustrating an example method 400 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 400 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200  described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 400 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 400. Depending on the embodiment, an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 400 can be performed dependently or independently from operations in method 300.
At operation 402, a processor retrieves the monitoring data associated with the deviation by executing a query against a database. The query is included in the contextual data described herein. The queried database stores time series monitoring data using a custom format optimized for efficient querying and storage. In the database, monitoring data can be organized as time series, where each time series represents a stream of timestamped metric values associated with a specific metric and set of labels.
At operation 404, a processor performs a similarity search (e.g., a time series similarity search) based on the monitoring data associated with the deviation. In response to performing the search, each correlated metric is assigned a similarity value representing a degree of similarity between the correlated metric and the deviating metric. Each similarity value can be associated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
At operation 406, a processor identifies a plurality of correlated metrics based on the similarity search.
At operation 408, a processor assigns a similarity value to each of the plurality of correlated metrics. A similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation. Each similarity value can be generated with a direction indicator corresponding to a deviation type, such as a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
Though not illustrated, method 400 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 402 through 408 or, alternatively, form part of one or more of operations 402 through 408.
FIG. 5 is a flowchart illustrating an example method 500 for facilitating incident triage and root cause analysis in cloud computing environments, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 500 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 500 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 500. Depending on the embodiment, an  operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 500 can be performed dependently or independently from operations in method 300 and method 400.
At operation 502, a processor applies a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm. The weights influence the calculation of scores and can reflect the importance, relevance, or quality of the connections between nodes.
At operation 504, a processor uses the link analysis algorithm to generate a plurality of scores for the plurality of correlated metrics based on the causal graph and the plurality of similarity values described herein. Each similarity value represents a degree of similarity between a correlated metric and the metric associated with the deviation. The scores indicate the relative importance of each correlated metric. Higher scores suggest greater importance or influence.
At operation 506, a processor ranks the plurality of scores in descending order.
At operation 508, a processor identifies a correlated metric with a top rank as the root cause of the incident.
Though not illustrated, method 500 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 502 through 508 or, alternatively, form part of one or more of operations 502 through 508.
FIG. 6 is a diagram illustrating an example data management system 600 that facilitates incident triage and root cause analysis in cloud computing  environments, according to various embodiments of the present disclosure. As shown, Alerts 608 is generated based on anomaly detection 606. Anomaly detection involves generating incident alerts in response to detecting anomaly events, such as a payment failure, or a shipping label failure. In various embodiments, alerts 608 can be generated after a similarity search 604 is performed. Causal graphs 614 are constructed before causal inferencing 616, where a weighted link analysis algorithm is used to determine the root cause of an incident. Root causes can be provided as the basis for generating actionable insights 612. A topology graph 610 can be generated based on data from tracing/log 622 and data from CMDB (Configuration Management Database) 624. CMDB data can include details about hardware, software, network devices, applications, and their relationships and dependencies in an organization’s IT infrastructure. Tracing and log data include monitoring data that records invocation and dependency relationships between microservices.
FIG. 7 is a diagram 700 illustrating an example topology graph and an example causal graph, according to various embodiments of the present disclosure. As shown, an example topology graph includes microservice S1, microservice S2, microservice S3, microservice S4. A topology graph shows invocation and dependency relationships between microservices and the associated infrastructure components (not shown) . Each microservice can be associated with (or include) one or more metrics used to assess the services' effectiveness, efficiency, and quality. As shown, microservice S0 is associated with metric 702. Microservice S1 is associated with metrics 704 and 706. Microservice S2 is associated with metrics 708, 710, and 712. Microservice S3 is associated with metric 714. The numbers and types of metrics needed for a given service are determined based on various considerations, including the nature of the service (e.g., the level of availability, required response time, data processing reliability) and business needs.
Each metric is assigned a label and a direction representing a deviation. As shown, the label “ERROR” and the direction “+” are generated for the metric 702. Microservice S1 includes metric 704 and metric 706. Metric 704 is also assigned the label “ERROR” and the direction “+. ” Metric 706 is assigned the label  “LATENCY” and the direction “+. ” Direction “+” represents a spike or a surge in deviation of a metric, whereas direction “-” represents a dip or a drop.
As shown, an example causal graph includes metrics 702, 704, 706, 708, 710, 712, and 714, connected by arrows marked with “CAUSE. ” For example, a saturation in metric 708 causes a latency in metric 706, which in turn causes an error in metric 704.
FIG. 8 is a diagram illustrating an example data table 800 that stores causal links derived from historical data and/or human experience, according to various embodiments of the present disclosure. As shown, row 802 includes a causal link between the metric “Downstream Service” and the metric “Current Service, ” where an error in the metric “Downstream Service” causes an error in the metric “Current Service. ” Causal links can be stored in a data table and updated periodically based on the outcomes of previous incident triage, for example.
FIG. 9 is a diagram 900 illustrating line graphs representing monitoring data associated with business metrics and correlated metrics, according to various embodiments of the present disclosure. A time series similarity search is performed to identify the monitoring data of the incident and filter out noise by focusing on the residuals of the time series, considering factors like offset and direction. As shown, line graph 902 represents monitoring data associated with a business metric between 3: 10 am and 6: 10 am on 3/25. Surge 910 is observed at 6: 10 am and is determined to be an incident (e.g., Shipping Label Failure) that triggers an incident alert. Line graph 906 represents monitoring data associated with a correlated metric identified as the result of a similarity search described herein. Line graph 906 illustrates that surge 912 occurred at 6: 12 am, around the same time surge 910 is observed. The deviations in both line graphs (i.e., 902 and 906) are in the same direction.
Line graphs 904 and 908 are illustrated to show metrics with deviation in opposite directions that can also be determined as similar (or correlated) based on a similarity search described herein. As shown, line graph 904 represents monitoring data associated with a business metric between 12: 38 pm and 3: 38 pm on 4/16. Dip 914 in line graph 904 is observed at 3: 38 pm. Line graph 908 represents monitoring  data associated with a correlated metric identified as the result of a similarity search. As shown in line graph 908, surge 916 is observed at the same time as dip 914. The surge 916 and the dip 914 are deviations in opposite directions.
Described implementations of the subject matter can include one or more features, alone or in combination as illustrated below by way of examples.
Example 1. A system comprising: one or more hardware processors; and at least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
Example 2. The system of Example 1, wherein the operation of generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises: mapping the plurality of correlated metrics into the plurality of categories of signals.
Example 3. The system of any one or more of Examples 1 or 2, wherein each signal in a category of signals corresponds to a key performance indicator for monitoring and managing performance of a service.
Example 4. The system of any one or more of Examples 1-3, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
Example 5. The system of any one or more of Examples 1-4, wherein the contextual data comprises a query associated with the deviation in the metric.
Example 6. The system of any one or more of Examples 1-5, wherein the operations comprise: executing the query against a database to retrieve monitoring data associated with the deviation; performing a similarity search based on the monitoring data associated with the deviation; and in response to performing the similarity search, identifying the plurality of correlated metrics.
Example 7. The system of any one or more of Examples 1-6, wherein the operations comprise: in response to performing the similarity search, assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
Example 8. The system of any one or more of Examples 1-7, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
Example 9. The system of any one or more of Examples 1-8, wherein the operation of determining the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises: applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm; generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph; ranking the plurality of scores in a descending order; and identifying a correlated metric with a top rank as the root cause of the incident.
Example 10. The system of any one or more of Examples 1-9, wherein the operations comprise: constructing the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics, a  causal link representing a cause-and-effect relationship between a pair of correlated metrics.
Example 11. A method comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
Example 12. The method of Example 11, the generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises: mapping the plurality of correlated metrics into the plurality of categories of signals.
Example 13. The method of any one or more of Examples 11 or12, wherein each signal in a category of signals corresponds to a key performance indicator for monitoring and managing performance of a service.
Example 14. The method of any one or more of Examples 11-13, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
Example 15. The method of any one or more of Examples 11-14, wherein the contextual data comprises a query associated with the deviation in the metric.
Example 16. The method of any one or more of Examples 11-15, comprising: executing the query against a database to retrieve monitoring data associated with the deviation; performing a similarity search based on the  monitoring data associated with the deviation; and in response to performing the similarity search: identifying the plurality of correlated metrics, and assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
Example 17. The method of any one or more of Examples 11-16, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
Example 18. The method of any one or more of Examples 11-17, the determining of the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises: applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm; generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph; ranking the plurality of scores in a descending order; and identifying a correlated metric with a top rank as the root cause of the incident.
Example 19. The method of any one or more of Examples 11-18, comprising: constructing the causal graph further based on tabular data that includes a plurality of causal links between the plurality of correlated metrics, a causal link representing a cause-and-effect relationship between a pair of correlated metrics.
Example 20. A machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: receiving an incident alert associated with a deviation in a metric, the incident alert including contextual data representing the deviation; identifying a plurality of correlated metrics based on the contextual data; generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals; constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of  correlated metrics; and determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
FIG. 10 is a block diagram illustrating an example of a software architecture 1002 that may be installed on a machine, according to some example embodiments. FIG. 10 is merely a non-limiting example of a software architecture, and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 1002 may be executing on hardware such as a machine 1100 of FIG. 11 that includes, among other things, processors 1110, memory 1130, and input/output (I/O) components 1150. A representative hardware layer 1004 is illustrated and can represent, for example, the machine 1100 of FIG. 11. The representative hardware layer 1004 comprises one or more processing units 1006 having associated executable instructions 1008. The executable instructions 1008 represent the executable instructions of the software architecture 1002. The hardware layer 1004 also includes memory or storage modules (storage components) 1010, which also have the executable instructions 1008. The hardware layer 1004 may also comprise other hardware 1012, which represents any other hardware of the hardware layer 1004, such as the other hardware illustrated as part of the machine 1200.
In the example architecture of FIG. 10, the software architecture 1002 may be conceptualized as a stack of layers, where each layer provides particular functionality. For example, the software architecture 1002 may include layers such as an operating system 1014, libraries 1016, frameworks/middleware 1018, applications 1020, and a presentation layer 1044. Operationally, the applications 1020 or other components within the layers may invoke API calls 1024 through the software stack and receive a response, returned values, and so forth (illustrated as messages 1026) in response to the API calls 1024. The layers illustrated are representative in nature, and not all software architectures have all layers. For  example, some mobile or special-purpose operating systems may not provide a frameworks/middleware 1018 layer, while others may provide such a layer. Other software architectures may include additional or different layers.
The operating system 1014 may manage hardware resources and provide common services. The operating system 1014 may include, for example, a kernel 1028, services 1030, and drivers 1032. The kernel 1028 may act as an abstraction layer between the hardware and the other software layers. For example, the kernel 1028 may be responsible for memory management, processor management (e.g., scheduling) , component management, networking, security settings, and so on. The services 1030 may provide other common services for the other software layers. The drivers 1032 may be responsible for controlling or interfacing with the underlying hardware. For instance, the drivers 1032 may include display drivers, camera drivers, drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers) , drivers, audio drivers, power management drivers, and so forth depending on the hardware configuration.
The libraries 1016 may provide a common infrastructure that may be utilized by the applications 1020 and/or other components and/or layers. The libraries 1016 typically provide functionality that allows other software components/modules to perform tasks in an easier fashion than by interfacing directly with the underlying operating system 1014 functionality (e.g., kernel 1028, services 1030, or drivers 1032) . The libraries 1016 may include system libraries 1034 (e.g., C standard library) that may provide functions such as memory allocation functions, string manipulation functions, mathematic functions, and the like. In addition, the libraries 1016 may include API libraries 1036 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as MPEG4, H. 264, MP3, AAC, AMR, JPG, and PNG) , graphics libraries (e.g., an OpenGL framework that may be used to render 2D and 3D graphic content on a display) , database libraries (e.g., SQLite that may provide various relational database functions) , web libraries (e.g., WebKit that may provide web browsing functionality) , and the like. The libraries 1016 may also include a wide variety of  other libraries 1038 to provide many other APIs to the applications 1020 and other software components/modules.
The frameworks 1018 (also sometimes referred to as middleware) may provide a higher-level common infrastructure that may be utilized by the applications 1020 or other software components/modules. For example, the frameworks 1018 may provide various graphical user interface functions, high-level resource management, high-level location services, and so forth. The frameworks 1018 may provide a broad spectrum of other APIs that may be utilized by the applications 1020 and/or other software components/modules, some of which may be specific to a particular operating system or platform.
The applications 1020 include built-in applications 1040 and/or third-party applications 1042. Examples of representative built-in applications 1040 may include, but are not limited to, a home application, a contacts application, a browser application, a book reader application, a location application, a media application, a messaging application, or a game application.
The third-party applications 1042 may include any of the built-in applications 1040, as well as a broad assortment of other applications. In a specific example, the third-party applications 1042 (e.g., an application developed using the AndroidTM or iOSTM software development kit (SDK) by an entity other than the vendor of the particular platform) may be mobile software running on a mobile operating system such as iOSTM, AndroidTM, or other mobile operating systems. In this example, the third-party applications 1042 may invoke the API calls 1024 provided by the mobile operating system such as the operating system 1014 to facilitate functionality described herein.
The applications 1020 may utilize built-in operating system functions (e.g., kernel 1028, services 1030, or drivers 1032) , libraries (e.g., system libraries 1034, API libraries 1036, and other libraries 1038) , or frameworks/middleware 1018 to create user interfaces to interact with users of the system. Alternatively, or additionally, in some systems, interactions with a user may occur through a presentation layer, such as the presentation layer 1044. In these systems, the  application/module “logic” can be separated from the aspects of the application/module that interact with the user.
Some software architectures utilize virtual machines. In the example of FIG. 10, this is illustrated by a virtual machine 1048. The virtual machine 1048 creates a software environment where applications/modules can execute as if they were executing on a hardware machine (e.g., the machine 1100 of FIG. 11) . The virtual machine 1048 is hosted by a host operating system (e.g., the operating system 1014) and typically, although not always, has a virtual machine monitor 1046, which manages the operation of the virtual machine 1048 as well as the interface with the host operating system (e.g., the operating system 1014) . A software architecture executes within the virtual machine 1048, such as an operating system 1050, libraries 1052, frameworks 1054, applications 1056, or a presentation layer 1058. These layers of software architecture executing within the virtual machine 1048 can be the same as corresponding layers previously described or may be different.
FIG. 11 illustrates a diagrammatic representation of a machine 1100 in the form of a computer system within which a set of instructions may be executed for causing the machine 1100 to perform any one or more of the methodologies discussed herein, according to an embodiment. Specifically, FIG. 11 shows a diagrammatic representation of the machine 1100 in the example form of a computer system, within which instructions 1116 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1100 to perform any one or more of the methodologies discussed herein may be executed. For example, the instructions 1116 may cause the machine 1100 to execute the method 300 described above with respect to FIG. 3, the method 400 described above with respect to FIG. 4, and the method 500 described above with respect to FIG. 5. The instructions 1116 transform the general, non-programmed machine 1100 into a particular machine 1100 programmed to carry out the described and illustrated functions in the manner described. In alternative embodiments, the machine 1100 operates as a standalone device or may be coupled (e.g., networked)  to other machines. In a networked deployment, the machine 1100 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1100 may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC) , a tablet computer, a laptop computer, a netbook, a personal digital assistant (PDA) , an entertainment media system, a cellular telephone, a smart phone, a mobile device, or any machine capable of executing the instructions 1116, sequentially or otherwise, that specify actions to be taken by the machine 1100. Further, while only a single machine 1100 is illustrated, the term “machine” shall also be taken to include a collection of machines 1100 that individually or jointly execute the instructions 1116 to perform any one or more of the methodologies discussed herein.
The machine 1100 may include processors 1110, memory 1130, and I/O components 1150, which may be configured to communicate with each other such as via a bus 1102. In an embodiment, the processors 1110 (e.g., a hardware processor, such as a central processing unit (CPU) , a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU) , a digital signal processor (DSP) , an application-specific integrated circuit (ASIC) , a radio-frequency integrated circuit (RFIC) , another processor, or any suitable combination thereof) may include, for example, a processor 1112 and a processor 1114 that may execute the instructions 1116. The term “processor” is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores” ) that may execute instructions contemporaneously. Although FIG. 11 shows multiple processors 1110, the machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor) , multiple processors with a single core, multiple processors with multiples cores, or any combination thereof.
The memory 1130 may include a main memory 1132, a static memory 1134, and a storage unit 1136 including machine-readable medium 1138,  each accessible to the processors 1110 such as via the bus 1102. The main memory 1132, the static memory 1134, and the storage unit 1136 store the instructions 1116 embodying any one or more of the methodologies or functions described herein. The instructions 1116 may also reside, completely or partially, within the main memory 1132, within the static memory 1134, within the storage unit 1136, within at least one of the processors 1110 (e.g., within the processor’s cache memory) , or any suitable combination thereof, during execution thereof by the machine 1100.
The I/O components 1150 may include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I/O components 1150 that are included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones will likely include a touch input device or other such input mechanisms, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I/O components 1150 may include many other components that are not shown in FIG. 11. The I/O components 1150 are grouped according to functionality merely for simplifying the following discussion, and the grouping is in no way limiting. In some examples, the I/O components 1150 may include output components 1152 and input components 1154. The output components 1152 may include visual components (e.g., a display such as a plasma display panel (PDP) , a light-emitting diode (LED) display, a liquid crystal display (LCD) , a projector, or a cathode ray tube (CRT) ) , acoustic components (e.g., speakers) , haptic components (e.g., a vibratory motor, resistance mechanisms) , other signal generators, and so forth. The input components 1154 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components) , point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument) , tactile input components (e.g., a physical button, a touch screen that provides location and/or force of touches or touch gestures, or other tactile input components) , audio input components (e.g., a microphone) , and the like.
In further embodiments, the I/O components 1150 may include biometric components 1156, motion components 1158, environmental components 1160, or position components 1162, among a wide array of other components. The motion components 1158 may include acceleration sensor components (e.g., accelerometer) , gravitation sensor components, rotation sensor components (e.g., gyroscope) , and so forth. The environmental components 1160 may include, for example, illumination sensor components (e.g., photometer) , temperature sensor components (e.g., one or more thermometers that detect ambient temperature) , humidity sensor components, pressure sensor components (e.g., barometer) , acoustic sensor components (e.g., one or more microphones that detect background noise) , proximity sensor components (e.g., infrared sensors that detect nearby objects) , gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere) , or other components that may provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 1162 may include location sensor components (e.g., a Global Positioning System (GPS) receiver component) , altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude may be derived) , orientation sensor components (e.g., magnetometers) , and the like.
Communication may be implemented using a wide variety of technologies. The I/O components 1150 may include communication components 1164 operable to couple the machine 1100 to a network 1180 or devices 1170 via a coupling 1182 and a coupling 1172, respectively. For example, the communication components 1164 may include a network interface component or another suitable device to interface with the network 1180. In further examples, the communication components 1164 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, components (e.g., Low Energy) , components, and other communication components to provide communication via other modalities. The devices 1170 may be another machine or  any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a USB) .
Moreover, the communication components 1164 may detect identifiers or include components operable to detect identifiers. For example, the communication components 1164 may include radio frequency identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes) , or acoustic detection components (e.g., microphones to identify tagged audio signals) . In addition, a variety of information may be derived via the communication components 1164, such as location via Internet Protocol (IP) geolocation, location viasignal triangulation, location via detecting an NFC beacon signal that may indicate a particular location, and so forth.
Certain embodiments are described herein as including logic or a number of components, components, elements, or mechanisms. Such components can constitute either software components (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a certain physical manner. In various example embodiments, one or more computer systems (e.g., a standalone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) are configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.
In some examples, a hardware component is implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic that is permanently configured to perform certain operations. For example, a hardware component can  be a special-purpose processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component can include software encompassed within a general-purpose processor or other programmable processor. It will be appreciated that the decision to implement a hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.
Accordingly, the phrase “component” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired) , or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware components are temporarily configured (e.g., programmed) , each of the hardware components need not be configured or instantiated at any one instance in time. For example, where a hardware component comprises a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as respectively different special-purpose processors (e.g., comprising different hardware components) at different times. Software can accordingly configure a particular processor or processors, for example, to constitute a particular hardware component at one instance of time and to constitute a different hardware component at a different instance of time.
Hardware components can provide information to, and receive information from, other hardware components. Accordingly, the described hardware components can be regarded as being communicatively coupled. Where multiple hardware components exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) between or among two or more of the hardware components. In embodiments in which multiple hardware components are configured or instantiated at different times, communications between or among such hardware components may be achieved,  for example, through the storage and retrieval of information in memory structures to which the multiple hardware components have access. For example, one hardware component performs an operation and stores the output of that operation in a memory device to which it is communicatively coupled. A further hardware component can then, at a later time, access the memory device to retrieve and process the stored output. Hardware components can also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information) .
The various operations of example methods described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, “processor-implemented component” refers to a hardware component implemented using one or more processors.
Similarly, the methods described herein can be at least partially processor-implemented, with a particular processor or processors being an example of hardware. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented components. Moreover, the one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS) . For example, at least some of the operations may be performed by a group of computers (as examples of machines 1100 including processors 1110) , with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API) . In certain embodiments, for example, a client device may relay or operate in communication with cloud computing systems and may access circuit design information in a cloud environment.
The performance of certain of the operations may be distributed among the processors, not only residing within a single machine 1100, but deployed  across a number of machines 1100. In some example embodiments, the processors 1110 or processor-implemented components are located in a single geographic location (e.g., within a home environment, an office environment, or a server farm) . In other example embodiments, the processors or processor-implemented components are distributed across a number of geographic locations.
The various memories (i.e., 1130, 1132, 1134, and/or the memory of the processor (s) 1110) and/or the storage unit 1136 may store one or more sets of instructions 1116 and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. These instructions (e.g., the instructions 1116) , when executed by the processor (s) 1110, cause various operations to implement the disclosed embodiments.
As used herein, the terms “machine-storage medium, ” “device-storage medium, ” and “computer-storage medium” mean the same thing and may be used interchangeably. The terms refer to a single or multiple storage devices and/or media (e.g., a centralized or distributed database, and/or associated caches and servers) that store executable instructions 1116 and/or data. The terms shall accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media, including memory internal or external to processors. Specific examples of machine-storage media, computer-storage media and/or device-storage media include non-volatile memory, including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine-storage media, ” “computer-storage media, ” and “device-storage media” specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term “signal medium” discussed below.
In some examples, one or more portions of the network 1180 may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN) , a LAN,  a wireless LAN (WLAN) , a WAN, a wireless WAN (WWAN) , a metropolitan-area network (MAN) , the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN) , a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, anetwork, another type of network, or a combination of two or more such networks. For example, the network 1180 or a portion of the network 1180 may include a wireless or cellular network, and the coupling 1182 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or another type of cellular or wireless coupling. In this example, the coupling 1182 may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT) , Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS) , High-Speed Packet Access (HSPA) , Worldwide Interoperability for Microwave Access (WiMAX) , Long-Term Evolution (LTE) standard, others defined by various standard-setting organizations, other long-range protocols, or other data transfer technology.
The instructions may be transmitted or received over the network using a transmission medium via a network interface device (e.g., a network interface component included in the communication components) and utilizing any one of a number of well-known transfer protocols (e.g., hypertext transfer protocol (HTTP) ) . Similarly, the instructions may be transmitted or received using a transmission medium via the coupling (e.g., a peer-to-peer coupling) to the devices 1170. The terms “transmission medium” and “signal medium” mean the same thing and may be used interchangeably in this disclosure. The terms “transmission medium” and “signal medium” shall be taken to include any intangible medium that is capable of storing, encoding, or carrying the instructions for execution by the machine, and include digital or analog communications signals or other intangible media to facilitate communication of such software. Hence, the terms “transmission  medium” and “signal medium” shall be taken to include any form of modulated data signal, carrier wave, and so forth. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
The terms “machine-readable medium, ” “computer-readable medium, ” and “device-readable medium” mean the same thing and may be used interchangeably in this disclosure. The terms are defined to include both machine-storage media and transmission media. Thus, the terms include both storage devices/media and carrier waves/modulated data signals. For instance, an embodiment described herein can be implemented using a non-transitory medium (e.g., a non-transitory computer-readable medium) .
Throughout this specification, plural instances may implement resources, components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components.
As used herein, the term “or” may be construed in either an inclusive or exclusive sense. The terms “a” or “an” should be read as meaning “at least one, ” “one or more, ” or the like. The presence of broadening words and phrases such as “one or more, ” “at least, ” “but not limited to, ” or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent. Additionally, boundaries between various resources, operations, components, engines, and data stores are somewhat arbitrary, and particular operations are illustrated in a context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within a scope of various embodiments of the present disclosure. The  specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
It will be understood that changes and modifications may be made to the disclosed embodiments without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure.

Claims (20)

  1. A system comprising:
    one or more hardware processors; and
    at least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising:
    receiving an incident alert associated with a deviation in a metric, the incident alert comprising contextual data representing the deviation;
    identifying a plurality of correlated metrics based on the contextual data;
    generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals;
    constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and
    determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
  2. The system of claim 1, wherein the generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises:
    mapping the plurality of correlated metrics into the plurality of categories of signals.
  3. The system of claim 2, wherein each signal in a category of signals corresponds to a performance indicator for monitoring and managing performance of a service.
  4. The system of claim 1, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
  5. The system of claim 1, wherein the contextual data comprises a query associated with the deviation in the metric.
  6. The system of claim 5, wherein the operations comprise:
    executing the query against a database to retrieve monitoring data associated with the deviation;
    performing a similarity search based on the monitoring data associated with the deviation; and
    in response to performing the similarity search, identifying the plurality of correlated metrics.
  7. The system of claim 6, wherein the operations comprise:
    in response to performing the similarity search, assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
  8. The system of claim 7, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  9. The system of claim 7, wherein the operation of determining the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises:
    applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm;
    generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph;
    ranking the plurality of scores in a descending order; and
    identifying a correlated metric with a top rank as the root cause of the incident.
  10. The system of claim 1, wherein the operations comprise:
    constructing the causal graph further based on tabular data that comprises a plurality of causal links between the plurality of correlated metrics, a causal link representing a cause-and-effect relationship between a pair of correlated metrics.
  11. A method comprising:
    receiving an incident alert associated with a deviation in a metric, the incident alert comprising contextual data representing the deviation;
    identifying a plurality of correlated metrics based on the contextual data;
    generating, by at least one hardware processor, a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals;
    constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and
    determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
  12. The method of claim 11, the generating of the plurality of labels for the plurality of correlated metrics in accordance with the plurality of categories of signals comprises:
    mapping the plurality of correlated metrics into the plurality of categories of signals.
  13. The method of claim 12, wherein each signal in a category of signals corresponds to a performance indicator for monitoring and managing performance of a service.
  14. The method of claim 11, wherein the deviation in the metric represents a departure from an expected behavior of the metric, and wherein the link analysis algorithm comprises a weighted PageRank algorithm.
  15. The method of claim 11, wherein the contextual data comprises a query associated with the deviation in the metric.
  16. The method of claim 15, comprising:
    executing the query against a database to retrieve monitoring data associated with the deviation;
    performing a similarity search based on the monitoring data associated with the deviation; and
    in response to performing the similarity search:
    identifying the plurality of correlated metrics, and
    assigning a similarity value to each of the plurality of correlated metrics, the similarity value representing a degree of similarity between a correlated metric and the metric associated with the deviation.
  17. The method of claim 16, wherein the similarity value is associated with a direction indicator of a type of deviation associated with a correlated metric, the type of deviation corresponding to one of a spike, a surge, a dip, a drop, a fluctuation, or a pattern.
  18. The method of claim 16, the determining of the root cause of the incident associated with the incident alert based on the causal graph and the plurality of weights assigned to the plurality of correlated metrics comprises:
    applying a plurality of similarity values assigned to the plurality of correlated metrics as the plurality of weights to the link analysis algorithm;
    generating, using the link analysis algorithm, a plurality of scores for the plurality of correlated metrics based on the causal graph;
    ranking the plurality of scores in a descending order; and
    identifying a correlated metric with a top rank as the root cause of the incident.
  19. The method of claim 11, comprising:
    constructing the causal graph further based on tabular data that comprises a plurality of causal links between the plurality of correlated metrics, a causal link representing a cause-and-effect relationship between a pair of correlated metrics.
  20. A machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising:
    receiving an incident alert associated with a deviation in a metric, the incident alert comprising contextual data representing the deviation;
    identifying a plurality of correlated metrics based on the contextual data;
    generating a plurality of labels for the plurality of correlated metrics in accordance with a plurality of categories of signals;
    constructing a causal graph based on the plurality of labels and a topology graph representing the plurality of correlated metrics; and
    determining, using a link analysis algorithm, a root cause of an incident associated with the incident alert based on the causal graph and a plurality of weights assigned to the plurality of correlated metrics, the root cause of the incident corresponding to a correlated metric from the plurality of correlated metrics.
PCT/CN2024/093949 2024-05-17 2024-05-17 Incident triage and root cause analysis Pending WO2025236282A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/093949 WO2025236282A1 (en) 2024-05-17 2024-05-17 Incident triage and root cause analysis

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/093949 WO2025236282A1 (en) 2024-05-17 2024-05-17 Incident triage and root cause analysis

Publications (1)

Publication Number Publication Date
WO2025236282A1 true WO2025236282A1 (en) 2025-11-20

Family

ID=97719193

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/093949 Pending WO2025236282A1 (en) 2024-05-17 2024-05-17 Incident triage and root cause analysis

Country Status (1)

Country Link
WO (1) WO2025236282A1 (en)

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20020083371A1 (en) * 2000-12-27 2002-06-27 Srinivas Ramanathan Root-cause approach to problem diagnosis in data networks
US8031634B1 (en) * 2008-03-31 2011-10-04 Emc Corporation System and method for managing a virtual domain environment to enable root cause and impact analysis
US20200371857A1 (en) * 2018-11-25 2020-11-26 Aloke Guha Methods and systems for autonomous cloud application operations
CN113240139A (en) * 2021-06-03 2021-08-10 南京中兴新软件有限责任公司 Alarm cause and effect evaluation method, fault root cause positioning method and electronic equipment
US20210311996A1 (en) * 2020-04-03 2021-10-07 International Business Machines Corporation Providing causality augmented information responses in a computing environment
US11321885B1 (en) * 2020-10-29 2022-05-03 Adobe Inc. Generating visualizations of analytical causal graphs
US20230069074A1 (en) * 2021-08-20 2023-03-02 Nec Laboratories America, Inc. Interdependent causal networks for root cause localization
CN115941446A (en) * 2022-12-27 2023-04-07 中国联合网络通信集团有限公司 Alarm root cause location method, apparatus, electronic device and computer readable medium

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20020083371A1 (en) * 2000-12-27 2002-06-27 Srinivas Ramanathan Root-cause approach to problem diagnosis in data networks
US8031634B1 (en) * 2008-03-31 2011-10-04 Emc Corporation System and method for managing a virtual domain environment to enable root cause and impact analysis
US20200371857A1 (en) * 2018-11-25 2020-11-26 Aloke Guha Methods and systems for autonomous cloud application operations
US20210311996A1 (en) * 2020-04-03 2021-10-07 International Business Machines Corporation Providing causality augmented information responses in a computing environment
US11321885B1 (en) * 2020-10-29 2022-05-03 Adobe Inc. Generating visualizations of analytical causal graphs
CN113240139A (en) * 2021-06-03 2021-08-10 南京中兴新软件有限责任公司 Alarm cause and effect evaluation method, fault root cause positioning method and electronic equipment
US20230069074A1 (en) * 2021-08-20 2023-03-02 Nec Laboratories America, Inc. Interdependent causal networks for root cause localization
CN115941446A (en) * 2022-12-27 2023-04-07 中国联合网络通信集团有限公司 Alarm root cause location method, apparatus, electronic device and computer readable medium

Similar Documents

Publication Publication Date Title
US20250117701A1 (en) Machine learning platform
US20240211106A1 (en) User interface based variable machine modeling
US11876837B2 (en) Network privacy policy scoring
US12417229B2 (en) Managing database offsets with time series
US11720601B2 (en) Active entity resolution model recommendation system
US11940897B2 (en) Contextualized notifications for verbose application errors
CN108370324B (en) Distributed database operation data tilt detection
US12603906B2 (en) Alert monitoring of data based on recommended attribute values
WO2019173202A1 (en) Systems and methods for decision tree ensembles for selecting actions
US20230350922A1 (en) Determining zone identification reliability
KR20240052035A (en) Verification of crowdsourced field reports based on user trust
US20250272552A1 (en) Machine learning model training on risk prediction using graph knowledge distillation
US20240127306A1 (en) Generation and management of data quality scores using multimodal machine learning
US20240054571A1 (en) Matching influencers with categorized items using multimodal machine learning
WO2023209638A1 (en) Determining zone identification reliability
KR20240089013A (en) Depletion modeling to estimate survey completeness by region
US20260003853A1 (en) Data conflict resolution and management
US12476890B2 (en) Incomplete matrix profile-based anomaly detection in time series data
WO2026020263A1 (en) Machine learning model training using feature augmentation
US20220383223A1 (en) Vendor profile data processing and management
WO2026036312A1 (en) Machine learning model training for data annotation
EP4733959A1 (en) Search optimization using query-based contextual features
US20250363408A1 (en) Machine learning model training using a self-training approach for knowledge distillation
US20260003997A1 (en) User consent management and coordination
US12596585B2 (en) Data processing and management

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24938305

Country of ref document: EP

Kind code of ref document: A1