WO2020157339A1 - Remediation of real-world systems - Google Patents
Remediation of real-world systems Download PDFInfo
- Publication number
- WO2020157339A1 WO2020157339A1 PCT/EP2020/052643 EP2020052643W WO2020157339A1 WO 2020157339 A1 WO2020157339 A1 WO 2020157339A1 EP 2020052643 W EP2020052643 W EP 2020052643W WO 2020157339 A1 WO2020157339 A1 WO 2020157339A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- elements
- remediation
- state
- positive
- variables
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/06—Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
- G06Q10/067—Enterprise or organisation modelling
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/10—Office automation; Time management
- G06Q10/105—Human resources
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/01—Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
Definitions
- identifying one or more changes to the variables comprises identifying changes that would map the element into the cluster containing the model element. This provides confidence that the remediation action will be successful, as there is a positive track record of achieving a positive state for an element as a final state for an element in a positive state cluster.
- a set of attributes may be established for each element. These may have predictive power, but they may not be changeable or may change according to existing rules (for example, age and gender would typically be attributes).
- the vector for each element may then represent the set of variables and the set of attributes.
- the step of identifying elements in a negative state for remediation comprises ranking the elements in order of need of remediation.
- Such ranking may comprise obtaining a clustering score for the element determined from the proximity of the element to elements in a positive state and to elements in a negative state. It also may comprise obtaining a pattern matching score for the element determined from identification of positive patterns associated with a positive state and negative patterns identified with a negative state and matching the elements against the positive patterns and the negative patters.
- the ranking may comprise assigning each element a ranking score formed from a combination of the clustering score and the pattern matching score.
- determining clusters of elements may comprise selecting clusters to optimise segregation of elements in positive state from elements in a negative state.
- the real-world system is an organisation and the elements are employees of the organisation.
- the negative state of the element may then be departure of an employee, with the variables describing aspects of the job roles and employment conditions of the employees.
- the remediation action may then comprise varying one or more of said aspects or employment conditions.
- Employee retention and workforce attrition are major issues for employers. While there are existing tools for analysing HR data in organisations, such as IBM HR Analytics and SAP Workforce Analytics, this typically have some predictive capability - for example, by identifying characteristics that may indicate an attrition risk - but do not propose any form of remediation action, which is left to HR professionals.
- the real-world system is a logistics system and the elements are items operated on by the logistics system, such as items to be transported from a producer to a retailer.
- the negative state of an element may be failure to meet conditions of service for the item, such as delivery within a particular time and at a particular level of quality.
- the variables may relate to location and condition of the items and the remediation action may comprise moving the item to a different location or changing the condition of the item.
- Condition here may include features such as temperature or exposure to sunlight - in the case of produce or perishable items, varying of such conditions may be as significant as speed of transit in ensuring that the items will be at a sale location at saleable quality for a sufficient length of time to meet conditions of service.
- the real-world system is a task performance environment and each element is a task to be performed. This may involve, for example, carrying out a large number of computational tasks on a distributed computing system.
- the negative state may here be failure or predicted failure of a task, and the remediation action may be changing performance of the task to reduce the risk of failure.
- the real-world system is a customer contract database, and each element is a customer contract. The negative state may then be termination or predicted termination of the contract by the customer, and the remediation action may comprise changing one or more contract terms or conditions to improve the likelihood of customer retention.
- the disclosure provides a computing system comprising a processor and a memory, wherein the processor of the computing system is programmed to execute the method of the first aspect using the memory.
- Figure 1 shows a computer-implemented system suitable for determining remediation actions in a real-world system
- Figure 2 is a flow diagram indicating a method in according to the disclosure for determining a remediation action to perform on an element of the real-world system;
- Figure 3 indicates clustering of elements in the method of Figure 2;
- Figure 4 illustrates mapping of an element for remediation to a model element in the method of Figure 2, shown in the context of an HR management system;
- Figure 5 illustrates further the making of a mapping choice in the mapping of Figure 4
- Figure 6 illustrates system elements in a system for performing the method of Figure 4
- Figure 7 shows comparative success of a model according an embodiment of the disclosure in identifying a negative state and performing a remediation action considered against conventional approaches in the context of an HR management system.
- FIG. 1 shows in schematic terms a system suitable for implementing embodiments of the disclosure.
- a real-world system 1 comprises a plurality of elements 2. Each of these elements 2 has its own set of attributes and variables - these can be measured by an appropriate measurement system 3. These characterisations of each element 2 in terms of attributes and variables are input to an analysis system 4.
- the analysis system 4 analyses the input data and provides outputs to an implementation system 5.
- the implementation system 5 implements changes to the variables of the real-world system 1 , in this case to one element 2a of the set of elements 2 in the real-world system 1.
- At least the analysis system 4 is a computing system programmed to perform analysis of input data and provision of outputs for implementation - in embodiments of the disclosure, these outputs include remediation actions for implementation by the implementation system 5.
- the measurement system 3 and the implementation system 5 may be implemented in whole or in part as computing systems, and they may be implemented as a collection of discrete interacting elements or as a single entity. In embodiments, some or all of the measurement system 3, the analysis system 4 and the implementation system 5 may be implemented within the same computing system.
- Embodiments of the disclosure relate to situations in which individual elements may be in a negative state - for example, in an organisation where the HR system is being modelled an element may be an individual and a negative state may indicate that an employee has left, or is determined to be at risk of leaving, the organisation - but where there is a possibility of making one or more changes that may move an individual element from a negative state to a positive state.
- the positive state may be an employee at a low risk of attrition, and the changes may involve changing specific aspects of the job (salary, work group or grade, for example).
- the philosophy in remediation in elements of the disclosure is to understand which elements are in the positive state and which in the negative state. Where a final state has been reached - for example, an employee has left the company - the state may is determined, but in other cases it will be necessary to determine from measured data whether the element should be placed in a positive or negative state.
- One approach involves determining how elements can be clustered - this may use changeable variables (and possibly also fixed attributes) to characterise the elements as vectors in a multidimensional space, and this can be used to identify clusters of similar elements. Fixed state elements can help to identify whether clusters as a whole can be recognised as having a positive or a negative state.
- Pattern matching may also be used to assign values to individual elements based on recognition of patterns common to known negative state elements or known positive state elements. These two approaches - cluster analysis and pattern matching - can be combined together to identify elements in a positive and a negative state, and also to identify elements in a particular need of remediation.
- model element is chosen to be the nearest element in the multidimensional space - this can easily determine by vector analysis - provided that the model element is both in a positive state and in a cluster in a positive state.
- the differences between the variables of the identified element and the model element are established as real-world differences, and a remediation action is proposed that will bring the identified element into proximity with the model element in the multidimensional space.
- the first step is to establish 21 a set of variables for each of a plurality of elements of the real-world system.
- the purpose of this is to enable the key characteristics of the elements to be represented in the model used for analysis so that the model will have real predictive power.
- the next step is to model 22 these elements in a multidimensional space.
- the elements are represented by vectors determined by their variables (and, if used, attributes). This mapping on to a multidimensional space allows elements to be considered as clusters. Determination of these clusters is, as will be shown below, a highly significant part of this modelling process.
- each of the elements is in a positive or a negative state.
- an element will already have a defined state - for example, in the employee database example, some employees will have left the company (or left within a short timescale) and be identified as lost to attrition and so marked in the negative state, whereas others might be identified as long-time employees (for example, having remained with the company until retirement or beyond a long-service threshold) and placed in a positive state.
- other elements will be assessed as being in a positive or negative state by virtue of their resemblance to other elements, using both clustering and pattern matching approaches.
- clusters are in a positive or a negative state, in the sense that they exceed a threshold value of elements in a positive state or elements in a negative state.
- One way to do this is to establish a critical percentage of negative elements, and to identify a cluster as a negative state cluster if the critical percentage is exceeded and as a positive state cluster if the critical percentage is not exceeded. In this way, all clusters will be identified as being in either a positive state or a negative state.
- the first step in determining the remediation action to be performed is to map 26 the element for remediation on to a model element.
- This model element is the nearest positive model element in the multidimensional space that is also located in a positive state cluster. This approach is chosen as it is minimally disruptive (it is the least distance in the multidimensional space) and as it stands the best chance of achieving the desired result (as there will be a track record of elements achieving a positive state as a final state in a positive state cluster).
- One or more changes to the variables of the element for remediation can in this way be identified 28 that would map it more closely to the model element.
- a suitable result would be one that, if implemented, would move the element from remediation from its current cluster to the cluster containing the model element.
- the real-world implementation of those changes is then proposed 29 as the remediation action.
- E L E the subset of employees who have left the organization in the past 12 months
- E A E ⁇ EL be the set of employees who are still active in the company.
- Attrition prevention problem finding the least disruptive set of remedial actions on f maximizing the retention of k active employees with the highest risk of attrition.
- Predictive scoring implements a novel core methodology to assess and predict the probability of voluntary attrition for an active employee.
- the system utilizes an aggregated scoring function based on (a) clustering - giving the employee’s proximity to other employees, and (b) frequent pattern mining - giving the similarity between any employee and the traits of a typical leaver. The score is then used to predict the chances of attrition for each individual, and subsequently recommend remedial steps to improve retention.
- the two major components of this level are the employee clustering module and the frequent feature pattern matching module, both of which are described below.
- each employee record (active or leaver) is represented as a high-dimensional vector.
- the system performs clustering to extract clusters or employee groups demonstrating a high degree of feature similarity.
- L 2 norm Euclidean distance
- the clustering-based score assigned to an active employee e is given by the average of these two scores:
- AF and LF be the set of frequent patterns extracted from the active E A and leaver E L data partitions respectively.
- the commonality of frequent patterns between the active and leaver employees provides an important connecting factor in predicting active employee attrition.
- certain patterns are more informative than others, and hence should contribute with a higher weight in the process of prediction.
- the system ranks the frequent patterns in FP and assigns weights based on the computed ranks. Specifically, for each frequent pattern p j e FP, the relative frequencies of p j in E A and E L data partitions are computed.
- the relative frequency of a pattern is defined as its frequency of occurrence in the particular dataset over the total number of records in the dataset (expressed in percentage).
- the system computes an aggregated score based on matching the employee features with the frequent patterns, taking into account the ranks assigned to the frequent patterns. Specifically, if a frequent pattern is matched in an employee feature vector, the pattern matching score of the employee is incremented by the inverse of the rank of the frequent pattern.
- the final predictive attrition score, PS, assigned by the system to an active employee e, is computed by a linear combination of the clustering score and the pattern matching score of the predictive scoring module.
- Our framework uses the averaged score, that is,
- the employee predictive attrition score provides a proxy for the probability of voluntary attrition of an active employee, and it is larger if the predicted chances of separation are higher.
- PS the system generates an ordered list of active employees in the decreasing order of the scores. This provides the attrition rank list for the organization with probable candidates at the highest risk of attrition due to varying factors. It is for these employees that our framework subsequently recommends remedial actions to help improve their retention.
- Cluster Purity One significant consideration here is the nature of the clusters.
- An important tuneable parameter here is k, the number of clusters (or employee groups) to be created during the predictive attrition score computation phase. To this end, we empirically select the best value optimizing the performance of the system. The value of k was varied between 5 and 50 to select the setting that provides the best clusters in terms of“cluster purity”. Traditional clustering approaches tend to optimize the intra- and inter- cluster distances. However, for our problem setting, we prefer a clean segregation of leaver and active employees to reduce noise and capture high quality feature patterns within the clusters.
- a key novel feature of this framework is the recommendation of personalized remedial steps on an individual level to improve the rate of attrition - moving from predictive to remedial HR analytics.
- the above ranked list of employees are read in the order of their scores, and based on the employee feature clustering and the active employee feature, recommendations are generated and presented to the management.
- a measure of the deviation is computed between the feature vector of e, and that of the closest active employee - providing the basis for the remedial recommendations.
- the system computes the vector difference where e ' is the
- Figure 4 shows an example of a remedial action in a cut-down system.
- a bi-dimensional space here: Salary Band, Team Size.
- the system needs to create a remedial action recommendation for the active employee in the bottom left cluster where the three vectors start from.
- the system considers all the active people in all the active clusters, and finds the closest active to compute the remedial vector.
- This multi-dimension view of the computed feature vector difference is translated to a natural language based recommendation.
- a salary difference in the feature vectors provides a simple remedial action of “Increase salary”, while a difference in the features salary and job grade might correspond to an “Increase responsibility” recommendation.
- a few remedial actions might not be physically possible or logically practical (e.g., move to smaller groups where there are no smaller groups, or decrease salary).
- a final layer of inexpensive human expert intervention for manual filtering might be required. This step is normally performed by HR at the time of browsing and applying the results. Incorporation of natural language processing and common sense knowledge extraction into our current framework provides an interesting direction of future work for creating a fully automated pipeline.
- Figure 6 shows a more detailed breakdown of the modules of a computational system suitable for carrying out the method described in detail above.
- the data importer module imports the data needed by the system.
- Employee data collected during the onboarding process or voluntarily disclosed by an individual, typically resides in an enterprise’s encrypted data store (dedicated data warehouse or cloud storage).
- the data importer module provides an interaction layer between the proposed framework and the raw storage. This layer is particularly dedicated to protecting the security and sensitivity aspects of the raw information, so that unauthorized access or leakage can be prevented.
- the data importer module enables raw data ingestion and its“masking” by performing the following steps: (i) reads the raw data from storage, (ii) decrypts the data, (iii) removes sensitive information pertaining to employees, and (iv) anonymizes employee details with random but unique identifiers. The last two steps are needed to enable maintenance, further development, and analysis by the developers. Additionally, data cleaning involving duplicate removal and handling of corrupt/missing information is also performed by the data importer. This forms the final clean input data that is visible to the higher layers of the system architecture.
- the feature selection and engineering module is equipped for pre-processing the input date received from the importer module, and for augmenting it with new features computed on top of the available ones.
- the different aspects present in the employee data that describe an employee, such as salary band, education, seniority, etc., are treated as employee features in this framework.
- this module then performs an analysis of the employee features across the entire dataset to evaluate the“goodness” of each of the obtained features based on its discriminative power for partitioning the data into (dis-)similar groups.
- the Pearson correlation coefficient may be used to identify and retain features (which may themselves be uncorrelated and demonstrate a high variance) for extracting meaningful patterns to predict individual employee attrition chances.
- Other feature selection techniques such as Principal Component Analysis (PCA), can easily be adopted in this module, making the system dynamic to several different application domains and requirements.
- PCA Principal Component Analysis
- team-based features For augmentation of an employee feature set, team-based features (specifically computed on new hires joining each team) can be considered. Team-based features include parameters like team size, average team seniority, difference between employee seniority and average team seniority, team diversity score, etc.
- new-hire features include the team-based features, but these are re-computed for each new hire to highlight the evolution of a team and the hiring strategy of the company’s leadership.
- Figure 7 indicates the results of an evaluation on a historical dataset showing the effectiveness of an embodiment of the current disclosure (identified as CLARA) against other approaches.
- Other machine learning approaches used for comparison and variants of the current approach are the following:
- XGBoost Trees Similar to the Random Forest approach, but instead the boost tree prediction model is used.
- Support Vector Machines Trained to classify employees into the binary category of active or leaver based on the different weights learnt on employee features. The predicted probability of an employee belonging to one of the classifications is used to rank currently active employees
- K-Means Clustering (without FPM) - This provide an ablation study for our proposed model, by using only the clustering-based scoring module (without frequent pattern mining). Active employees are ranked based on the clustering score CS Spectral Clustering - Provides a variation of the clustering technique used based on eigen-space decomposition. The distance and correlation-based score is computed between each pair of employee points to find the clusters. Frequent Pattern Mining (without K-Means) - The final ablation baseline considers only the frequent pattern overlap score PO to rank the active employees to obtain those at the highest risk of attrition.
- the CLARA embodiment is markedly more successful in identifying the employees most at risk of attrition. Effect of remedial actions was considered by running the system on the entire data set, identifying remedial actions and applying them all and evaluating the data set to see (i) what happened to the highest ranked elements and (ii) how do the highest ranked elements in the new data set compare to those in the old data set (ie how does the attrition risk compare)? Note the subtle difference between i) and ii): the aim of the first is to“follow” the first x employee from R 1 to R 2 . The aim of the second is to“follow” the first x ranks from R 1 to R 2 .
- PS R1 ( ) be function returning the score of an employee in R 1
- PS R 2 ( ) be the function returning the score of an employee in R 2
- map ( ) be the function returning the rank that the argument has in R 2
- ⁇ R1 be the employee with the lowest score (i.e. ranked last) in R 1 .
- Embodiments of this approach may be used in a logistics system.
- Elements in the logistics system may be items to be transported, for example from a producer to a retailer.
- the negative state of an element may be failure to meet conditions of service for the item, such as delivery within a particular time and at a particular level of quality.
- the variables may relate to location and condition of the items and the remediation action may comprise moving the item to a different location or changing the condition of the item.
- Condition here may include features such as temperature or exposure to sunlight - in the case of produce or perishable items, varying of such conditions may be as significant as speed of transit in ensuring that the items will be at a sale location at saleable quality for a sufficient length of time to meet conditions of service.
- the real-world system may be a task performance environment and each element is a task to be performed. This may involve, for example, carrying out many computational tasks on a distributed computing system.
- the negative state may here be failure or predicted failure of a task, and the remediation action may be changing performance of the task to reduce the risk of failure.
- An implementation of the approach described here could be run across the full set of tasks to be carried out, and those at greatest risk of failure identified. These tasks may be re-allocated by an appropriate remediation action - changing their place in a queue, or assigning them to a particular computation engine, for example - to reduce this risk of failure.
- the real-world system is a customer contract database, and each element is a customer contract.
- the negative state may then be termination or predicted termination of the contract by the customer, and the remediation action may comprise changing one or more contract terms or conditions to improve the likelihood of customer retention. This is relatively closely aligned to the employee retention case - the problem is now customer retention rather than employee retention - and the set of variables directly comparable.
Landscapes
- Engineering & Computer Science (AREA)
- Business, Economics & Management (AREA)
- Human Resources & Organizations (AREA)
- Strategic Management (AREA)
- Entrepreneurship & Innovation (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Economics (AREA)
- Physics & Mathematics (AREA)
- Software Systems (AREA)
- Marketing (AREA)
- General Business, Economics & Management (AREA)
- Tourism & Hospitality (AREA)
- Quality & Reliability (AREA)
- Operations Research (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Educational Administration (AREA)
- Artificial Intelligence (AREA)
- Game Theory and Decision Science (AREA)
- Development Economics (AREA)
- General Engineering & Computer Science (AREA)
- Medical Informatics (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
A computer-implemented method of determining a remediation action for an element of a real-world system comprises the following steps. First of all, a set of variables are established (21) for each of a plurality of elements of the real-world system. These are used to determine (22) clusters of elements. These clusters are based on proximity of elements in a space in which each element is represented by a vector comprising the set of variables for that element. It is then established (23) whether each of the elements is in a positive or a negative state, and it is also established (24) whether each cluster is in a positive state or a negative state. This determination is based on the composition of that cluster of elements, specifically whether elements in the cluster are in a positive state or in a negative state. Elements are then identified (25) which are in a negative state and which are candidates for remediation. An element identified for remediation is mapped (26) on to a model element. The model element is the nearest element in a positive state that is also in a cluster in a positive state. Differences are then determined (27) in the variables between the element for remediation and its model element, and one or more changes identified (28) to the variables of the element for remediation that would map it more closely to the model element. The real-world implementation of those changes is determined (29) as the remediation action. A suitable computing system for implementing this method is also described.
Description
REMEDIATION OF REAL-WORLD SYSTEMS
Field of Disclosure
The present disclosure relates to remediation of real-world systems, and it describes a computer-implemented method to achieve such remediation.
Background to Disclosure
Many real-world systems are very complex, and even when there is rich data available to describe them, it may be difficult to determine the best actions to improve the state of the system. Some systems contain a number of elements that can be well- described by variables, but which may transition from a desired to a non-desired state for reasons that are not simply deterministic (or which cannot be readily determined). Analysis of the elements and their variables is often used in such systems to determine which elements are at risk of transitioning from a desired to a non-desired state. It is typically much more difficult to identify how such transitions could be prevented.
Summary of Disclosure
In a first aspect, the disclosure provides a computer-implemented method of determining a remediation action for an element of a real-world system, the method comprising: establishing a set of variables for each of a plurality of elements of the real-world system; determining clusters of elements based on proximity of elements in a space in which each element is represented by a vector comprising the set of variables for that element; establishing whether each of the elements is in a positive or a negative state; establishing whether each cluster is in a positive state or a negative state based on the composition of that cluster of elements in a positive state and of elements in a negative state; identifying elements in a negative state for remediation; mapping each element identified for remediation on to a model element, being the nearest element in a positive state that is also in a cluster in a positive state; determining differences in the variables between the element for remediation and its model element; identifying one or more changes to the variables of the element for
remediation that would map it more closely to the model element; and determining the real-world implementation of those changes as the remediation action.
By taking this approach, it is possible not only to analyse a complex system to determine where an element is a negative state, but also to identify a remediation action to move that element to a positive state. Moreover, this approach achieves this with the least disturbance to the overall system, meaning that it is the least disruptive approach to achieving this remediation. This is thus the remediation step least likely to affect other elements of the system, and so also least likely to result in unintended consequences.
In embodiments, identifying one or more changes to the variables comprises identifying changes that would map the element into the cluster containing the model element. This provides confidence that the remediation action will be successful, as there is a positive track record of achieving a positive state for an element as a final state for an element in a positive state cluster.
In embodiments, in addition to establishing a set of variables for each element, a set of attributes may be established for each element. These may have predictive power, but they may not be changeable or may change according to existing rules (for example, age and gender would typically be attributes). The vector for each element may then represent the set of variables and the set of attributes.
In embodiments, the step of identifying elements in a negative state for remediation comprises ranking the elements in order of need of remediation. Such ranking may comprise obtaining a clustering score for the element determined from the proximity of the element to elements in a positive state and to elements in a negative state. It also may comprise obtaining a pattern matching score for the element determined from identification of positive patterns associated with a positive state and negative patterns identified with a negative state and matching the elements against the positive patterns and the negative patters. In some such cases, the ranking may comprise assigning each element a ranking score formed from a combination of the clustering score and the pattern matching score.
In embodiments, determining clusters of elements may comprise selecting clusters to optimise segregation of elements in positive state from elements in a negative state.
This approach to remediation actions can be used in a wide variety of real-world systems.
In certain embodiments, the real-world system is an organisation and the elements are employees of the organisation. The negative state of the element may then be departure of an employee, with the variables describing aspects of the job roles and employment conditions of the employees. The remediation action may then comprise varying one or more of said aspects or employment conditions. Employee retention and workforce attrition are major issues for employers. While there are existing tools for analysing HR data in organisations, such as IBM HR Analytics and SAP Workforce Analytics, this typically have some predictive capability - for example, by identifying characteristics that may indicate an attrition risk - but do not propose any form of remediation action, which is left to HR professionals.
In other embodiments, the real-world system is a logistics system and the elements are items operated on by the logistics system, such as items to be transported from a producer to a retailer. Here, the negative state of an element may be failure to meet conditions of service for the item, such as delivery within a particular time and at a particular level of quality. The variables may relate to location and condition of the items and the remediation action may comprise moving the item to a different location or changing the condition of the item. Condition here may include features such as temperature or exposure to sunlight - in the case of produce or perishable items, varying of such conditions may be as significant as speed of transit in ensuring that the items will be at a sale location at saleable quality for a sufficient length of time to meet conditions of service.
In still other embodiments, the real-world system is a task performance environment and each element is a task to be performed. This may involve, for example, carrying out a large number of computational tasks on a distributed computing system. The negative state may here be failure or predicted failure of a task, and the remediation action may be changing performance of the task to reduce the risk of failure.
In yet further embodiments, the real-world system is a customer contract database, and each element is a customer contract. The negative state may then be termination or predicted termination of the contract by the customer, and the remediation action may comprise changing one or more contract terms or conditions to improve the likelihood of customer retention.
In a second aspect, the disclosure provides a computing system comprising a processor and a memory, wherein the processor of the computing system is programmed to execute the method of the first aspect using the memory.
Description of Specific Embodiments
Specific embodiments of the disclosure are now described, by way of example, with reference to the accompanying drawings, of which:
Figure 1 shows a computer-implemented system suitable for determining remediation actions in a real-world system;
Figure 2 is a flow diagram indicating a method in according to the disclosure for determining a remediation action to perform on an element of the real-world system; Figure 3 indicates clustering of elements in the method of Figure 2;
Figure 4 illustrates mapping of an element for remediation to a model element in the method of Figure 2, shown in the context of an HR management system;
Figure 5 illustrates further the making of a mapping choice in the mapping of Figure 4; Figure 6 illustrates system elements in a system for performing the method of Figure 4; and
Figure 7 shows comparative success of a model according an embodiment of the disclosure in identifying a negative state and performing a remediation action considered against conventional approaches in the context of an HR management system.
Figure 1 shows in schematic terms a system suitable for implementing embodiments of the disclosure. A real-world system 1 comprises a plurality of elements 2. Each of these elements 2 has its own set of attributes and variables - these can be measured
by an appropriate measurement system 3. These characterisations of each element 2 in terms of attributes and variables are input to an analysis system 4. The analysis system 4 analyses the input data and provides outputs to an implementation system 5. The implementation system 5 implements changes to the variables of the real-world system 1 , in this case to one element 2a of the set of elements 2 in the real-world system 1.
At least the analysis system 4 is a computing system programmed to perform analysis of input data and provision of outputs for implementation - in embodiments of the disclosure, these outputs include remediation actions for implementation by the implementation system 5. In embodiments, the measurement system 3 and the implementation system 5 may be implemented in whole or in part as computing systems, and they may be implemented as a collection of discrete interacting elements or as a single entity. In embodiments, some or all of the measurement system 3, the analysis system 4 and the implementation system 5 may be implemented within the same computing system.
Embodiments of the disclosure relate to situations in which individual elements may be in a negative state - for example, in an organisation where the HR system is being modelled an element may be an individual and a negative state may indicate that an employee has left, or is determined to be at risk of leaving, the organisation - but where there is a possibility of making one or more changes that may move an individual element from a negative state to a positive state. In the HR example, the positive state may be an employee at a low risk of attrition, and the changes may involve changing specific aspects of the job (salary, work group or grade, for example).
The philosophy in remediation in elements of the disclosure is to understand which elements are in the positive state and which in the negative state. Where a final state has been reached - for example, an employee has left the company - the state may is determined, but in other cases it will be necessary to determine from measured data whether the element should be placed in a positive or negative state. One approach involves determining how elements can be clustered - this may use changeable variables (and possibly also fixed attributes) to characterise the elements as vectors in a multidimensional space, and this can be used to identify clusters of similar elements.
Fixed state elements can help to identify whether clusters as a whole can be recognised as having a positive or a negative state. Pattern matching may also be used to assign values to individual elements based on recognition of patterns common to known negative state elements or known positive state elements. These two approaches - cluster analysis and pattern matching - can be combined together to identify elements in a positive and a negative state, and also to identify elements in a particular need of remediation.
Once an element in need of remediation is identified, it is then mapped on to a model element. The model element is chosen to be the nearest element in the multidimensional space - this can easily determine by vector analysis - provided that the model element is both in a positive state and in a cluster in a positive state. The differences between the variables of the identified element and the model element are established as real-world differences, and a remediation action is proposed that will bring the identified element into proximity with the model element in the multidimensional space.
Individual steps of this process will now be described in more detail with reference to Figures 2 and 3. A specific example in the context of employee retention in an organisation will then be discussed in detail with reference to Figures 4 to 7.
The method shown in Figure 2 is necessarily computer-implemented - it requires a complex analysis of a large dataset, and it could not be implemented practically otherwise. The computer implementation works from real-world system inputs, and it establishes as an output remediation actions for elements of the real-world system to bring them into a positive state.
The first step is to establish 21 a set of variables for each of a plurality of elements of the real-world system. The purpose of this is to enable the key characteristics of the elements to be represented in the model used for analysis so that the model will have real predictive power. In some cases, it may be desirable to include attributes which are either invariable or not capable of change by a remediation action, but which may still have predictive power. For example, in the case of an employee database, salary,
job grade, job role and team size may all be variables, but age and gender may be taken as attributes.
The next step is to model 22 these elements in a multidimensional space. In the multidimensional space, the elements are represented by vectors determined by their variables (and, if used, attributes). This mapping on to a multidimensional space allows elements to be considered as clusters. Determination of these clusters is, as will be shown below, a highly significant part of this modelling process.
It is then determined 23 whether each of the elements is in a positive or a negative state. In some cases, an element will already have a defined state - for example, in the employee database example, some employees will have left the company (or left within a short timescale) and be identified as lost to attrition and so marked in the negative state, whereas others might be identified as long-time employees (for example, having remained with the company until retirement or beyond a long-service threshold) and placed in a positive state. As will be described in the detailed example provided below, other elements will be assessed as being in a positive or negative state by virtue of their resemblance to other elements, using both clustering and pattern matching approaches.
In some cases, there may be elements in one defined state but not in the other - for example, the only defined state may be the negative state (in the case of an employee database, an employee who has left the company for whatever reason), in which case assigning elements to a negative state may be a matter of identifying elements that are sufficiently similar to those that are in the negative state and placing the others in the positive state. In the case of an employee database, this would amount to identifying a cohort of “likely leavers” in the negative state by similarity to those who had already left.
It is also established 24 whether clusters are in a positive or a negative state, in the sense that they exceed a threshold value of elements in a positive state or elements in a negative state. One way to do this is to establish a critical percentage of negative elements, and to identify a cluster as a negative state cluster if the critical percentage is exceeded and as a positive state cluster if the critical percentage is not exceeded.
In this way, all clusters will be identified as being in either a positive state or a negative state.
Figure 3 illustrates clusters before determination of the state of all elements, showing elements in a final state that have a positive state 31 or a negative state 32, with the other elements initially being undetermined 33 (as discussed above, it may be that all elements other than those in a defined negative state may initially be considered as positive - for example, leavers and active employees). As can be seen, clusters 35 have been established. The undetermined elements 33 will then be determined as described. While it is not yet certain whether clusters will be positive or negative, it can be expected from the presence of elements with a determined state that the rightmost cluster 35a will be in a negative state whereas the topmost cluster 35b will be in a positive state, for example.
Elements are then identified 25 for remediation. This may be done, for example, by a ranking process, in which each of the elements that is not in a determined final state is ranked from“most negative” to“most positive”. These“most negative” elements will be those most in need of a remediation action.
The first step in determining the remediation action to be performed is to map 26 the element for remediation on to a model element. This model element is the nearest positive model element in the multidimensional space that is also located in a positive state cluster. This approach is chosen as it is minimally disruptive (it is the least distance in the multidimensional space) and as it stands the best chance of achieving the desired result (as there will be a track record of elements achieving a positive state as a final state in a positive state cluster).
When this mapping is achieved, the differences in the variables between the element for remediation and its model element are determined 27. Differences in attribute - as attributes cannot be changed by a remediation action - are not so significant at this point - though such differences may be determined at this stage and inappropriate remediation actions removed by filtering at a later point. These differences are then interpreted as real world differences - in the case of the employee database example,
these may include changes in salary, changes in seniority, changes in work group and other aspects of a job that could realistically be changed.
One or more changes to the variables of the element for remediation can in this way be identified 28 that would map it more closely to the model element. A suitable result would be one that, if implemented, would move the element from remediation from its current cluster to the cluster containing the model element. The real-world implementation of those changes is then proposed 29 as the remediation action.
A detailed implementation of this approach is described with reference to the employee attrition example with reference to Figures 4 to 7. This implementation is described in greater detail in Brockett, Clarke, Berlingerio and Dutta “CLARA: Clustering for Analysis and Remedial of Attrition” in ACM SIGKDD Ί9, August 04-08, 2019, Alaska, USA, the contents of which are incorporated by reference herein to the extent possible under applicable law. The implementation is described with reference to its performance on an employee data set as input data. This input data is historical by nature, that is, it contains employee records of current employees (hereafter “active”) of an organization, as well as of employees who have separated (hereafter “leaver”) from the organization.
The problem assessed is as follows: given a set of employees E who have been a part of the company during the last 12 months, each employee e, e E is represented by an associated feature vector based on the employee characteristics. That is, e, = (ei ei2, ·
· · , ein), where the features (eij ) include age, gender, highest education level, salary band, salary, length of service, length of previous experience, team and time since last promotion. Let EL E be the subset of employees who have left the organization in the past 12 months, and EA = E \ EL be the set of employees who are still active in the company.
The goal was established of ranking the active employees in EA by their descending likelihood of leaving, and, for each employee compute the least disruptive remedial action that would reduce the risk of attrition. These remedial actions are suggestions, but should be actionable - for example, a company can’t make people younger, but can surely move a person to a larger team. Also note, the trivial solutions of just
promoting every employee, or increase all salaries are neither valid nor practical. Promoting an employee too early may propagate the risk of attrition to the other team members, who may become upset or lose confidence in the company’s leadership. Similarly, increasing salaries for all employees is typically not feasible due to the budgetary constraints of an organization.
The objective is thus to find employees at the highest risk of attrition, and to identify least disruptive set of changes that would improve the situation. Assessment of if and to what degree the attrition situation has improved following the remedial actions is difficult given a static dataset, though an evaluation mechanism is also discussed further below.
Formally, given a set of employees E, a set of employee features f (like salary band and team size), and a number k of employees to retain (considered as the“budget”), we define attrition prevention problem as finding the least disruptive set of remedial actions on f maximizing the retention of k active employees with the highest risk of attrition.
Predictive scoring will now be considered. Predictive attrition scoring implements a novel core methodology to assess and predict the probability of voluntary attrition for an active employee. To this end, for each active employee, the system utilizes an aggregated scoring function based on (a) clustering - giving the employee’s proximity to other employees, and (b) frequent pattern mining - giving the similarity between any employee and the traits of a typical leaver. The score is then used to predict the chances of attrition for each individual, and subsequently recommend remedial steps to improve retention. The two major components of this level are the employee clustering module and the frequent feature pattern matching module, both of which are described below.
Employee Clustering - Based on the employee feature values, each employee record (active or leaver) is represented as a high-dimensional vector. Considering the dataset as a collection of high-dimensional employee data points, the system performs clustering to extract clusters or employee groups demonstrating a high degree of feature similarity. To be fair during comparisons against baselines that do not deal with
categorical variables natively, we use the k-means clustering algorithm based on the Euclidean distance (L2 norm) to obtain the cluster groups.
Considering the transformed input employee dataset E (obtained here from the feature selection module, discussed with reference to Figure 6 below) to consist of m active employees and n leavers, assume C = {C1, c2, , ck } to be the set of k employee clusters obtained as above. Observe that the number of clusters k is a model parameter affecting the performance of the framework - this can be varied to achieve optimal performance. Assume cluster c, to contain m, active and n, leaver employees, such that.
A cluster c, is referred to as a“leaver cluster” iff the ratio of the leavers to the active in the cluster is greater than a tuneable thresholding parameter T, i.e. , if n, /m, ³ T. Otherwise the cluster is termed as an“active cluster”. Note that within this framework we set the parameter t to the global ratio of leaver to active in the input dataset, i.e., t = n/(n + m). Let represent the feature vector of an active employee e,, and let the function o(b, ) = j return the cluster j where e, is found. For each such active employee e,, the system computes two feature similarity scores using the obtained clusters, one distance-based and one correlation-based. Distance based - Computes the total inverse Euclidean distance between the feature vector and the vectors of all other neighbouring employees assigned to the same cluster
c(ei). The measure also takes into account if the neighbouring employee is active or leaver, assigning a higher aggregated score to e, if it is closer to leavers. Formally,
Correlation based - Computes the total linear relationship between the employee feature vector and the feature vectors of all other neighbouring employees assigned to the same cluster o(b,). We use the Pearson correlation coefficient between the vectors, and factor in if the neighbouring point is active or leaver as in the distance-based score above. Formally, we have,
The clustering-based score assigned to an active employee e, is given by the average of these two scores:
From the above, we see that the higher the similarity of features of an active employee to (historical) leavers present within the same cluster, the higher is the computed score - thereby demonstrating a larger probability of possible attrition. Frequent Feature Pattern Matching - This involves analysing the patterns in the feature vectors present among employees that have already left the organization. Intuitively, an active employee having a high similarity to such extracted frequent patterns are at a higher risk of attrition - and should thus be captured by this measure.
In this setting, we consider the input employee dataset E to be divided into two mutually exclusive parts - EA and EL containing the m active and n leaver employees respectively. The system next extracts frequent patterns of employee features (for active and leaver) separately from the two partitions EA and EL using Eclat’s algorithm. For our experimental setup, the support threshold for the pattern finding algorithm is set at 5%, while the minimum itemset (pattern) size considered is 3.
Assume AF and LF to be the set of frequent patterns extracted from the active EA and leaver EL data partitions respectively. The commonality of frequent patterns between the active and leaver employees provides an important connecting factor in predicting active employee attrition. Hence, we consider the frequent feature patterns that are present in both AF and LF , and construct the set FP = AF P LF as the candidate discriminative feature patterns for our prediction model. However, observe that certain patterns are more informative than others, and hence should contribute with a higher weight in the process of prediction.
To take into account this importance factor (of patterns), the system ranks the frequent patterns in FP and assigns weights based on the computed ranks. Specifically, for each frequent pattern pj e FP, the relative frequencies of pj in EA and EL data partitions are computed. The relative frequency of a pattern is defined as its frequency of occurrence in the particular dataset over the total number of records in the dataset (expressed in percentage).
Intuitively, a pattern belonging to the intersection of the frequent patterns found in the active population and the frequent patterns found in the leavers population is more discriminative (in terms of predicting attrition) if its relative frequency (or support) is higher in one of the two population and significantly lower in the other partition. If each pattern belonging to said intersection is represented as a point in a scatter plot, the more discriminative patterns would be farther away from the diagonal bisector. Points along or close to the bisector instead represent near equal presence of the pattern in both datasets. These patterns are not particularly representative of either population.
We formalize this intuitition by computing the distance from a pattern having as coordinates (x,y) = ( relative frequency in the leavers, relative frequency in the active ), and the bisector y = x. We are interested in the patterns typically describing leavers, so instead of using the absolute, we subtract y from x, thus computing:
We then sort all patterns p by d and obtain a ranked list of patterns.
Using the above sorted list of frequent patterns, for each active employee ei, the system computes an aggregated score based on matching the employee features with the frequent patterns, taking into account the ranks assigned to the frequent patterns. Specifically, if a frequent pattern is matched in an employee feature vector, the pattern matching score of the employee is incremented by the inverse of the rank of the frequent pattern. Formally, we have,
Similar to the clustering scores, observe that an employee exhibiting a large number of discriminative frequent leaver patterns (featuring at the top of the ordered list) will exhibit a high pattern matching score - demonstrating a larger tendency towards probable attrition.
Ranking - The final predictive attrition score, PS, assigned by the system to an active employee e, is computed by a linear combination of the clustering score and the pattern matching score of the predictive scoring module. Our framework uses the averaged score, that is,
Other combinations can easily be used according to scoring requirements. The employee predictive attrition score provides a proxy for the probability of voluntary attrition of an active employee, and it is larger if the predicted chances of separation are higher. Hence, based on the total predictive scores, PS, the system generates an ordered list of active employees in the decreasing order of the scores. This provides the attrition rank list for the organization with probable candidates at the highest risk of attrition due to varying factors. It is for these employees that our framework subsequently recommends remedial actions to help improve their retention.
Cluster Purity - One significant consideration here is the nature of the clusters. An important tuneable parameter here is k, the number of clusters (or employee groups) to be created during the predictive attrition score computation phase. To this end, we empirically select the best value optimizing the performance of the system. The value of k was varied between 5 and 50 to select the setting that provides the best clusters in terms of“cluster purity”. Traditional clustering approaches tend to optimize the intra- and inter- cluster distances. However, for our problem setting, we prefer a clean segregation of leaver and active employees to reduce noise and capture high quality feature patterns within the clusters. We define“cluster purity” as a measure that captures and aggregates the ratio of leavers to noise (active employees) in leaver clusters and vice-versa for active clusters. Formally, consider Lj and Aj to be a leaver and active clusters respectively, obtained with the number of clusters k = j. The cluster purity is computed as
The final value of k is set to the one providing the maximum cluster purity, i.e. , k = arg maxj pure( k = j), where, j e [5, 50] It is found that the cluster purity initially increases and then tends to oscillate over a range of values. For the dataset used here, we set k = 23 providing a good trade-off between purity and computational efficiency.
Remedial Action Recommendation - A key novel feature of this framework is the recommendation of personalized remedial steps on an individual level to improve the rate of attrition - moving from predictive to remedial HR analytics. To this end, the above ranked list of employees (with high predictive attrition scores) are read in the order of their scores, and based on the employee feature clustering and the active employee feature, recommendations are generated and presented to the management.
Each employee record (active or leaver) is represented as a high-dimensional vector (based on the features), and the dataset is clustered to extract employee groups with a high degree of feature similarity. Consider an active employee e, (ranked high in the candidate employee rank list) to be assigned to cluster q. For e,, the system searches for the closest active employee that belongs to an active cluster. Since employees in an active cluster are possibly less prone to attrition, comparing the associated features with that of e, provides avenues to improve the retention of e,. Also note that this closest possible vector difference provides the least disruptive change, optimizing operational cost of an organization.
A measure of the deviation is computed between the feature vector of e, and that of the closest active employee - providing the basis for the remedial recommendations.
closest active employee in an active cluster.
Figure 4 shows an example of a remedial action in a cut-down system. For simplicity, we use a bi-dimensional space here: Salary Band, Team Size. We have 3 active clusters and 1 leaver cluster. Letters Ά’ mark active employee in a cluster, while we use ’L’ for leavers. In this image, the system needs to create a remedial action recommendation for the active employee in the bottom left cluster where the three vectors start from. The system considers all the active people in all the active clusters,
and finds the closest active to compute the remedial vector. We depict here two valid remedial actions, however one of them is the least disruptive (as it’s the shortest vector), the other one requires bigger changes in either dimension. We also depict with a dot-dashed line a non-valid remedial action: this vector would bring the active employee to an active in a leaver cluster, which is not our desired output.
This multi-dimension view of the computed feature vector difference is translated to a natural language based recommendation. For example, a salary difference in the feature vectors provides a simple remedial action of “Increase salary”, while a difference in the features salary and job grade might correspond to an “Increase responsibility” recommendation. Note, in certain cases a few remedial actions might not be physically possible or logically practical (e.g., move to smaller groups where there are no smaller groups, or decrease salary). Hence, a final layer of inexpensive human expert intervention for manual filtering might be required. This step is normally performed by HR at the time of browsing and applying the results. Incorporation of natural language processing and common sense knowledge extraction into our current framework provides an interesting direction of future work for creating a fully automated pipeline.
Figure 6 shows a more detailed breakdown of the modules of a computational system suitable for carrying out the method described in detail above. The data importer module imports the data needed by the system. Employee data, collected during the onboarding process or voluntarily disclosed by an individual, typically resides in an enterprise’s encrypted data store (dedicated data warehouse or cloud storage). The data importer module provides an interaction layer between the proposed framework and the raw storage. This layer is particularly dedicated to protecting the security and sensitivity aspects of the raw information, so that unauthorized access or leakage can be prevented.
Specifically, the data importer module enables raw data ingestion and its“masking” by performing the following steps: (i) reads the raw data from storage, (ii) decrypts the data, (iii) removes sensitive information pertaining to employees, and (iv) anonymizes employee details with random but unique identifiers. The last two steps are needed to enable maintenance, further development, and analysis by the developers.
Additionally, data cleaning involving duplicate removal and handling of corrupt/missing information is also performed by the data importer. This forms the final clean input data that is visible to the higher layers of the system architecture.
The feature selection and engineering module is equipped for pre-processing the input date received from the importer module, and for augmenting it with new features computed on top of the available ones. The different aspects present in the employee data that describe an employee, such as salary band, education, seniority, etc., are treated as employee features in this framework.
During the pre-processing stage, the different categorical features are identified and are represented appropriately (e.g., representation of feature seniority = D/M/E might reflect whether an employee is of level Director/Manager/Employee respectively) for uniformity and computational ease in the subsequent modules. Further, continuous features depicting a range of real values are suitably categorized - for example, employee salaries between 2000 - 3000 might be denoted by salary category S3.
Given the transformed data, this module then performs an analysis of the employee features across the entire dataset to evaluate the“goodness” of each of the obtained features based on its discriminative power for partitioning the data into (dis-)similar groups.
For example, the Pearson correlation coefficient may be used to identify and retain features (which may themselves be uncorrelated and demonstrate a high variance) for extracting meaningful patterns to predict individual employee attrition chances. Other feature selection techniques, such as Principal Component Analysis (PCA), can easily be adopted in this module, making the system dynamic to several different application domains and requirements. The final clean, categorized and selected employee features form the pre-processed input to the next stage.
For augmentation of an employee feature set, team-based features (specifically computed on new hires joining each team) can be considered. Team-based features include parameters like team size, average team seniority, difference between employee seniority and average team seniority, team diversity score, etc. On the other
hand, new-hire features include the team-based features, but these are re-computed for each new hire to highlight the evolution of a team and the hiring strategy of the company’s leadership.
The clustering, frequency pattern matching (FPM) and ranking modules are described in detail above, as is the remedial actions recommender. The final module provides the external interface API, here presenting to the HR and management a list of employees at a high risk of attrition along with the suggested remedial steps. It is to be noted that the recommendations are mere suggestions aimed at augmenting the understanding of probable reasons and steps against employee attrition in the organization. As such, proper chain of organizational structure, latest feedback from manager, overall performance, and other associated data pertaining to an employee would typically also be presented along with the remedial steps via a simple interface. In summary, all the available information regarding employees, their performances and achievements should be discussed and taken into consideration for gauging the feasibility of the remedial steps and the next course of action.
Figure 7 indicates the results of an evaluation on a historical dataset showing the effectiveness of an embodiment of the current disclosure (identified as CLARA) against other approaches. Other machine learning approaches used for comparison and variants of the current approach are the following:
Random Forests - Trained to classify an employee (using the features) to either be active / leaver in the near-future. The confidence of the classification is used to rank employees based on the predicted risk of attrition.
XGBoost Trees - Similar to the Random Forest approach, but instead the boost tree prediction model is used.
Support Vector Machines - Trained to classify employees into the binary category of active or leaver based on the different weights learnt on employee features. The predicted probability of an employee belonging to one of the classifications is used to rank currently active employees
K-Means Clustering (without FPM) - This provide an ablation study for our proposed model, by using only the clustering-based scoring module (without frequent pattern mining). Active employees are ranked based on the clustering score CS
Spectral Clustering - Provides a variation of the clustering technique used based on eigen-space decomposition. The distance and correlation-based score is computed between each pair of employee points to find the clusters. Frequent Pattern Mining (without K-Means) - The final ablation baseline considers only the frequent pattern overlap score PO to rank the active employees to obtain those at the highest risk of attrition.
As can be seen from Figure 7, the CLARA embodiment is markedly more successful in identifying the employees most at risk of attrition. Effect of remedial actions was considered by running the system on the entire data set, identifying remedial actions and applying them all and evaluating the data set to see (i) what happened to the highest ranked elements and (ii) how do the highest ranked elements in the new data set compare to those in the old data set (ie how does the attrition risk compare)? Note the subtle difference between i) and ii): the aim of the first is to“follow” the first x employee from R1 to R2. The aim of the second is to“follow” the first x ranks from R1 to R2. More formally, let PSR1 ( ) be function returning the score of an employee in R1 and PS R2 ( ) be the function returning the score of an employee in R2. Let also map ( ) be the function returning the rank that the argument has in R2. Lastly, let ±R1 be the employee with the lowest score (i.e. ranked last) in R1.
We can then define
as the average decrease (in percentage) of predictive attrition score of the first x employees in R1 when“followed” through the second ranking R2, and
as the average decrease (in percentage) of predictive attrition score of the people at the top x ranks in both rankings. Both scores are normalised by the score of the employee ranked last in R1 to manage expectations with respect to achievable actions.
On the dataset evaluated, there was found to be a 22.5% reduction in the first score if the top 5 remedial actions were applied and a 13% reduction if 10 actions were applied, with a reduction of 8% in the second score if we apply the first 5 actions, and a 5.4% reduction if we apply the top 10.
As discussed above, this approach is applicable to other real-world systems. Examples are briefly discussed below.
Embodiments of this approach may be used in a logistics system. Elements in the logistics system may be items to be transported, for example from a producer to a retailer. Here, the negative state of an element may be failure to meet conditions of service for the item, such as delivery within a particular time and at a particular level of quality. The variables may relate to location and condition of the items and the remediation action may comprise moving the item to a different location or changing the condition of the item. Condition here may include features such as temperature or exposure to sunlight - in the case of produce or perishable items, varying of such conditions may be as significant as speed of transit in ensuring that the items will be at a sale location at saleable quality for a sufficient length of time to meet conditions of service. Typically in a logistics system, information concerning an item to be transported - its nature and its current location and projected itinerary - will be stored, so it is then not a significant step to take this data and perform a regular evaluation of items most at risk of failing to meet conditions of service. The items most at risk could then be given a remediation action generated by the system to minimise the risk of breaching conditions of service.
In other implementations, the real-world system may be a task performance environment and each element is a task to be performed. This may involve, for example, carrying out many computational tasks on a distributed computing system.
The negative state may here be failure or predicted failure of a task, and the remediation action may be changing performance of the task to reduce the risk of failure. An implementation of the approach described here could be run across the full set of tasks to be carried out, and those at greatest risk of failure identified. These tasks may be re-allocated by an appropriate remediation action - changing their place in a queue, or assigning them to a particular computation engine, for example - to reduce this risk of failure.
In another type of implementation, the real-world system is a customer contract database, and each element is a customer contract. The negative state may then be termination or predicted termination of the contract by the customer, and the remediation action may comprise changing one or more contract terms or conditions to improve the likelihood of customer retention. This is relatively closely aligned to the employee retention case - the problem is now customer retention rather than employee retention - and the set of variables directly comparable.
As the skilled person will appreciate, all the embodiments described above are exemplary, and further embodiments falling within the spirit and scope of the disclosure may be developed by the skilled person working from the principles and examples set out above.
Claims
1. A computer-implemented method of determining a remediation action for an element of a real-world system, the method comprising:
establishing a set of variables for each of a plurality of elements of the real- world system;
determining clusters of elements based on proximity of elements in a space in which each element is represented by a vector comprising the set of variables for that element;
establishing whether each of the elements is in a positive or a negative state; establishing whether each cluster is in a positive state or a negative state based on the composition of that cluster of elements in a positive state and of elements in a negative state;
identifying elements in a negative state for remediation;
mapping each element identified for remediation on to a model element, being the nearest element in a positive state that is also in a cluster in a positive state;
determining differences in the variables between the element for remediation and its model element;
identifying one or more changes to the variables of the element for remediation that would map it more closely to the model element; and
determining the real-world implementation of those changes as the remediation action.
2. The method of claim 1 , wherein identifying one or more changes to the variables comprises identifying changes that would map the element into the cluster containing the model element.
3. The method of claim 1 or claim 2, further comprising establishing a set of invariable attributes for each element, wherein the vector for each element represents the set of variables and the set of attributes.
4. The method of any preceding claim, wherein the step of identifying elements in a negative state for remediation comprises ranking the elements in order of need of remediation.
5. The method of claim 4, wherein ranking comprises obtaining a clustering score for the element determined from the proximity of the element to elements in a positive state and to elements in a negative state.
6. The method of claim 4 or claim 5, wherein the ranking comprises obtaining a pattern matching score for the element determined from identification of positive patterns associated with a positive state and negative patterns identified with a negative state and matching the elements against the positive patterns and the negative patterns.
7. The method of claim 6 where dependent on claim 5, wherein the ranking comprises assigning each element a ranking score formed from a combination of the clustering score and the pattern matching score.
8. The method of any preceding claim, wherein determining clusters of elements comprises selecting clusters to optimise segregation of elements in positive state from elements in a negative state.
9. The method of any preceding claim, wherein the real-world system is an organisation and the elements are employees of the organisation, wherein the negative state of an element is departure of an employee, wherein the variables describe aspects of the job roles and employment conditions of the employees, and the remediation action comprises varying one or more of said aspects or employment conditions.
10. The method of any of claims 1 to 8, wherein the real-world system is a logistics system and the elements are items operated on by the logistics system, wherein the negative state of an element is failure to meet conditions of service for the item, wherein the variables relate to location and condition of the items and the remediation action comprises moving the item to a different location or changing the condition of the item.
11. The method of any of claims 1 to 8, wherein the real-world system is a task performance environment and wherein each element is a task to be performed, wherein the negative state is failure or predicted failure of a task and the remediation action is changing performance of the task to reduce the risk of failure.
12. The method of any of claims 1 to 8, wherein the real-world system is a customer contract database, wherein each element is a customer contract, and wherein the negative state is termination or predicted termination of the contract by the customer, wherein the remediation action comprises changing one or more contract terms or conditions.
13. A computing system comprising a processor and a memory, wherein the processor of the computing system is programmed to execute the method of any of claims 1 to 12 using the memory.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201962799807P | 2019-02-01 | 2019-02-01 | |
| US62/799,807 | 2019-02-01 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020157339A1 true WO2020157339A1 (en) | 2020-08-06 |
Family
ID=69645912
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2020/052643 Ceased WO2020157339A1 (en) | 2019-02-01 | 2020-02-03 | Remediation of real-world systems |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2020157339A1 (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110307413A1 (en) * | 2010-06-15 | 2011-12-15 | Oracle International Corporation | Predicting the impact of a personnel action on a worker |
| US20150269244A1 (en) * | 2013-12-28 | 2015-09-24 | Evolv Inc. | Clustering analysis of retention probabilities |
-
2020
- 2020-02-03 WO PCT/EP2020/052643 patent/WO2020157339A1/en not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110307413A1 (en) * | 2010-06-15 | 2011-12-15 | Oracle International Corporation | Predicting the impact of a personnel action on a worker |
| US20150269244A1 (en) * | 2013-12-28 | 2015-09-24 | Evolv Inc. | Clustering analysis of retention probabilities |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7570812B2 (en) | AI-driven transaction management system | |
| Adeniyi et al. | Automated web usage data mining and recommendation system using K-Nearest Neighbor (KNN) classification method | |
| Lael et al. | Use of data mining for the analysis of consumer purchase patterns with the fpgrowth algorithm on motor spare part sales transactions data | |
| US20220075762A1 (en) | Method for classifying an unmanaged dataset | |
| Meisel et al. | Synergies of operations research and data mining | |
| US8805836B2 (en) | Fuzzy tagging method and apparatus | |
| Kasemsap | Multifaceted applications of data mining, business intelligence, and knowledge management | |
| Tamilselvi et al. | An overview of data mining techniques and applications | |
| Akerkar | Advanced data analytics for business | |
| Zhou et al. | Resolution recommendation for event tickets in service management | |
| Maquee et al. | Clustering and association rules in analyzing the efficiency of maintenance system of an urban bus network | |
| Weng et al. | Mining time series data for segmentation by using Ant Colony Optimization | |
| Verma | Optimizing Database Performance for Big Data Analytics and Business Intelligence | |
| Otten et al. | Towards decision analytics in product portfolio management | |
| JP2025066082A (en) | Retention Management System | |
| Boyapati et al. | Predicting sales using Machine Learning Techniques | |
| Fan et al. | Spatially enabled customer segmentation using a data classification method with uncertain predicates | |
| Yan et al. | Dynamic matching algorithm of human resource allocation based on big data mining | |
| Preethi et al. | Data Mining In Banking Sector | |
| Mary et al. | Data mining and business intelligence trends | |
| Saraiya et al. | Study of clustering techniques in the data mining domain | |
| Bresciani et al. | Data management | |
| Marikkannu et al. | Classification of customer credit data for intelligent credit scoring system using fuzzy set and MC2—Domain driven approach | |
| Chaudhary | Data mining system, functionalities and applications: a radical review | |
| Mehdizade | BIG DATA ANALYTICS |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20706387 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20706387 Country of ref document: EP Kind code of ref document: A1 |








