EP4706227A1 - Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems - Google Patents

Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems

Info

Publication number
EP4706227A1
EP4706227A1 EP24726687.7A EP24726687A EP4706227A1 EP 4706227 A1 EP4706227 A1 EP 4706227A1 EP 24726687 A EP24726687 A EP 24726687A EP 4706227 A1 EP4706227 A1 EP 4706227A1
Authority
EP
European Patent Office
Prior art keywords
communications network
performance indicator
layer
output
timestamp
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24726687.7A
Other languages
German (de)
French (fr)
Inventor
Haoyu LIU
Alec DIALLO
Samuel Knight
Marco Fiore
Paul PATRAS
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Net Ai Tech Ltd
Original Assignee
Net Ai Tech Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Net Ai Tech Ltd filed Critical Net Ai Tech Ltd
Publication of EP4706227A1 publication Critical patent/EP4706227A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/14Network analysis or design
    • H04L41/147Network analysis or design for predicting network behaviour
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/16Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks using machine learning or artificial intelligence
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04WWIRELESS COMMUNICATION NETWORKS
    • H04W24/00Supervisory, monitoring or testing arrangements
    • H04W24/04Arrangements for maintaining operational condition
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/06Management of faults, events, alarms or notifications
    • H04L41/0631Management of faults, events, alarms or notifications using root cause analysis; using analysis of correlation between notifications, alarms or events based on decision criteria, e.g. hierarchy, tree or time analysis
    • H04L41/064Management of faults, events, alarms or notifications using root cause analysis; using analysis of correlation between notifications, alarms or events based on decision criteria, e.g. hierarchy, tree or time analysis involving time analysis
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L41/00Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
    • H04L41/14Network analysis or design
    • H04L41/142Network analysis or design using statistical or mathematical methods

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Biophysics (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Biomedical Technology (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

There is disclosed a computer implemented method of training an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) leveraging a data-driven approach to generate in an unsupervised way baselines describing normal daily patterns of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate), followed by a measure of deviation from the baseline to label each segment of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate); (ii) using a data augmentation technique that increases both the volume and the variability of anomalous performance indicator values, including taking as input normal/benign time series and randomly scaling them, with a scaling factor sampled from a pre-set uniform distribution; (iii) utilizing a Seq2Seq network to encode the historical data and forecast the future, further fusing the historical information with predicted upcoming performance indicator values, which is supplied to the final classifier; (iv) to achieve both classification and forecasting simultaneously, training the neural models involved with respective first (e.g. Cross Entropy (CE)) and second (e.g. Mean Squared Error (MSE)) loss functions, employing a balancing factor.

Description

METHODS OF TRAINING AN ARTIFICIAL INTELLIGENCE MODEL FOR OPERATIONAL ANOMALY PREDICTION IN A COMMUNICATIONS NETWORK, AND SYSTEMS BACKGROUND OF THE INVENTION 1. Field of the Invention The field of the invention relates to computer implemented methods of generating training data for an artificial intelligence model for operational anomaly prediction in a communications network, to computer implemented methods of training an artificial intelligence model for operational anomaly prediction in a communications network, to computer implemented methods of operational anomaly prediction in a communications network, and to related computer systems. 2. Technical Background Mobile operators currently encounter inefficient infrastructure management and high costs when it comes to hiring specialized engineering staff. As a result, as much as 65.3% of these operators’ revenue goes into operating costs, as discussed for example in M. Newman, Global trends in telco OpEx, 2021. Operators face challenges when responding to network performance problems manually, using weak static thresholds and ineffective time granularity. Due to these flaws, service quality is impacted, with failures or misconfigurations not addressed in a timely manner or implemented changes not matching network conditions. Customer churn may be high, as users often feel dissatisfied with their service. Operational anomalies in communications networks may be caused, for example, by equipment failure or radio communication interrupted by weather; by network service outage caused by misconfiguration or an accidentally disconnected network cable; or an external attack where an adversary tries to disable the network. What is needed is improved prediction of operational anomalies in communications networks. 3. Discussion of Related Art EP3346666(A1) and EP3346666(B1) disclose a prediction system configured for modelling the expected number of attacks on a computer or a communication network using data obtained by sensors of an attack monitoring system. The prediction system is capable of modelling the attack rates and predicts the expected number of attacks at different resolutions, granularity levels and horizons. The core modelling module of the system requires basic and minimal information about the observed attacks which makes the prediction system applicable for a variety of attack monitoring systems such as low, medium and high interaction honeypot systems or network telescopes. EP3357271(A1) and EP3357271(B1) disclose a method for managing a wireless network, comprising: - collecting a sequence of traffic data samples ordered in time, and arranging said collected data samples in at least one level-0 residual matrix having at least one dimension, said dimension of said level-0 residual matrix corresponding to a respective time scale comprising an ordered sequence of time units, said ordered sequences of time units defining a first time window; - performing at least once a cycle, each n-th iteration of the cycle, starting from n = 0, comprising a sequence of phases A), B), C), D), E): A) for at least one dimension of a level-(n) residual matrix, subdividing the corresponding time scale in such a way to group the time units thereof in a respective level-(n+1) partition of time units so as to subdivide the traffic data samples in corresponding level-(n+1) traffic data sample sets; B) for each level-(n+1) traffic data sample set, calculating a corresponding functional which fits said level-(n+1) traffic data sample set; C) for each level-(n+1) traffic data sample set, calculating a corresponding approximation of the level-(n+1) traffic data sample set by applying the corresponding functional to the corresponding level-(n+1) partition of time units; D) joining together the approximations of the level-(n+1) traffic data sample sets to calculate a level-(n+1) approximated matrix, said level-(n+1) approximated matrix being an approximated version of the level-(n) residual matrix; E) calculating the difference between the level-(n) residual matrix and the calculated level-(n+1) approximated matrix so as to obtain a level-(n+1) residual matrix; - forecasting traffic data trend in a second time window different from the first time window by generating predicted data samples by applying the calculated functional to a partition of time units comprising an ordered sequence of time units corresponding to at least one among said second time window and said first time window; - using said forecasted traffic data trend to manage the wireless network.
SUMMARY OF THE INVENTION According to a first aspect of the invention, there is provided a computer implemented method of generating training data for an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) receiving a communications network historical traffic timeseries which includes one or more performance indicators (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) as a function of timestamp; (ii) for each timestamp, computing respective timestamp features denoting relative position in time, e.g. within a year, the month, week of the month, weekday, hour, minute and second; (iii) generating a function which receives timestamp features as inputs and outputs the closest performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) at the same time (e.g. the same weekday, the same hour, the same minute and the same second), for any predefined time period (e.g. week) except for the predefined time period (e.g. week) in the input timestamp features; (iv) using the function generated in (iii), labelling each chunk of the performance indicator (e.g. communications network historical traffic, accessibility, physical resource block (PRB) utilisation, call drop rate) timeseries based on deviation of the chunk from a baseline, e.g. using a Wasserstein distance to measure the deviation using the function generated in (iii); and (v) storing training data including the label of each chunk of the performance indicator (e.g. communications network historical traffic, accessibility, physical resource block (PRB) utilisation, call drop rate) timeseries. An advantage is improved training data for an artificial intelligence model for operational anomaly prediction in a communications network. The method may be one in which the one or more performance indicators is or includes communications network traffic volume, accessibility, physical resource block (PRB) utilisation, or call drop rate. The method may be one in which the predefined time period is week, hour, or several hours, or a plurality of hours. The method may be one wherein a baseline is a common profile of the performance indicator of the historical timeseries, that describes how the performance indicator normally evolves over each day, month or year. The method may be one wherein the baseline is a daily historical average. The method may be one wherein the baseline is generated by a neural network, e.g. which incorporates some hidden features of the performance indicator timeseries, such as seasonality. An advantage is improved training data for an artificial intelligence model for operational anomaly prediction in a communications network. The method may be one wherein the function generated in step (iii) is used to produce a baseline. The method may be one wherein step (iii) includes: (a) Receiving the timestamp features as inputs and extracting features of the timestamp features, and forwarding the extracted timestamp features to a MLP layer to generate first MLP output, and forwarding the extracted timestamp features to a first FC layer to generate first FC layer output, and forwarding the extracted timestamp features to a second FC layer to generate second FC layer output; (b) Performing element-wise product transformation on the first MLP output and on the first FC layer output to generate element-wise product output; (c) Performing element-wise addition transformation on the element-wise product output and on the second FC layer output to generate element-wise addition output; (d) Inputting the element-wise addition output to the MLP layer to generate second MLP output, and evaluating a (e.g. MAE) loss of the second MLP output compared to performance indicator value (e.g. communications network traffic volume) for the timestamp which corresponds to the timestamp features; (e) Repeating steps (b) to (d), while modulating the first FC layer and the second FC layer, until a (e.g. MAE) loss minimization criterion is satisfied. An advantage is improved training data for an artificial intelligence model for operational anomaly prediction in a communications network. The method may be one including storing weights of the first FC layer and of the second FC layer. The method may be one wherein the MLP layer, the first FC layer and the second FC layer are independent. The method may be one wherein the modulation is an element-wise affine transformation on the first MLP output. The method may be one wherein the first FC layer output and the second FC layer output^ influence the output of the MLP by scaling and shifting each intermediate feature based on the first FC layer input and on the second FC layer input. The method may be one including constructing a first sample of the training set, the first sample including P steps up to the present of the training set, and repeating for each P; constructing a second sample of the training set, the second sample including a future F steps of the training set, and repeating for each F; identifying a random variable sampled from a uniform distribution, and scaling each pair of the first sample and the second sample using the random variable, to create a scaled set of pairs of the first sample and of the second sample, and performing steps (iv) and (v) on the scaled set of pairs of the first sample and of the second sample, to label each pair of the scaled set of pairs, and to store augmented training data including the label of each pair of the scaled set of pairs. An advantage is improved training data for an artificial intelligence model for operational anomaly prediction in a communications network. The method may be one wherein the variability of anomalous performance indicator values (e.g. network traffic) from which the neural networks learn is enhanced. The method may be one wherein the volume and the variability of anomalous traffic is increased. The method may be one wherein normal/benign time series are input, which are randomly scaled, using a scaling factor sampled from a pre-set uniform distribution. The method may be one wherein the scaling results in larger deviations from the baseline, and thereby more synthetic anomalies are generated. According to a second aspect of the invention, there is provided a storage medium storing the training data or the augmented training data of any aspect of the first aspect of the invention. According to a third aspect of the invention, there is provided a computer implemented method of training an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) receiving training data including a communications network historical performance indicator timeseries which includes performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) as a function of timestamp, and a label of each chunk of the communications network historical performance indicator timeseries; (ii) processing the training data using a positional encoding layer which encodes the timestamp of each datum and fuses that with a corresponding performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) to produce processed training data; (iii) encoding the processed training data in a first set of LSTMs as an encoded representation; (iv) passing the encoded representation to a decoder including a second set of LSTMs which forecasts performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) of next steps to produce a future performance indicator value forecast; (v) outputting the encoded representation from the first set of LSTMs via a first FC layer to a first FiLM layer; (vi) outputting the future performance indicator value forecast from the second set of LSTMs via a second FC layer to a second FiLM layer; (vii) the first FiLM layer fusing the received encoded representation with received input conditional features (e.g. the mean, standard deviation, maximum, minimum, and skewness of the communications network historical performance indicator timeseries) to generate output of the first FiLM layer; (viii) the second FiLM layer fusing the received future performance indicator value forecast with received target conditional features to generate output of the second FiLM layer; (ix) concatenating the output of the first FiLM layer and the output of the second FiLM layer to produce a concatenated vector; (x) evaluating a first (e.g. MSE) loss function of the future performance indicator value forecast compared to actual future performance indicator value, and passing the concatenated vector through a MLP and evaluating the MLP output using a second (e.g. BCE, e.g. including squared error (SE)) loss function which compares predicted labels with ground truth labels, and (xi) adjusting the first FC layer and the second FC layer while repeating steps (v) to (x) until a convergence criterion with respect to the first loss function and the second loss function is satisfied, to produce a trained model. An advantage is an improved artificial intelligence model for operational anomaly prediction in a communications network. The method may be one in which the performance indicator is communications network traffic volume, accessibility, physical resource block (PRB) utilisation, or call drop rate. The method may be one wherein the training data includes training data of any of aspect of the first aspect of the invention, or wherein the training data includes augmented training data of any respective aspect of the first aspect of the invention. An advantage is an improved artificial intelligence model for operational anomaly prediction in a communications network. The method may be one including storing weights of the trained model. The method may be one wherein anomaly prediction is treated as a combined task of both classification and forecasting. The method may be one wherein a prediction is made whether an anomaly will occur (e.g. soon) in the network, while forecasting is also needed since an ‘anomaly’ in our task is defined as a period of performance indicator values (e.g. traffic) that deviates from the baseline, i.e., forecasting the trend of upcoming performance indicator values (e.g. traffic) plays an important role in the decision-making process. The method may be one wherein a Seq2Seq network is used to encode the historical data and forecast the future performance indicator values (e.g. traffic), further fusing the historical information with predicted upcoming traffic, which is supplied to the final classifier. The method may be one wherein a Seq2Seq architecture is used, with a stacked LSTM block used as the encoder. The method may be one wherein to achieve both classification and forecasting simultaneously, the neural models involved are trained with Cross Entropy (CE) and Mean Squared Error (MSE) loss functions, employing a balancing factor. The method may be one wherein predictions are based on the provided input and the predicted future performance indicator values (e.g. traffic) together. The method may be one wherein a positional encoding layer is added before the encoder, which encodes the timestamp of each datum and fuses that with the original input. The method may be one wherein a timestamp is encoded as a d-dimensional vector in which d elements have varying frequencies depending on the index of the vector. The method may be one wherein conditional features are features obtained solely from the associated timeseries data. The method may be one wherein FiLM acts as a feature-wise affine transformation on its inputs. According to a fourth aspect of the invention, there is provided a storage medium storing the weights of the trained model of any aspect of the third aspect of the invention. According to a fifth aspect of the invention, there is provided a computer system, including a trained artificial intelligence model for operational anomaly prediction in a communications network produced by a method of any aspect of the third aspect of the invention. An advantage is a computer system including an improved artificial intelligence model for operational anomaly prediction in a communications network. According to a sixth aspect of the invention, there is provided a computer implemented method of operational anomaly prediction in a communications network, including the step of receiving input of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) time series data into the computer system including the trained artificial intelligence model for operational anomaly prediction in a communications network of the fifth aspect of the invention, and outputting from the computer system including the trained artificial intelligence model an operational anomaly prediction for the communications network. An advantage is improved operational anomaly prediction in a communications network. The method may be one in which associations between antennas and upper-level network nodes of the communications network are determined by using a k- partitioning algorithm over a Delaunay graph of antennas. The method may be one including computing rates for true positive, true negative, false positive and false negative. According to a seventh aspect of the invention, there is provided a computer implemented method of training an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) leveraging a data-driven approach to generate in an unsupervised way baselines describing normal daily patterns of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate), followed by a measure of deviation from the baseline to label each segment of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate); (ii) using a data augmentation technique that increases both the volume and the variability of anomalous performance indicator values, including taking as input normal/benign time series and randomly scaling them, with a scaling factor sampled from a pre-set uniform distribution; (iii) utilizing a Seq2Seq network to encode the historical data and forecast the future, further fusing the historical information with predicted upcoming performance indicator values, which is supplied to the final classifier; (iv) to achieve both classification and forecasting simultaneously, training the neural models involved with respective first (e.g. Cross Entropy (CE)) and second (e.g. Mean Squared Error (MSE)) loss functions, employing a balancing factor. An advantage is an improved artificial intelligence model for operational anomaly prediction in a communications network. The method may be one including a method of any aspect of the first aspect of the invention, and/or any aspect of the third aspect of the invention. According to an eighth aspect of the invention, there is provided a computer system including a trained artificial intelligence model for operational anomaly prediction in a communications network, trained using a method of any aspect of the seventh aspect of the invention. An advantage is a computer system including an improved artificial intelligence model for operational anomaly prediction in a communications network. According to a ninth aspect of the invention, there is provided a computer implemented method of operational anomaly prediction in a communications network, including the step of receiving input of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) time series data into the computer system including the trained artificial intelligence model for operational anomaly prediction in a communications network of the eighth aspect of the invention, and outputting from the computer system including the trained artificial intelligence model an operational anomaly prediction for the communications network. An advantage is improved operational anomaly prediction in a communications network. Aspects of the invention may be combined.
BRIEF DESCRIPTION OF THE FIGURES Aspects of the invention will now be described, by way of example(s), with reference to the following Figures, in which: Figure 1 shows an example structure of a baseline generation model. In this Figure, MLP is a multilayer perceptron layer and FC is a fully connected layer; MAE loss is a Mean Absolute Error loss function. Figure 2 shows an example structure of a predictive alert model. In this Figure, MLP is a multilayer perceptron layer and FC is a fully connected layer; LSTM is a Long short-term memory network; MSE loss is a Mean Squared Error loss function; BCE is Binary Cross Entropy loss function; SE is Squared Error loss because when computing the loss for a single event, it is just (y - ŷ)^2, so there is no need to calculate the mean. Figure 3 shows an example of a segment of an actual traffic timeseries and the associated neural network (NN)-generated baseline (dashed line). Black, red, orange and blue represents True Negative (TN), True Positive (TP), False Positive (FP), and False Negative (FN), respectively. Figure 4 shows an example of a single FiLM layer for a CNN. The dot signifies a Hadamard product. Various combinations of γ and β can modulate individual feature maps in a variety of ways.
DETAILED DESCRIPTION Deep Learning Pipeline for Predictive Operational Anomaly Detection in Communications Networks Herein, Seq2Seq, or Sequence To Sequence, is a model used in sequence prediction tasks; such models are used in elsewhere in language modelling and/or in machine translation. In an example Seq2Seq, the approach is to use a Long short-term memory (LSTM), the encoder, to read the input sequence one timestep at a time, to obtain a large fixed dimensional vector representation (a context vector), and then to use another LSTM, the decoder, to extract the output sequence from that vector. The second LSTM may be thought of as a recurrent neural network language model except that it is conditioned on the input sequence. Herein, we refer to a multilayer perceptron (MLP) layer and to a fully connected (FC) layer. Herein, the Wasserstein distance, sometimes referred to as the Kantorovich– Rubinstein metric, is a distance function defined between probability distributions on a given metric space. Herein, MEC stands for mobile edge computing; C-RAN stands for cloud radio access network. In computational geometry, a Delaunay triangulation of a set of points in the plane subdivides their convex hull into triangles whose circumcircles do not contain any of the points. This maximizes the size of the smallest angle in any of the triangles. Although we state that the timestamps of snapshots can be represented by a list with seconds granularity, the granularity used could be different such as with minutes granularity or hours granularity, as would be clear to the skilled person. 1 Introduction Problem formulation: Consider an antenna/computing facility that handles traffic gener- ated by a particular service ^^. The traffic generated by service at the current timestep ^^ is denoted by ^^ ^^ . Then ^^ ^^+1, ^^ ^^+2, ..., ^^ ^^+ ^^ is the traffic timeseries in the next ^^ steps with a label ^^ ^^∼ ^^ ∈ {0, 1} denoting whether or not the segment of traffic is anomalous. Anomaly prediction seeks to determine the label of the traffic in the upcoming ^^ steps, by leveraging historical information { ^^ ^^− ^^+1, ^^ ^^− ^^+2, ..., which is formally expressed as: arg Traffic anomalies, i.e., the traffic trends that significantly deviate from historical obser- vations are difficult to predict, since they do not follow common periodic patterns and are therefore unlikely to be learned by any machine learning (ML)/deep learning (DL) models directly in a forecasting task. Traditional anomaly detection approaches are themselves unable to solve this particular issue, as they have to observe the actual network traffic, to determine whether anomalies exist; by the time the mobile operator realizes that e.g. an abnormal surge in network traffic has occurred, a Service Level Agreement (SLA) violation may be inevitable. Solution: To solve this problem, we have developed IdentifAI, a DL-based pipeline that in an example predicts the probability of observing operational network anomalies in the near future, effectively providing insights and a time buffer for mobile operators to manage their networks in advance. The pipeline includes or consists of the following key components: 1. Data Labeling Module: anomalies in time-series data sometimes have various defini- tions, depending on the degree of deviation of each datum or the persistence of ab- normal behaviors. To obtain a clear and unified objective for classification/prediction, we first leverage a data-driven approach to generate in an unsupervised way baselines describing the normal daily patterns of network traffic, followed by a measure of deviation from the baseline to label each segment of network traffic. 2. Anomaly Augmentation Module: Compared with normal traffic, anomalies by their nature only account for a small proportion of samples in a dataset, which poses difficulty for neural models in identifying such occurrences. Traditional techniques, such as oversampling, only alter the proportion of the imbalanced class (anomaly) but does not produce distinct, and potentially new samples. This leads to less-than- desirable detection performance. In an example, we overcome this issue by introducing a data augmentation technique that increases both the volume and the variability of anomalous traffic. Specifically, the module may take as input normal/benign time series and randomly scale them, with a scaling factor sampled from a pre-set uniform distribution. The scaling may result in larger deviations from the baseline, and thereby more synthetic anomalies are generated. 3. Forecasting-aided Seq2Seq Anomaly Prediction Engine: We treat anomaly prediction as a combined task of both classification and forecasting. The classification task is straightforward to understand as the model aims to predict whether anomaly would occur soon in the network, while forecasting is also needed since an ‘anomaly’ in our task is defined as a period of traffic that deviates from the baseline, i.e., forecasting the trend of upcoming traffic plays an important role in the decision-making process. In an example, we utilize a Seq2Seq network to encode the historical data and forecast the future, further fusing the historical information with predicted upcoming traffic, which is supplied to the final classifier. 4. Multi-objective Optimization: To achieve both classification and forecasting simul- taneously, we train the neural models involved with Cross Entropy (CE) and Mean Squared Error (MSE) loss functions, employing a balancing factor. 2 Framework Design In what follows, we detail the inner working of the IdentifAI framework. 2.1 Baseline Generation and Data Labeling Data labeling is usually based on manual configurations by mobile operators, given that they may particularly attach importance to some abnormal patterns (e.g. persistently low/high traffic volumes, high call drop rates, etc.) over others factors (e.g. network traffic spikes and jitter). However, it is not always straightforward to obtain labels directly from timeseries data without additional support from field technicians. We offer a pure data-driven approach to automated labelling of timeseries data, that includes or consists of two stages, namely, baseline generation and deviation-based labelling. Baseline Generation. A baseline is considered as a common profile of historical traffic timeseries, that describes how traffic normally evolves over each day, month or year. We generate this baseline by a neural network, which incorporates some hidden features of traffic timeseries, such as seasonality. This can be achieved by approximating a function: ^^ ^^ = ^^( ^^) which takes as input a timestamp ^^ and outputs the most likely traffic volume at this timestamp, without actually observing it. We first introduce how we generate a dataset ({ ^^, ^^}) for this function, and then explain why by this approach a neural network that approximates this function can learn a baseline. Assume we have a dataset that logs the traffic volume over ^^ days (where ^^ = ^^/7 weeks), each data including or consisting of ^^ steps. Then, the timestamps of all the snapshots can be represented by a list with seconds granularity: { } 60 × 60 × 24 0, , 2 × 60 × 60 × 24 , ..., 60 × 60 × 24 × ^^ . ^^ ^^ For each timestamp ^^ ^^ above we compute the a series of features to describe its relative position in a year, including second ( ^^ ^^ mod 60), minute ( ^^ ^^/60 mod 60), hour ( ^^ ^^/(60× 60) mod 24), day of a week, week of a month, month, semester and quarter, which becomes the input of a function ^^(·) at timestamp ^^ ^^. For target values, we denote the traffic dataset by a ^^ × 7 × ^^ matrix. At timestep ^^ ^ ^ ^^, ^^ (at week ^^ , day ^^ and step ^^), the target value is the closest traffic volume at the same step, the same day of a week, in the rest of the weeks, i.e., ^^ ^ ^ ^^, ^^ = ^^ ^^, ^^ ^^ s.t., arg min | ^^ ^ ^ ^^, ^^ − ^^ ^ ^ ^^, ^^ |, ^^ That said, we approximate a function which takes timestamp features as input and outputs the closest data point at the same timestep in the rest of the weeks. It is expected that most of the data are not anomalies so that the closest data point does not differ much from the traffic volume at the target timestep. If the target timestep happens to observe an anomaly, the possibility that the closest data point is also an anomaly is relatively small, avoiding the neural networks from fitting extreme patterns. We design a baseline generation model, for example as depicted in Figure 1 to approx- imate function ^^(·). In an example, the model receives the timestamp features extracted above as inputs, forwarding to three independent layers (a MLP layer and two FC layers) respectively. Denote the output of the MLP as x ^^ ^^ ^^, and the outputs of FC layers as a and b. The network modulates an element-wise affine transformation on x ^^ ^^ ^^ by a⊙ x ^^ ^^ ^^ +b. a and b can influence the output of the MLP by scaling and shifting each intermediate feature based on the original input. In this case, the network is more flexible and less likely to be trapped in mediocre local optima during training. This model is optimized with a Mean Absolute Error (MAE) function that minimizes the gap between outputs and target baseline values. Labeling by Deviation. We label each chunk of timeseries based on how much ‘deviation’ it has from the baseline, which can be defined in several ways. Here, we employ a method based on Wasserstein distance to measure the said ‘deviation’, for which we leverage the daily historical average as the baseline for simplicity. Denote the chunk of traffic at day ^^ between ^^ + 1 and ^^ + ^^ as X ^ ^ ^^ ^^ := { ^^ ^ ^^ ^+1, The collection of the same period of traffic across all days in the training data is expressed as X ^^∼ ^^ := {X ^^ 1 ^^ , ...,X ^ ^ ^^ ^^ , ...,X ^^ ^^ ^^}. B ^^∼ ^^ indicates the associated chunk of timeseries in the baseline. Define a function ^^ ^^ ^^ ^^ (·, ·) : → R that measures the Wasserstein distance between two vectors. between ea chunk in X ^^∼ ^^ and B ^^∼ ^^ yields a list of distances: ^^ ^^∼ ^^ := represents the standard deviation of the distances between actual traffic and the baseline during ^^+1 ∼ ^^+ ^^. Each window of the timeseries is labeled with the following conditional equation: ^^ ^^ ^^∼ ^^ > ^^, ∀ ^^ ∈ ^^, otherwise, in which ^^ is a global hyperparameter that controls the proportion of anomalies. Note that the standard deviation ^^ ^^∼ ^^ is only acquired based on the training data ( ^^ days) and used to label test data directly without re-computation. 2.2 Anomaly Augmentation We introduce an anomaly augmentation technique to be used during training, in order to enhance the variability of anomalous traffic from which the neural networks can learn. This is independent from the labeling method introduced previously and flexible in controlling the anomaly ratio in the training set. Denote ( ^^ ^^− ^^∼ ^^, ^^ ^^∼ ^^), ^^ ^^∼ ^^ ∈ ^^ ^^ ^^ ^^ ^^ ^^, ^^ ^^ ^^ ^^ ^^ ^^ as a sample in the training set, where ^^ ^^− ^^∼ ^^ the input timeseries with P steps, ^^ ^^∼ ^^ the future F-step timeseries, and ^^ ^^∼ ^^ the label of the future timeseries. Let ^^ ∼ U(1 − ^^, 1 + ^^) be a random variable sampled from a uniform distribution. For each ( ^^ ^^− ^^∼ ^^, ^^ ^^∼ ^^) pair, we scale both of them as and re-label the scaled future timeseries chunk by The operation is repeated for every sample in the training set and ( ^˜^ ^^− ^^∼ ^^, ^˜^ ^^∼ ^^), ^˜^ ^^∼ ^^ is fed for training instead. Note that ^^ ^^∼ ^^ is the standard deviation of the distances computed based on the original traffic. It is potentially more easy to label the scaled timeseries as anomalous, while scaling the input timeseries equally provides the model a relatively consistent view between the past and the future. In practice, we set ^^ = 0.5 and find our augmentation technique notably effective. 2.3 Anomaly Prediction Engine The anomaly prediction engine in IdentifAI leverages a combination of neural networks to achieve accurate predictions of anomalies. Since a label is associated with a chunk of timeseries data instead of a single datum, it is crucial to first forecast the upcoming traffic, and then make predictions based on the provided input and the predicted future traffic together. In our design we build on a Seq2Seq architecture, with a stacked LSTM block used as the encoder, for example as shown in Figure 2. Different from the simpler architectures, a positional encoding layer is added before the encoder, which encodes the timestamp of each datum and fuses that with the original input. The encoding function is expressed as follows: where ^^ is the timestamp of a datum, ^^ the encoded dimension of the timestamp, and ^^ denotes the ^^-th element in the encoded timestamp. In other words, a timestamp is encoded as a ^^-dimensional vector in which ^^ elements have varying frequencies depending on the index of the vector. Positional encoding was originally designed for Transformers, which assign order to each word in a sequence. Although LSTM does not struggle with the relative order of data in a sequence, the model has no access to the timestamp of a day wrt. a chunk of input, thus may experience difficulty in telling apart the data captured in the morning and in the evening, where positional encoding can address this issue effectively. The encoded representation is passed to a decoder which forecasts traffic of the next ^^ steps (referred to as ‘output timeseries’ in the rest of this document). The encoded input and the encoded output timeseries are fed to two FiLM layers respec- tively, which fuse these with the conditional features. Conditional features in our task are features obtained solely from the associated timeseries data. For example, the mean, stan- dard deviation, maximum, minimum, and skewness of the input timeseries are considered as conditional features of the input. Since these features are highly correlated with the input but less relevant to the output, we leverage the FiLM layer to explicitly fuse them with the encoded input timeseries, preventing the model from exploring meaningless relationships between features of the input and the encoded output timeseries. Specifically, denote con- ditional features of input timeseries as c ^^, and let ^^ and ℎ be two arbitrary functions that output ^^ dimensional vectors: Let F ^^ ∈ R ^^ be an intermediate representation of the input timeseries in ^^-dimensional space. FiLM acts as a feature-wise affine transformation on F ^^ via: ^^ ^^ ^^ ^^ (F ^^ | ^^ ^^, ^^ ^^) = ^^ ^^ ⊙ F ^^ + ^^ ^^, where ⊙ represents element-wise product. FiLM is also applied to the encoded output timeseries, followed by a concatenation of two FiLM outputs. In an example, target features are computed in the same way as input features, with the difference being that target features are derived from predicted, future timeseries data, whereas input features from historical timeseries. The concatenated vector includes or consists of not only historical information, but also predicted future information for the final classification. 2.4 Loss Functions In an example, to train this model, we adopt two loss functions simultaneously. Denote ˆ ^^∼ ^^ = { ˆ ^^+1, ˆ ^^+2, ..., ˆ ^^+ ^^} as the predicted future traffic and ^^ ^^∼ ^^ = { ^^ ^^+1, ^^ ^^+2, ..., ^^ ^^+ ^^} the actual future traffic. The Seq2Seq part of the model is trained with a Mean Squared Error (MSE) function: Let ^ˆ^ ^^∼ ^^ ∈ [0, 1] be the predicted label and ^^ ^^∼ ^^ ∈ {0, 1} the ground truth label. The entire model including the Seq2Seq part is trained via Binary Cross Entropy (BCE) for the classification task: ^^ ^^ ^^ ^^ ( ^ˆ^ ^^∼ ^^ , ^^ ^^∼ ^^) = ^^ ^^∼ ^^ ^^ ^^ ^^( ^ˆ^ ^^∼ ^^) + (1 − ^^ ^^∼ ^^) ^^ ^^ ^^(1 − ^ˆ^ ^^∼ ^^) 3 Preliminary Evaluations 3.1 Datasets We evaluated our model on two large-scale real-world mobile network datasets collected in two major European cities, each representative of deployments with over 700 antennas, and monitoring over 20 mobile services. To evaluate the performance of our model under different network architectures, we consider three extra cases apart from antenna-level analyses: (1) 50 MEC facilities deployed at the edge handling traffic betweem 10–20 antennas; (2) 30 C-RAN datacenters, handling the traffic from 20–40 antennas; (3) 10 core datacenters (DC) handling the traffic from over 50 antennas. Following the approach in C. Zhang, M. Fiore, and P. Patras, “Multi-service mobile traffic forecasting via convolutional long short-term memories,” IEEE International Symposium on Measurements & Networking (M&N), 2019, we determine the associations between antennas and upper-level network nodes by using a k-partitioning algorithm over the Delaunay graph of antennas. We keep weekday traffic in both datasets for evaluations due to the distinct temporal patterns that exist over weekends. Predicting anomalies for weekend traffic is possible but requires a dataset over a longer time span for acceptable performance. We further split weekday traffic into a training set (40 days) and a test set (15/20 days). 3.2 Evaluation Metrics We formulate anomaly prediction as a classification task and therefore use F1 score as the evaluation metric, defined as: where ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ = ^^ ^^/( ^^ ^^ + ^^ ^^) and ^^ ^^ ^^ ^^ ^^ ^^ = ^^ ^^/( ^^ ^^ + ^^ ^^). F1 score computes the harmonic mean between precision and recall. The former represents how likely an algorithm would give true alarms, and the latter indicates how sensitive an algorithm is towards positive samples. Additionally, mobile operators may be concerned about the False Positive Rate, ^^ ^^ ^^ = ^^ ^^/( ^^ ^^ + ^^ ^^). A high FPR could result in a waste of resources as every predicted positive instance needs to be checked by on-site engineers. Conversely, the False Negative Rate is defined as ^^ ^^ ^^ = 1 − ^^ ^^ ^^ ^^ ^^ ^^. 3.3 Efficacy of Data Augmentation We demonstrate that our data augmentation technique improves performance, with regard to all evaluation metrics considered, on both datasets at multiple clustering levels. Table 1 compares the metrics achieved with and without using data augmentation. On average, data augmentation results in an average F1 score increase of 10%.
Network without augmentation with augmentation Dataset Level Precision Recall F1 FPR(%) Precision Recall F1 FNR(%) DC 0.828 0.763 0.79 3.0 0.872 0.878 0.874 4.1 CRAN 0.808 0.487 0.601 3.1 0.793 0.823 0.807 6.1 City 1 MEC 0.701 0.473 0.547 4.7 0.794 0.779 0.786 5.4 antenna 0.414 0.177 0.235 5.0 0.481 0.423 0.448 6.4 DC 0.762 0.554 0.632 3.0 0.757 0.714 0.731 3.7 CRAN 0.718 0.477 0.562 4.6 0.682 0.643 0.658 6.5 City 2 MEC 0.646 0.479 0.541 4.5 0.679 0.633 0.652 4.9 antenna 0.379 0.260 0.302 7.6 0.392 0.434 0.405 12.0 Table 1: The average Precision, Recall, F1 scores and FPR of our model on two datasets with/without the anomaly augmentation technique applied. The anomalies are labeled with NN-generated baseline and the distance-based method.
4 Feature-wise Linear Modulation (FiLM) FiLM stands for Feature-wise Linear Modulation. The reference E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” 2017 is referred to. FiLM is a general-purpose conditioning method for neural networks. FiLM layers influence neural network (NN) computation via a simple, feature-wise affine transformation based on conditioning information. A FiLM layer carries out a simple, feature-wise affine transformation on a neural network’s intermediate features, conditioned on an arbitrary input.^ FiLM layers may enable a Recurrent Neural Network (RNN) over an input question to influence a Convolutional Neural Network (CNN) computation. This process may adaptively and radically alter the CNN’s behavior as a function of the input question, allowing the overall model to carry out a variety of reasoning tasks, ranging from counting to comparing, for example. FiLM operates in a coherent manner. It learns a complex, underlying structure and manipulates the conditioned network’s features in a selective manner. It also enables the CNN to properly localize question-referenced objects. FiLM models may learn from little data to generalize to more complex and/or substantially different data than seen during training. FiLM learns to adaptively influence the output of a neural network by applying an affine transformation, or FiLM, to the network’s intermediate features, based on some input. More formally, FiLM learns functions f and h which output^ γ and β as a function of input. f and h can be arbitrary functions such as neural networks. Modulation of a target neural network’s processing can be based on the same input to that neural network or some other input, as in the case of multi-modal or conditional tasks. As FiLM only requires two parameters per modulated feature map, it is a scalable and computationally efficient conditioning method. Figure 4 shows an example of a single FiLM layer for a CNN. Note It is to be understood that the above-referenced arrangements are only illustrative of the application for the principles of the present invention. Numerous modifications and alternative arrangements can be devised without departing from the spirit and scope of the present invention. While the present invention has been shown in the drawings and fully described above with particularity and detail in connection with what is presently deemed to be the most practical and preferred example(s) of the invention, it will be apparent to those of ordinary skill in the art that numerous modifications can be made without departing from the principles and concepts of the invention as set forth herein.

Claims

CLAIMS 1. A computer implemented method of generating training data for an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) receiving a communications network historical timeseries which includes one or more performance indicators (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) as a function of timestamp; (ii) for each timestamp, computing respective timestamp features denoting relative position in time, e.g. within a year, the month, week of the month, weekday, hour, minute and second; (iii) generating a function which receives timestamp features as inputs and outputs the closest performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) at the same time (e.g. the same weekday, the same hour, the same minute and the same second), for any predefined time period (e.g. week) except for the predefined time period (e.g. week) in the input timestamp features; (iv) using the function generated in (iii), labelling each chunk of the performance indicator (e.g. communications network historical traffic, accessibility, physical resource block (PRB) utilisation, call drop rate) timeseries based on deviation of the chunk from a baseline, e.g. using a Wasserstein distance to measure the deviation using the function generated in (iii); and (v) storing training data including the label of each chunk of the performance indicator (e.g. communications network historical traffic, accessibility, physical resource block (PRB) utilisation, call drop rate) timeseries.
2. The method of Claim 1, in which the one or more performance indicators is or includes communications network traffic volume, accessibility, physical resource block (PRB) utilisation, or call drop rate.
3. The method of Claims 1 or 2, in which the predefined time period is week, hour, or several hours, or a plurality of hours.
4. The method of any previous Claim, wherein a baseline is a common profile of the performance indicator of the historical timeseries, that describes how the performance indicator normally evolves over each day, month or year.
5. The method of any previous Claim, wherein the baseline is a daily historical average.
6. The method of any of Claims 1 to 4, wherein the baseline is generated by a neural network, e.g. which incorporates some hidden features of the performance indicator timeseries, such as seasonality.
7. The method of any previous Claim, wherein the function generated in step (iii) is used to produce a baseline.
8. The method of any previous Claim, wherein step (iii) includes: (a) Receiving the timestamp features as inputs and extracting features of the timestamp features, and forwarding the extracted timestamp features to a MLP layer to generate first MLP output, and forwarding the extracted timestamp features to a first FC layer to generate first FC layer output, and forwarding the extracted timestamp features to a second FC layer to generate second FC layer output; (b) Performing element-wise product transformation on the first MLP output and on the first FC layer output to generate element-wise product output; (c) Performing element-wise addition transformation on the element-wise product output and on the second FC layer output to generate element-wise addition output; (d) Inputting the element-wise addition output to the MLP layer to generate second MLP output, and evaluating a (e.g. MAE) loss of the second MLP output compared to performance indicator value (e.g. communications network traffic volume) for the timestamp which corresponds to the timestamp features; (e) Repeating steps (b) to (d), while modulating the first FC layer and the second FC layer, until a (e.g. MAE) loss minimization criterion is satisfied.
9. The method of Claim 8, including storing weights of the first FC layer and of the second FC layer.
10. The method of Claims 8 or 9, wherein the MLP layer, the first FC layer and the second FC layer are independent.
11. The method of any of Claims 8 to 10, wherein the modulation is an element- wise affine transformation on the first MLP output.
12. The method of any of Claims 8 to 11, wherein the first FC layer output and the second FC layer output^ influence the output of the MLP by scaling and shifting each intermediate feature based on the first FC layer input and on the second FC layer input.
13. The method of any previous Claim, including constructing a first sample of the training set, the first sample including P steps up to the present of the training set, and repeating for each P; constructing a second sample of the training set, the second sample including a future F steps of the training set, and repeating for each F; identifying a random variable sampled from a uniform distribution, and scaling each pair of the first sample and the second sample using the random variable, to create a scaled set of pairs of the first sample and of the second sample, and performing steps (iv) and (v) on the scaled set of pairs of the first sample and of the second sample, to label each pair of the scaled set of pairs, and to store augmented training data including the label of each pair of the scaled set of pairs.
14. The method of Claim 13, wherein the variability of anomalous performance indicator values (e.g. network traffic) from which the neural networks learn is enhanced.
15. The method of Claims 13 or 14, wherein the volume and the variability of anomalous performance indicator values (e.g. network traffic) is increased.
16. The method of any of Claims 13 to 15, wherein normal/benign time series are input, which are randomly scaled, using a scaling factor sampled from a pre-set uniform distribution.
17. The method of any of Claims 13 to 16, wherein the scaling results in larger deviations from the baseline, and thereby more synthetic anomalies are generated.
18. A storage medium storing the training data or the augmented training data of any previous Claim.
19. A computer implemented method of training an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) receiving training data including a communications network historical performance indicator timeseries which includes performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) as a function of timestamp, and a label of each chunk of the communications network historical performance indicator timeseries; (ii) processing the training data using a positional encoding layer which encodes the timestamp of each datum and fuses that with a corresponding performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) to produce processed training data; (iii) encoding the processed training data in a first set of LSTMs as an encoded representation; (iv) passing the encoded representation to a decoder including a second set of LSTMs which forecasts performance indicator value (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) of next steps to produce a future performance indicator value forecast; (v) outputting the encoded representation from the first set of LSTMs via a first FC layer to a first FiLM layer; (vi) outputting the future performance indicator value forecast from the second set of LSTMs via a second FC layer to a second FiLM layer; (vii) the first FiLM layer fusing the received encoded representation with received input conditional features (e.g. the mean, standard deviation, maximum, minimum, and skewness of the communications network historical performance indicator timeseries) to generate output of the first FiLM layer; (viii) the second FiLM layer fusing the received future performance indicator value forecast with received target conditional features to generate output of the second FiLM layer; (ix) concatenating the output of the first FiLM layer and the output of the second FiLM layer to produce a concatenated vector; (x) evaluating a first (e.g. MSE) loss function of the future performance indicator value forecast compared to actual future performance indicator value, and passing the concatenated vector through a MLP and evaluating the MLP output using a second (e.g. BCE, e.g. including squared error (SE)) loss function which compares predicted labels with ground truth labels, and (xi) adjusting the first FC layer and the second FC layer while repeating steps (v) to (x) until a convergence criterion with respect to the first loss function and the second loss function is satisfied, to produce a trained model.
20. The method of Claim 19, in which the performance indicator is communications network traffic volume, accessibility, physical resource block (PRB) utilisation, or call drop rate.
21. The method of Claims 19 or 20, wherein the training data includes training data of any of Claims 1 to 11, or augmented training data of any of Claims 12 to 16.
22. The method of any of Claims 19 to 21, including storing weights of the trained model.
23. The method of any of Claims 19 to 22, wherein anomaly prediction is treated as a combined task of both classification and forecasting.
24. The method of any of Claims 19 to 23, wherein a prediction is made whether an anomaly will occur (e.g. soon) in the network, while forecasting is also needed since an ‘anomaly’ in our task is defined as a period of performance indicator values (e.g. traffic) that deviates from the baseline, i.e., forecasting the trend of upcoming performance indicator values (e.g. traffic) plays an important role in the decision- making process.
25. The method of any of Claims 19 to 24, wherein a Seq2Seq network is used to encode the historical data and forecast the future performance indicator values (e.g. traffic), further fusing the historical information with predicted upcoming performance indicator values (e.g. traffic), which is supplied to the final classifier.
26. The method of any of Claims 19 to 25, wherein a Seq2Seq architecture is used, with a stacked LSTM block used as the encoder.
27. The method of any of Claims 19 to 26, wherein to achieve both classification and forecasting simultaneously, the neural models involved are trained with Cross Entropy (CE) and Mean Squared Error (MSE) loss functions, employing a balancing factor.
28. The method of any of Claims 19 to 27, wherein predictions are based on the provided input and the predicted future performance indicator values (e.g. traffic) together.
29. The method of any of Claims 19 to 28, wherein a positional encoding layer is added before the encoder, which encodes the timestamp of each datum and fuses that with the original input.
30. The method of any of Claims 19 to 29, wherein a timestamp is encoded as a d- dimensional vector in which d elements have varying frequencies depending on the index of the vector.
31. The method of any of Claims 19 to 30, wherein conditional features are features obtained solely from the associated timeseries data.
32. The method of any of Claims 19 to 31, wherein FiLM acts as a feature-wise affine transformation on its inputs.
33. A storage medium storing the weights of the trained model of any of Claims 19 to 32.
34. A computer system, including a trained artificial intelligence model for operational anomaly prediction in a communications network produced by a method of any of Claims 19 to 32.
35. Computer implemented method of operational anomaly prediction in a communications network, including the step of receiving input of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) time series data into the computer system including the trained artificial intelligence model for operational anomaly prediction in a communications network of Claim 34, and outputting from the computer system including the trained artificial intelligence model an operational anomaly prediction for the communications network.
36. The method of Claim 35, in which associations between antennas and upper- level network nodes of the communications network are determined by using a k- partitioning algorithm over a Delaunay graph of antennas.
37. The method of Claims 35 or 36, including computing rates for true positive, true negative, false positive and false negative.
38. A computer implemented method of training an artificial intelligence model for operational anomaly prediction in a communications network, the method including the steps of: (i) leveraging a data-driven approach to generate in an unsupervised way baselines describing normal daily patterns of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate), followed by a measure of deviation from the baseline to label each segment of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate); (ii) using a data augmentation technique that increases both the volume and the variability of anomalous performance indicator values, including taking as input normal/benign time series and randomly scaling them, with a scaling factor sampled from a pre-set uniform distribution; (iii) utilizing a Seq2Seq network to encode the historical data and forecast the future, further fusing the historical information with predicted upcoming performance indicator values, which is supplied to the final classifier; (iv) to achieve both classification and forecasting simultaneously, training the neural models involved with respective first (e.g. Cross Entropy (CE)) and second (e.g. Mean Squared Error (MSE)) loss functions, employing a balancing factor.
39. The method of Claim 38, including a method of any of Claims 1 to 17, or 19 to 32.
40. A computer system including a trained artificial intelligence model for operational anomaly prediction in a communications network, trained using a method of Claims 38 or 39.
41. Computer implemented method of operational anomaly prediction in a communications network, including the step of receiving input of performance indicator values (e.g. communications network traffic volume, accessibility, physical resource block (PRB) utilisation, call drop rate) time series data into the computer system including the trained artificial intelligence model for operational anomaly prediction in a communications network of Claim 40, and outputting from the computer system including the trained artificial intelligence model an operational anomaly prediction for the communications network.
EP24726687.7A 2023-05-02 2024-05-02 Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems Pending EP4706227A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GBGB2306469.4A GB202306469D0 (en) 2023-05-02 2023-05-02 Methods
PCT/GB2024/051153 WO2024228021A1 (en) 2023-05-02 2024-05-02 Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems

Publications (1)

Publication Number Publication Date
EP4706227A1 true EP4706227A1 (en) 2026-03-11

Family

ID=86691985

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24726687.7A Pending EP4706227A1 (en) 2023-05-02 2024-05-02 Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems

Country Status (3)

Country Link
EP (1) EP4706227A1 (en)
GB (1) GB202306469D0 (en)
WO (1) WO2024228021A1 (en)

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119675990B (en) * 2025-02-19 2025-10-28 国家计算机网络与信息安全管理中心江西分中心 A multi-functional monitoring platform for IPv6 traffic with multi-level management
CN120386981B (en) * 2025-04-10 2025-10-03 中国联合重型燃气轮机技术有限公司 Performance prediction method of gas turbine secondary air system based on transformer
CN120654233B (en) * 2025-05-30 2025-12-09 国能(肇庆)热电有限公司 Abnormal operation behavior analysis method and system based on business behavior distribution
CN120278527B (en) * 2025-06-03 2025-09-19 山东师范大学 A dynamic prediction method for financial risks based on deep learning
CN120750417B (en) * 2025-08-06 2026-04-14 深圳市光网世纪科技有限公司 Method, device, equipment and medium for detecting abnormal flow of optical fiber network
CN121056336B (en) * 2025-10-31 2026-03-20 国网上海市电力公司 Training method for network monitoring model of low-voltage distributed power monitoring system

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10506457B2 (en) 2015-09-30 2019-12-10 Telecom Italia S.P.A. Method for managing wireless communication networks by prediction of traffic parameters
US20170339022A1 (en) * 2016-05-17 2017-11-23 Brocade Communications Systems, Inc. Anomaly detection and prediction in a packet broker
IL249950A0 (en) 2017-01-05 2017-06-29 Shapira Bracha A prediction system configured for modeling the expected number of attacks on a computer or communication network
US11496353B2 (en) * 2019-05-30 2022-11-08 Samsung Electronics Co., Ltd. Root cause analysis and automation using machine learning
EP4073653A4 (en) * 2019-12-09 2022-12-14 Visa International Service Association Failure prediction in distributed systems
CN112532643B (en) * 2020-12-07 2024-02-20 长春工程学院 Traffic anomaly detection methods, systems, terminals and media based on deep learning

Also Published As

Publication number Publication date
GB202306469D0 (en) 2023-06-14
WO2024228021A1 (en) 2024-11-07

Similar Documents

Publication Publication Date Title
EP4706227A1 (en) Methods of training an artificial intelligence model for operational anomaly prediction in a communications network, and systems
Jagait et al. Load forecasting under concept drift: Online ensemble learning with recurrent neural network and ARIMA
Manias et al. Concept drift detection in federated networked systems
Su et al. Uncertainty quantification of collaborative detection for self-driving
Kondratenko et al. Multi-criteria decision making for selecting a rational IoT platform
Cazzanti et al. Mining maritime vessel traffic: Promises, challenges, techniques
WO2020164740A1 (en) Methods and systems for automatically selecting a model for time series prediction of a data stream
Eljabu et al. Anomaly detection in maritime domain based on spatio-temporal analysis of ais data using graph neural networks
Srivastava et al. Framework for ship trajectory forecasting based on linear stationary models using automatic identification system
WO2025034245A1 (en) Methods and processes to enable rnn-gnn-based network digital twin for o-ran
Kirmaz et al. Mobile network traffic forecasting using artificial neural networks
Moysen et al. Big data-driven automated anomaly detection and performance forecasting in mobile networks
Wang et al. Deep Bi-Directional Adaptive Gating Graph Convolutional Networks for Spatio-Temporal Traffic Forecasting
US11962475B2 (en) Estimating properties of units using system state graph models
Sharma et al. An Experimental Study for Comparing Different Method for Time Series Forecasting Prediction & Anomaly Detection
CN116522213A (en) Service state level classification and classification model training method and electronic equipment
Shoman et al. Graph Convolutional Gated Recurrent Unit Network for Traffic Prediction Using Loop Detector Data
Choi et al. The empirical evaluation of machine learning models predicting round-trip time in cellular network
Malhotra A new traffic prediction algorithm with machine learning for optical fiber communication system
Singh et al. TL-ConvLSTM: A Transfer-Learning-Based Convolutional LSTM to Identify and Forecast Traffic in the NextG Environments
Nguyen Anomaly detection in self-organizing network
Zhao et al. Long-term traffic flow prediction model based on SUTDGCN
Hofmann Developing a streaming-based architecture for demand prediction of taxi trips in the presence of concept drift
Bhimu Scalable Smart City Analytics Through SmartFusion-Grid: A Unified Data Engineering Approach
Tian A Software Defined Network Traffic Prediction Model Based on LSTM and Feature Compression

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251202

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR