WO2020005599A1 - Trend prediction based on neural network - Google Patents
Trend prediction based on neural network Download PDFInfo
- Publication number
- WO2020005599A1 WO2020005599A1 PCT/US2019/037408 US2019037408W WO2020005599A1 WO 2020005599 A1 WO2020005599 A1 WO 2020005599A1 US 2019037408 W US2019037408 W US 2019037408W WO 2020005599 A1 WO2020005599 A1 WO 2020005599A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- time series
- series data
- neural network
- vector representation
- recurrent neural
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/04—Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
- G06N3/0442—Recurrent networks, e.g. Hopfield networks characterised by memory or gating, e.g. long short-term memory [LSTM] or gated recurrent units [GRU]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/01—Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
Definitions
- data associated with one or more objects may be abrupt and time-relevant.
- news data associated with epidemiology and monitoring data within a large-scale data center may have the above characteristics. Therefore, it becomes quite difficult to predict the object trend based on such data.
- information sources are becoming more and more abundant and data are becoming more and more diverse, which makes it difficult to determine information quality and reliability. Hence, it is desirable to provide an improved trend prediction solution.
- a method comprises: obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
- FIG. 1 is a block diagram illustrating a computing device for implementing various implementations of the present disclosure
- FIG. 2 illustrates an architecture of a neural network in accordance with some implementations of the present disclosure
- FIG. 3 is a schematic diagram illustrating a recurrent neural network in accordance with some implementations of the present disclosure
- FIG. 4 is a flowchart illustrating a method for predicting a trend of an object in accordance with some implementations of the present disclosure.
- FIG. 5 is a flowchart illustrating a method for training a neural network in accordance with an implementation of the present disclosure.
- the term“includes” and its variants are to be read as open terms that mean“includes, but is not limited to.”
- the term“based on” is to be read as“based at least in part on.”
- the term“one implementation” and“an implementation” are to be read as “at least one implementation.”
- the term“another implementation” is to be read as“at least one other implementation.”
- the terms“first,”“second,” and the like may refer to different or same objects. Other definitions, explicit and implicit, may be included below.
- Fig. 1 illustrates a block diagram of a computing device 100 that can carry out a plurality of implementations of the present disclosure. It should be understood that the computing device 100 shown in Fig. 1 is only exemplary and shall not constitute any restrictions over functions and scopes of the implementations described by the present disclosure. According to Fig. 1, the computing device 100 includes a computing device 100 in the form of a general purpose computing device. Components of the computing device 100 can include, but not limited to, one or more processors or processing units 110, memory 120, storage device 130, one or more communication units 140, one or more input devices 150 and one or more output devices 160.
- the computing device 100 can be implemented as various user terminals or service terminals with computing power.
- the service terminals can be servers, large-scale computing devices and the like provided by a variety of service providers.
- the user terminal for example, is mobile terminal, fixed terminal or portable terminal of any types, including mobile phone, site, unit, device, multimedia computer, multimedia tablet, Internet nodes, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, Personal Communication System (PCS) device, personal navigation device, Personal Digital Assistant (PDA), audio/video player, digital camera/video, positioning device, television receiver, radio broadcast receiver, electronic book device, gaming device or any other combinations thereof consisting of accessories and peripherals of these devices or any other combinations thereof. It can also be predicted that the computing device 100 can support any types of user-specific interfaces (such as“wearable” circuit and the like).
- the processing unit 110 can be a physical or virtual processor and can execute various processing based on the programs stored in the memory 120. In a multi-processor system, a plurality of processing units executes computer-executable instructions in parallel to enhance parallel processing capability of the computing device 100.
- the processing unit 110 also can be known as central processing unit (CPU), microprocessor, controller and microcontroller.
- the computing device 100 usually includes a plurality of computer storage media. Such media can be any attainable media accessible by the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media.
- the memory 120 can be a volatile memory (e.g., register, cache, Random Access Memory (RAM)), a non-volatile memory (such as, Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash), or any combinations thereof.
- the memory 120 can include a predicting module 122 configured to execute functions of various implementations described herein. The predicting module 122 can be accessed and operated by the processing unit 110 to perform corresponding functions.
- the storage device 130 can be removable or non-removable medium, and can include machine readable medium, which can be used for storing information and/or data and can be accessed within the computing device 100.
- the computing device 100 can further include a further removable/non-removable, volatile/non-volatile storage medium.
- a disk drive for reading from or writing into a removable and non-volatile disk
- an optical disk drive for reading from or writing into a removable and non-volatile optical disk.
- each drive can be connected via one or more data medium interfaces to the bus (not shown).
- the communication unit 140 implements communication with another computing device through communication media. Additionally, functions of components of the computing device 100 can be realized by a single computing cluster or a plurality of computing machines, and these computing machines can communicate through communication connections. Therefore, the computing device 100 can be operated in a networked environment using a logic connection to one or more other servers, a Personal Computer (PC) or a further general network node.
- PC Personal Computer
- the input device 150 can be one or more various input devices, such as mouse, keyboard, trackball, voice-input device and the like.
- the output device 160 can be one or more output devices, e.g., display, loudspeaker and printer etc.
- the computing device 100 also can communicate through the communication unit 140 with one or more external devices (not shown) as required, wherein the external device, e.g., storage device, display device etc., communicates with one or more devices that enable the users to interact with the computing device 100, or with any devices (such as network card, modem and the like) that enable the computing device 100 to communicate with one or more other computing devices. Such communication can be executed via Input/Output (I/O) interface (not shown).
- I/O Input/Output
- the computing device 100 can be provided for implementing trend prediction for one or more objects in accordance with implementations of the present disclosure, for example, the object increases, reduces, or remains stable.
- the computing device 100 when performing the prediction, can receive time series data 170 associated with the one or more objects through the input device 150.
- the computing device 100 can process the time series data 170 and determine the trend of the object based on the time series data 170.
- the trend of the object can be provided to the output device 160 as an output 180 for the user and the like.
- Fig. 2 is a schematic diagram illustrating a neural network 200 for predicting a trend of an object in accordance with some implementations of the present disclosure.
- the neural network 200 includes a Recurrent Neural Network (RNN) 230, which can comprise one or more recurrent neural network units, such as Gated Recurrent Unit (GRU), Long Short Term Memory (LSTM) unit and/or the like.
- RNN Recurrent Neural Network
- GRU Gated Recurrent Unit
- LSTM Long Short Term Memory
- the recurrent neural network unit can receive the input data and latent state of the previous time step and determine the latent state of this time step based on the input data of this time step and the latent state of the previous time step.
- the latent state of the last time step usually serves as an output provided to the subsequent network.
- the neural network 200 can obtain the time series data associated with one or more objects.
- the time series data can be data in the text form, such as news etc., and also can be data in other forms, e.g., numbers.
- the time series data can include news at different time and from different sources.
- data items such as news can be converted through an embedding layer into corresponding vector representations, which are provided to the neural network 200 to determine a vector representation (e.g., vector V 250 shown in Fig. 2) of the time series data (also known as a sample for training).
- the vector representation can have the same dimension as the latent state of the recurrent neural network 230.
- a number of words can be selected from one data item such as news to determine a vector representation of each word through the embedding layer, and these vector representations are averaged to determine the vector representation of the data item.
- vector representations n 2,i 202, n 2 2 204 and n 2 L 206 of several data items are illustrated as examples in Fig. 2.
- the recurrent neural network 230 can be used for determining the vector representation of the obtained time series data.
- the vector representation can be the vector V 250 as shown in Fig. 2.
- a discriminative network 260 can determine, based on the vector V 250, a probability that an object is associated with a predefined label as an output 270, which is provided to facilitate prediction of the in which the object varies over time.
- the discriminative network 260 can output a label with the maximum probability as the prediction for the object trend.
- the label can indicate a class of trend in which the object varies over time, for example, up, down, stable, and/or the like.
- the time series data can be segmented or divided into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network 230 as shown in Fig. 2.
- the time series data is segmented into these portions based on a predefined time period.
- Each input of the recurrent neural network 230 can correspond to one time step, such as one day or one week. Therefore, each portion of the time series data can include a plurality of different data items within the respective time step.
- an attention layer 210 can determine weights of different data items for each time step and weight the respective data items with the weights to determine the weighted data item associated with the time step.
- the weighted data item can be provided to a respective input of the recurrent neural network 230.
- Equation (1) illustrates a mathematical expression for determining the weighted data item in accordance with some implementations of the present disclosure.
- the vector representation n ti of the data item i at the time step t is provided to the attention layer 210 to obtain a respective attention value ua, where sigmoid(-) represents a sigmoid function and W n and b n denote a parameter matrix and a parameter vector, respectively.
- the attention value uu is normalized through the soflmax function to obtain a weight a ti .
- the weight a ti can weight the vector representation n ti of the respective data item to determine the weighted data item d t at the time step t.
- the attention layer 210 can determine weights a 2 1 , a 2 2 , and a 2 L (also known as importance or attention) for the data items n 2 1 202, n 2 2 204, and n 2 L 206.
- the weighted data item d 2 224 can be provided to the respective input of the recurrent neural network 230.
- the recurrent neural network 230 can be a bidirectional recurrent neural network.
- Fig. 3 illustrates a recurrent neural network 230 in accordance with some implementations of the present disclosure.
- the recurrent neural network units 314, 324, and 334 in the recurrent neural network 230 can receive, at each time step, respective input data and the latent state at the previous time step.
- the recurrent neural network unit 324 receives, from the recurrent neural network unit 314, the latent state at the previous time step and respective input data (e.g., weighted data item d 2 ) and determines the latent state at the time step based on the latent state and the input data (e.g., weighted data item d 2 ).
- the recurrent neural network units 312, 322 and 332 in the recurrent neural network 230 can receive, at each time step, the respective input data and the latent state at the next time step.
- the recurrent neural network unit 312 receives, from the recurrent neural network unit 322, the latent state at the next time step and respective input data (e.g., weighted data item ⁇ c ) and determines the latent state at the time step based on the latent state and the input data (e.g., weighted data item d ⁇ ).
- respective input data e.g., weighted data item ⁇ c
- the neural network unit at each time step can be the same neural network unit and the neural network unit is shown at each time step for the purpose of illustration only.
- the recurrent neural network unit can be Gated Recurrent Unit (GRU).
- GRU Gated Recurrent Unit
- the GRU computes the state h t at the time step by linearly interpolating the state h t-1 of the previous time step and the current updated state h t as shown in equation (2):
- the updated current state h t can be computed by combining the input vector at the time step and the state at the previous time step as shown in equation (3):
- W r , U r , b r , W z , U z , and b z denote parameter matrices and vectors.
- the latent vector for the time step can be obtained via the GRU.
- the latent vectors in both directions can be concatenated to build a bidirectional encoded vector h t as an output of the convolutional neural network 230 according to equation (5):
- the bidirectional recurrent neural network is introduced above with reference to Figs. 2 and 3, it should be appreciated that implementations of the present disclosure also can be applied in a uni-directional recurrent neural network.
- the reverse recurrent neural network unit can be omitted and the latent state of each time step is provided for the subsequent neural network layers.
- the latent state at the last time step is provided for the subsequent neural network.
- the attention layer 240 which will be subsequently described in details, can be omitted in this implementation.
- the attention layer 240 can determine the weights of data items provided by a plurality of outputs of the recurrent neural network 230 and weight the corresponding data items using these weights. For example, more recent data may have a big influence over the prediction result while the less recent data may have a small influence over the prediction result.
- the data item provided by the output of the recurrent neural network 230 can be the bidirectional encoded vector as mentioned above.
- the attention layer 240 considers the factor that the importance of the information varies for the trend determination at the respective time step. By means of the attention layer 240, the vector representation (vector V) of the time series data can be calculated via the equation (6):
- W h and b h represent parameter matrix and vector
- 0 L denotes parameters at each time step in the soflmax layer indicating which time step is more important
- o L indicates the latent representation of the bidirectional encoded vector h ⁇ .
- the attention vector b is acquired to weight the outputs of different time. Accordingly, the attention vector b is employed to weight the bidirectional encoded vector hi to compute the vector representation V of the time series data for subsequent classification.
- the vector representation V is a weighted value of the bidirectional encoded vector the vector representation V can have the same dimension as the latent state of the recurrent neural network. The dimension can be pre-specified.
- the discriminative network 260 can be a classifier of Multi-Layer Perceptron (MLP), which receives the vector V as the input and outputs the classification result.
- the classification result can include a probability that the obj ect is associated with the predefined label. For example, the label with the maximum probability can be output as the classification result.
- MLP Multi-Layer Perceptron
- the training difficulty varies for different training data.
- some training data lacks sufficient information and the training of such training data can be quite challenging.
- some challenging training data can be skipped and the training data which can be easily trained is selected and the challenging training data can be progressively added to the training process.
- the neural network 200 can be trained using the self-paced learning, which can automatically choose the suitable training samples for different training phases, so as to enhance the model performance.
- E(w, v, ) represents an objective function
- a regularization term for self-paced learning or known as penalty
- l is a hyper-parameter that controls the pace at which the model learns new samples.
- the regularization term f(v: A) can be computed using the equation (8):
- ACS Alternative Convex Search
- equation (7) ACS divides the variables into two disjoint blocks. In each iteration, a block of variables is optimized while the other block is fixed. With the fixed w, the unconstrained solution for the linear regularization term (8) can be computed by the equation (9):
- the self-paced learning inputs the data set D, the linear regularizer f and the self-paced step m with the expectation of obtaining the model parameter w.
- l if being too small, can be increased based on the step m.
- w * can be output as the trained model parameter. As the training progresses and A increases, samples with larger loss will be gradually incorporated, thereby realizing the effective and efficient learning.
- Fig. 4 is a flowchart illustrating a method 400 for predicting the trend in accordance with some implementations of the present disclosure.
- the method 400 can be implemented by the computing device 100 or the predicting module 122 shown by Fig. 1.
- the concepts introduced above with reference to Figs. 2 and 3, alone or in combination, can be applied into the method 400.
- time series data associated with an object is obtained.
- a vector representation of the time series data is computed with a recurrent neural network.
- computing the vector representation of the time series data includes: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- the vector representation of the time series data can be determined by for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data includes: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network can be a bidirectional recurrent neural network, which also includes a plurality of outputs corresponding to a plurality of inputs.
- Determining the vector representation of the time series data includes determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- a probability that the object is associated with a predefined label is determined based on the vector representation of the time series data.
- the label indicates a class of a trend in which the object varies over time. In this way, the prediction problem can be transformed into a classification problem.
- Fig. 5 is a flowchart illustrating a method 500 for training a neural network in accordance with some implementations of the present disclosure.
- the method 500 can be implemented by the computing device 100 or the predicting module 122 shown by Fig. 1.
- the concepts introduced above with reference to Figs. 2 and 3, alone or in combination, can be applied into the method 500.
- the time series data associated with an object and a predefined label for the object are obtained.
- the label can indicate a class of a trend in which the object varies over time.
- the recurrent neural network for predicting the trend in which the object varies over time is updated based on the time series data and the label.
- the updating comprises computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
- a computer- implemented method comprises: The method comprises obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- determining the vector representation of the time series data based on the relationship of importance comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- a computer- implemented method comprises obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend that the object varies over time; and updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising: computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
- a device comprising: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- determining the vector representation of the time series data based on the relationship of importance comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- a device comprising: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend that the object varies over time; and updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising: computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
- determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
- computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
- the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
- updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
- a computer program product comprising instructions stored tangibly on a computer-readable medium, the instructions, when executed by a machine, causing the machine to execute the method described above.
- the functionality described herein can be performed, at least in part, by one or more hardware logic components.
- illustrative types of hardware logic components include Field-Programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
- FPGAs Field-Programmable Gate Arrays
- ASICs Application-specific Integrated Circuits
- ASSPs Application-specific Standard Products
- SOCs System-on-a-chip systems
- CPLDs Complex Programmable Logic Devices
- the functions described above herein can be at least partially executed by a Graphical Processing Unit (GPU).
- GPU Graphical Processing Unit
- Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions/operations specified in the flowcharts and/or block diagrams to be implemented.
- the program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
- a machine readable medium may be any tangible medium that may contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
- the machine readable medium may be a machine readable signal medium or a machine readable storage medium.
- a machine readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- machine readable storage medium More specific examples of the machine readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or Flash memory erasable programmable read-only memory
- CD-ROM portable compact disc read-only memory
- magnetic storage device or any suitable combination of the foregoing.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Artificial Intelligence (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Biomedical Technology (AREA)
- Business, Economics & Management (AREA)
- Human Resources & Organizations (AREA)
- Economics (AREA)
- Strategic Management (AREA)
- Entrepreneurship & Innovation (AREA)
- Game Theory and Decision Science (AREA)
- Development Economics (AREA)
- Marketing (AREA)
- Operations Research (AREA)
- Quality & Reliability (AREA)
- Tourism & Hospitality (AREA)
- General Business, Economics & Management (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
Implementations of the present disclosure relate to trend prediction based on neural network. In some implementations, a method comprises: obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
Description
TREND PREDICTION BASED ON NEURAL NETWORK
BACKGROUND
[0001] In some trend predicting applications, data associated with one or more objects may be abrupt and time-relevant. For example, news data associated with epidemiology and monitoring data within a large-scale data center may have the above characteristics. Therefore, it becomes quite difficult to predict the object trend based on such data. Furthermore, with the development of the network, information sources are becoming more and more abundant and data are becoming more and more diverse, which makes it difficult to determine information quality and reliability. Hence, it is desirable to provide an improved trend prediction solution.
SUMMARY
[0002] Various implementations of the present disclosure provide a solution for trend prediction based on a neural network. In some implementations, a method comprises: obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Fig. 1 is a block diagram illustrating a computing device for implementing various implementations of the present disclosure;
[0005] Figs. 2 illustrates an architecture of a neural network in accordance with some implementations of the present disclosure;
[0006] Fig. 3 is a schematic diagram illustrating a recurrent neural network in accordance with some implementations of the present disclosure;
[0007] Fig. 4 is a flowchart illustrating a method for predicting a trend of an object in accordance with some implementations of the present disclosure; and
[0008] Fig. 5 is a flowchart illustrating a method for training a neural network in
accordance with an implementation of the present disclosure.
[0009] Throughout the drawings, the same or similar reference symbols refer to the same or similar elements.
DETAILED DESCRIPTION
[0010] The subject matter described herein will now be discussed with reference to several example implementations. It is to be understood these implementations are discussed only for the purpose of enabling those skilled persons in the art to better understand and thus implement the subject matter described herein, rather than suggesting any limitations on the scope of the subject matter.
[0011] As used herein, the term“includes” and its variants are to be read as open terms that mean“includes, but is not limited to.” The term“based on” is to be read as“based at least in part on.” The term“one implementation” and“an implementation” are to be read as “at least one implementation.” The term“another implementation” is to be read as“at least one other implementation.” The terms“first,”“second,” and the like may refer to different or same objects. Other definitions, explicit and implicit, may be included below.
Example Environment
[0012] Basic principles and several example implementations of the present disclosure will be explained below with reference to the drawings. Fig. 1 illustrates a block diagram of a computing device 100 that can carry out a plurality of implementations of the present disclosure. It should be understood that the computing device 100 shown in Fig. 1 is only exemplary and shall not constitute any restrictions over functions and scopes of the implementations described by the present disclosure. According to Fig. 1, the computing device 100 includes a computing device 100 in the form of a general purpose computing device. Components of the computing device 100 can include, but not limited to, one or more processors or processing units 110, memory 120, storage device 130, one or more communication units 140, one or more input devices 150 and one or more output devices 160.
[0013] In some implementations, the computing device 100 can be implemented as various user terminals or service terminals with computing power. The service terminals can be servers, large-scale computing devices and the like provided by a variety of service providers. The user terminal, for example, is mobile terminal, fixed terminal or portable terminal of any types, including mobile phone, site, unit, device, multimedia computer, multimedia tablet, Internet nodes, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, Personal Communication System
(PCS) device, personal navigation device, Personal Digital Assistant (PDA), audio/video player, digital camera/video, positioning device, television receiver, radio broadcast receiver, electronic book device, gaming device or any other combinations thereof consisting of accessories and peripherals of these devices or any other combinations thereof. It can also be predicted that the computing device 100 can support any types of user-specific interfaces (such as“wearable” circuit and the like).
[0014] The processing unit 110 can be a physical or virtual processor and can execute various processing based on the programs stored in the memory 120. In a multi-processor system, a plurality of processing units executes computer-executable instructions in parallel to enhance parallel processing capability of the computing device 100. The processing unit 110 also can be known as central processing unit (CPU), microprocessor, controller and microcontroller.
[0015] The computing device 100 usually includes a plurality of computer storage media. Such media can be any attainable media accessible by the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 can be a volatile memory (e.g., register, cache, Random Access Memory (RAM)), a non-volatile memory (such as, Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash), or any combinations thereof. The memory 120 can include a predicting module 122 configured to execute functions of various implementations described herein. The predicting module 122 can be accessed and operated by the processing unit 110 to perform corresponding functions.
[0016] The storage device 130 can be removable or non-removable medium, and can include machine readable medium, which can be used for storing information and/or data and can be accessed within the computing device 100. The computing device 100 can further include a further removable/non-removable, volatile/non-volatile storage medium. Although not shown in Fig. 1, there can be provided a disk drive for reading from or writing into a removable and non-volatile disk and an optical disk drive for reading from or writing into a removable and non-volatile optical disk. In such cases, each drive can be connected via one or more data medium interfaces to the bus (not shown).
[0017] The communication unit 140 implements communication with another computing device through communication media. Additionally, functions of components of the computing device 100 can be realized by a single computing cluster or a plurality of computing machines, and these computing machines can communicate through communication connections. Therefore, the computing device 100 can be operated in a
networked environment using a logic connection to one or more other servers, a Personal Computer (PC) or a further general network node.
[0018] The input device 150 can be one or more various input devices, such as mouse, keyboard, trackball, voice-input device and the like. The output device 160 can be one or more output devices, e.g., display, loudspeaker and printer etc. The computing device 100 also can communicate through the communication unit 140 with one or more external devices (not shown) as required, wherein the external device, e.g., storage device, display device etc., communicates with one or more devices that enable the users to interact with the computing device 100, or with any devices (such as network card, modem and the like) that enable the computing device 100 to communicate with one or more other computing devices. Such communication can be executed via Input/Output (I/O) interface (not shown).
[0019] The computing device 100 can be provided for implementing trend prediction for one or more objects in accordance with implementations of the present disclosure, for example, the object increases, reduces, or remains stable. The computing device 100, when performing the prediction, can receive time series data 170 associated with the one or more objects through the input device 150. The computing device 100 can process the time series data 170 and determine the trend of the object based on the time series data 170. The trend of the object can be provided to the output device 160 as an output 180 for the user and the like.
[0020] Fig. 2 is a schematic diagram illustrating a neural network 200 for predicting a trend of an object in accordance with some implementations of the present disclosure. As shown in Fig. 2, the neural network 200 includes a Recurrent Neural Network (RNN) 230, which can comprise one or more recurrent neural network units, such as Gated Recurrent Unit (GRU), Long Short Term Memory (LSTM) unit and/or the like. At each time step, the recurrent neural network unit can receive the input data and latent state of the previous time step and determine the latent state of this time step based on the input data of this time step and the latent state of the previous time step. In a unidirectional recurrent neural network, the latent state of the last time step usually serves as an output provided to the subsequent network.
[0021] For example, the neural network 200 can obtain the time series data associated with one or more objects. The time series data can be data in the text form, such as news etc., and also can be data in other forms, e.g., numbers. For example, the time series data can include news at different time and from different sources.
[0022] In some implementations, data items such as news can be converted through an
embedding layer into corresponding vector representations, which are provided to the neural network 200 to determine a vector representation (e.g., vector V 250 shown in Fig. 2) of the time series data (also known as a sample for training). The vector representation can have the same dimension as the latent state of the recurrent neural network 230. A number of words can be selected from one data item such as news to determine a vector representation of each word through the embedding layer, and these vector representations are averaged to determine the vector representation of the data item. For example, vector representations n2,i 202, n2 2 204 and n2 L 206 of several data items are illustrated as examples in Fig. 2.
[0023] As described above, the recurrent neural network 230 can be used for determining the vector representation of the obtained time series data. For example, the vector representation can be the vector V 250 as shown in Fig. 2. A discriminative network 260 can determine, based on the vector V 250, a probability that an object is associated with a predefined label as an output 270, which is provided to facilitate prediction of the in which the object varies over time. In addition or alternatively, the discriminative network 260 can output a label with the maximum probability as the prediction for the object trend. The label can indicate a class of trend in which the object varies over time, for example, up, down, stable, and/or the like.
[0024] In some implementations, the time series data can be segmented or divided into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network 230 as shown in Fig. 2. For example, the time series data is segmented into these portions based on a predefined time period. Each input of the recurrent neural network 230 can correspond to one time step, such as one day or one week. Therefore, each portion of the time series data can include a plurality of different data items within the respective time step. As shown in Fig. 2, a plurality of different data items n2 l 202, n2 2 204 and h2c 206 are illustrated as examples when the time step t=2.
[0025] In some implementations, different data items such as different news within each time step or time period show different importance for prediction. For example, some data items can greatly affect the prediction result while some data items have less influence on the prediction result. In order to consider this factor, an attention layer 210 can determine weights of different data items for each time step and weight the respective data items with the weights to determine the weighted data item associated with the time step. The weighted data item can be provided to a respective input of the recurrent neural network 230.
[0026] Equation (1) illustrates a mathematical expression for determining the weighted data item in accordance with some implementations of the present disclosure. The vector
representation nti of the data item i at the time step t is provided to the attention layer 210 to obtain a respective attention value ua, where sigmoid(-) represents a sigmoid function and Wn and bn denote a parameter matrix and a parameter vector, respectively. The attention value uu is normalized through the soflmax function to obtain a weight ati. The weight ati can weight the vector representation nti of the respective data item to determine the weighted data item dt at the time step t. The input vectors over N time steps can be denoted as D = [dil i e [1, N].
[0027] As an example, as shown in Fig. 2, when the time step t=2, the attention layer 210 can determine weights a2 1, a2 2, and a2 L (also known as importance or attention) for the data items n2 1 202, n2 2 204, and n2 L 206. The weights a2 1 , a2 2 and a2 L can weight the respective data items n2 1 202, n2 2 204, and n2 L 206 to determine the weighted data item d2 224 associated with the time step t=2. The weighted data item d2 224 can be provided to the respective input of the recurrent neural network 230. In a similar way, respective weighted data items d± 222, . , dN 226 can be obtained for the time steps t=l, . , t=N.
[0028] In some implementations, the recurrent neural network 230 can be a bidirectional recurrent neural network. Fig. 3 illustrates a recurrent neural network 230 in accordance with some implementations of the present disclosure. The recurrent neural network units 314, 324, and 334 in the recurrent neural network 230 can receive, at each time step, respective input data and the latent state at the previous time step. For example, at the time step t=2, the recurrent neural network unit 324 receives, from the recurrent neural network unit 314, the latent state
at the previous time step and respective input data (e.g., weighted data item d2) and determines the latent state
at the time step based on the latent state and the input data (e.g., weighted data item d2). Likewise, at the time step t=N, the recurrent neural network unit 334 receives, from the recurrent neural network unit at the time step t=N-l, the latent state hN-1 and respective input data (e.g., weighted data item dN ) and determines the latent state h^ at the time step t=N based on the latent state hN-1 and the input data (e.g., weighted data item dN).
[0029] In addition, the recurrent neural network units 312, 322 and 332 in the recurrent neural network 230 can receive, at each time step, the respective input data and the latent state at the next time step. For example, at the time step t=l, the recurrent neural network unit 312 receives, from the recurrent neural network unit 322, the latent state
at the next time step and respective input data (e.g., weighted data item άc) and determines the latent state at the time step based on the latent state
and the input data (e.g., weighted data item d±). Similarly, at the time step t=N-l, the recurrent neural network unit can receive, from the recurrent neural network unit 332 at the time step t=N, the latent state
and the respective input data (e.g., weighted data item dN-1) and determine, based on the latent state and the input data (e.g., weighted data item dN-1), the latent state hN-1 at the time step t=N-l .
[0030] In Fig. 3, at the time step t=l, the recurrent neural network unit 310 is shown to include respective recurrent neural network units 312 and 314; at the time step t=2, the recurrent neural network unit 320 is shown to include respective recurrent neural network units 322 and 324; likewise, at the time step t=N, the recurrent neural network unit 330 is shown to include respective recurrent neural network units 332 and 334. However, it should be understood that the neural network unit at each time step can be the same neural network unit and the neural network unit is shown at each time step for the purpose of illustration only.
[0031] In some implementations, the recurrent neural network unit can be Gated Recurrent Unit (GRU). At the time step t, the GRU computes the state ht at the time step by linearly interpolating the state ht-1 of the previous time step and the current updated state ht as shown in equation (2):
where zt represents an update gate for determining how much historical information should be kept and how much new information should be added and * denotes element-wise multiplication.
[0032] The updated current state ht can be computed by combining the input vector at the time step and the state at the previous time step as shown in equation (3):
where rt represents a reset gate for controlling how many historical states should be used to update the new state and Wh, Uh, and bh denote parameter matrices and vectors.
[0033] The update gate zt and the reset gate rt can be computed through equation (4):
zt = a(Wzdt + Uzht- 1 + bz) ^
where s(·) indicates an activation function and Wr , Ur , br , Wz, Uz , and bz denote parameter matrices and vectors.
[0034] At each time step, the latent vector for the time step can be obtained via the GRU. In order to obtain historical and future information for the data, the latent vectors in both directions can be concatenated to build a bidirectional encoded vector ht as an output of the convolutional neural network 230 according to equation (5):
[0035] Although the bidirectional recurrent neural network is introduced above with reference to Figs. 2 and 3, it should be appreciated that implementations of the present disclosure also can be applied in a uni-directional recurrent neural network. For example, the reverse recurrent neural network unit can be omitted and the latent state of each time step is provided for the subsequent neural network layers. Alternatively, the latent state at the last time step is provided for the subsequent neural network. Accordingly, the attention layer 240, which will be subsequently described in details, can be omitted in this implementation.
[0036] In some implementations, as shown in Fig. 2, the attention layer 240 can determine the weights of data items provided by a plurality of outputs of the recurrent neural network 230 and weight the corresponding data items using these weights. For example, more recent data may have a big influence over the prediction result while the less recent data may have a small influence over the prediction result. The data item provided by the output of the recurrent neural network 230 can be the bidirectional encoded vector
as mentioned above. The attention layer 240 considers the factor that the importance of the information varies for the trend determination at the respective time step. By means of the attention layer 240, the vector representation (vector V) of the time series data can be calculated via the equation (6):
where Wh and bh represent parameter matrix and vector; 0L denotes parameters at each time step in the soflmax layer indicating which time step is more important; oL indicates the latent representation of the bidirectional encoded vector h^. By combining 0L and oL through the soflmax layer, the attention vector b is acquired to weight the outputs of different time. Accordingly, the attention vector b is employed to weight the bidirectional encoded vector hi to compute the vector representation V of the time series data for subsequent classification. As the vector representation V is a weighted value of the bidirectional encoded vector
the vector representation V can have the same dimension as the latent state of the recurrent neural network. The dimension can be pre-specified.
[0037] The discriminative network 260 can be a classifier of Multi-Layer Perceptron (MLP), which receives the vector V as the input and outputs the classification result. The classification result can include a probability that the obj ect is associated with the predefined label. For example, the label with the maximum probability can be output as the classification result.
[0038] In the training process, the training difficulty varies for different training data. For example, some training data lacks sufficient information and the training of such training data can be quite challenging. In some implementations, some challenging training data can be skipped and the training data which can be easily trained is selected and the challenging training data can be progressively added to the training process. For example, the neural network 200 can be trained using the self-paced learning, which can automatically choose the suitable training samples for different training phases, so as to enhance the model performance.
[0039] The self-paced learning process is introduced below with reference to several mathematical expressions. However, it should be understood that these mathematical expressions are provided as examples only, which enables those skilled in the art to better understand and implement implementations of the present disclosure without limiting the scope of the present disclosure. In some implementations, given a training set
where Xj & Rm denotes all inputs in the i-th sample and yL represents
a corresponding trend label. L(yi g(Xi, w)) represents the loss function between the label yi and the output g(xi, w) of the entire model. An importance weight vt can be assigned to each sample xL in the training set.
[0040] In some implementations, the goal of the self-paced learning is to jointly learn the model parameter w and the latent weight v = [v1, . , vn] via the equation (7):
where E(w, v, ) represents an objective function,
denotes a regularization term for self-paced learning (or known as penalty), which controls the learning solution for penalizing the latent weight variables and l is a hyper-parameter that controls the pace at which the model learns new samples. For example, the regularization term f(v: A) can be computed using the equation (8):
[0041] In some implementations, Alternative Convex Search (ACS) can be used to solve equation (7). ACS divides the variables into two disjoint blocks. In each iteration, a block of variables is optimized while the other block is fixed. With the fixed w, the unconstrained solution for the linear regularization term (8) can be computed by the equation (9):
where v- denotes the i-th element in the iterated optimal solution, and lt denotes the loss for each element L(yi, g(xi, w)). The latent weight for samples that are different to what model has already learned will receive a linear penalty.
[0042] In some implementations, it is required that the self-paced learning inputs the data set D, the linear regularizer f and the self-paced step m with the expectation of obtaining the model parameter w. In the initialization procedure, v can be initialized evenly and w* and v* can be iteratively updated in accordance with w*— argmin^E(w, v* , A) and V * = argminvE(w* , V, A). In each iteration, l, if being too small, can be increased based on the step m. Finally, w* can be output as the trained model parameter. As the training progresses and A increases, samples with larger loss will be gradually incorporated, thereby realizing the effective and efficient learning.
[0043] Fig. 4 is a flowchart illustrating a method 400 for predicting the trend in
accordance with some implementations of the present disclosure. The method 400 can be implemented by the computing device 100 or the predicting module 122 shown by Fig. 1. The concepts introduced above with reference to Figs. 2 and 3, alone or in combination, can be applied into the method 400. At 402, time series data associated with an object is obtained.
[0044] At 404, a vector representation of the time series data is computed with a recurrent neural network. In some implementations, computing the vector representation of the time series data includes: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion. For example, the vector representation of the time series data can be determined by for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0045] In some implementations, computing the vector representation of the time series data includes: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions. For example, the recurrent neural network can be a bidirectional recurrent neural network, which also includes a plurality of outputs corresponding to a plurality of inputs. Determining the vector representation of the time series data includes determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
[0046] At 406, a probability that the object is associated with a predefined label is determined based on the vector representation of the time series data. The label indicates a class of a trend in which the object varies over time. In this way, the prediction problem can be transformed into a classification problem.
[0047] At 408, a prediction for the trend in which the object varies over time is generated based on the probability. For example, if the probability associated with a certain label exceeds a predetermined threshold, it is considered that the trend of the object is identical to the label.
[0048] Fig. 5 is a flowchart illustrating a method 500 for training a neural network in accordance with some implementations of the present disclosure. The method 500 can be implemented by the computing device 100 or the predicting module 122 shown by Fig. 1. The concepts introduced above with reference to Figs. 2 and 3, alone or in combination, can be applied into the method 500.
[0049] At 502, the time series data associated with an object and a predefined label for the object are obtained. The label can indicate a class of a trend in which the object varies over time.
[0050] At 504, the recurrent neural network for predicting the trend in which the object varies over time is updated based on the time series data and the label. The updating comprises computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
[0051] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
[0052] In some implementations, determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0053] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
[0054] In some implementations, the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality
of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
[0055] In some implementations, updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
[0056] Some example implementations of the present disclosure are listed below.
[0057] In accordance with some implementations, there is provided a computer- implemented method. The method comprises: The method comprises obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
[0058] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
[0059] In some implementations, determining the vector representation of the time series data based on the relationship of importance comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0060] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
[0061] In some implementations, the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items
with the weights to determine the vector representation of the time series data.
[0062] In accordance with some implementations, there is provided a computer- implemented method. The method comprises obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend that the object varies over time; and updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising: computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
[0063] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
[0064] In some implementations, determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0065] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
[0066] In some implementations, the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
[0067] In some implementations, updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
[0068] In accordance with some implementations, there is provided a device. The device comprises: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: obtaining time series data associated with an object; computing a vector representation of the time series data with a recurrent neural network; determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and generating, based on the probability, a prediction for the trend in which the object varies over time.
[0069] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
[0070] In some implementations, determining the vector representation of the time series data based on the relationship of importance comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0071] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
[0072] In some implementations, the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
[0073] In accordance with some implementations, there is provided a device. The device comprises: a processing unit; and a memory coupled to the processing unit and comprising
instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend that the object varies over time; and updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising: computing a vector representation of the time series data with the recurrent neural network; and determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
[0074] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
[0075] In some implementations, determining, based on the relationship of importance, the vector representation of the time series data comprises: for the respective portion of the plurality of portions: determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
[0076] In some implementations, computing the vector representation of the time series data comprises: segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
[0077] In some implementations, the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises: determining weights of a plurality of data items provided by the plurality of outputs; and weighting the plurality of data items with the weights to determine the vector representation of the time series data.
[0078] In some implementations, updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
[0079] In accordance with some implementations, there is provided a computer program
product comprising instructions stored tangibly on a computer-readable medium, the instructions, when executed by a machine, causing the machine to execute the method described above.
[0080] The functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), and the like. Additionally, the functions described above herein can be at least partially executed by a Graphical Processing Unit (GPU).
[0081] Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions/operations specified in the flowcharts and/or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0082] In the context of this disclosure, a machine readable medium may be any tangible medium that may contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine readable medium may be a machine readable signal medium or a machine readable storage medium. A machine readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0083] Further, although operations are depicted in a particular order, it should be understood that the operations are required to be executed in the shown particular order or in a sequential order, or all shown operations are required to be executed to achieve the expected results. In certain circumstances, multitasking and parallel processing may be
advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation may also be implemented in multiple implementations separately or in any suitable sub- combination.
[0084] Although the subject matter has been described in language specific to structural features and/or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A computer-implemented method, comprising:
obtaining time series data associated with an object;
computing a vector representation of the time series data with a recurrent neural network;
determining, based on the vector representation of the times series data, a probability that the object is associated with a predefined label, wherein the label indicates a class of a trend in which the object varies over time; and
generating, based on the probability, a prediction for the trend in which the object varies over time.
2. The method of claim 1, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and
determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
3. The method of claim 2, wherein determining the vector representation of the time series data based on the relationship of importance comprises:
for the respective portion of the plurality of portions:
determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and
providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
4. The method of claim 1, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and
determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
5. The method of claim 4, wherein the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein
determining the vector representation of the time series data comprises:
determining weights of a plurality of data items provided by the plurality of outputs; and
weighting the plurality of data items with the weights to determine the vector representation of the time series data.
6. A computer-implemented method, comprising:
obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend that the object varies over time; and
updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising:
computing a vector representation of the time series data with the recurrent neural network; and
determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
7. The method of claim 6, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and
determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
8. The method of claim 7, wherein determining, based on the relationship of importance, the vector representation of the time series data comprises:
for the respective portion of the plurality of portions:
determining weights of the data items in the respective portion; weighting the data items in the respective portion with the weights to determine the weighted data items associated with the respective portion; and
providing the weighted data items to an input of the plurality of inputs associated with the respective portion.
9. The method of claim 6, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network; and
determining the vector representation of the time series data based on relationship of
importance between the plurality of portions.
10. The method of claim 9, wherein the recurrent neural network comprises a bidirectional recurrent neural network and the bidirectional recurrent neural network further comprises a plurality of outputs corresponding to the plurality of inputs, and wherein determining the vector representation of the time series data comprises:
determining weights of a plurality of data items provided by the plurality of outputs; and
weighting the plurality of data items with the weights to determine the vector representation of the time series data.
11. The method of claim 6, wherein updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
12. A device, comprising:
a processing unit; and
a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising:
obtaining time series data associated with an object and a predefined label for the object, the label indicating a class of a trend in which the object varies over time; and updating, based on the time series data and the label, a recurrent neural network for predicting the trend in which the object varies over time, the updating comprising:
computing a vector representation of the time series data with the recurrent neural network; and
determining, based on the vector representation of the time series data, a probability that the object is associated with the label.
13. The device of claim 12, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a plurality of inputs of the recurrent neural network, a respective portion of the plurality of portions comprising data items from a plurality of sources; and
determining the vector representation of the time series data based on relationship of importance between the data items in the respective portion.
14. The device of claim 12, wherein computing the vector representation of the time series data comprises:
segmenting the time series data into a plurality of portions corresponding to a
plurality of inputs of the recurrent neural network; and
determining the vector representation of the time series data based on relationship of importance between the plurality of portions.
15. The device of claim 12, wherein updating the recurrent neural network comprises updating the recurrent neural network through self-paced learning.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810716540.8A CN110659759A (en) | 2018-06-29 | 2018-06-29 | Neural network based trend prediction |
| CN201810716540.8 | 2018-06-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020005599A1 true WO2020005599A1 (en) | 2020-01-02 |
Family
ID=67211872
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2019/037408 Ceased WO2020005599A1 (en) | 2018-06-29 | 2019-06-17 | Trend prediction based on neural network |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110659759A (en) |
| WO (1) | WO2020005599A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117851823A (en) * | 2024-01-08 | 2024-04-09 | 山东大学 | A method and system for improving the sampling rate of a source measurement unit by interpolation based on time series prediction |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110993118A (en) * | 2020-02-29 | 2020-04-10 | 同盾控股有限公司 | Epidemic situation prediction method, device, equipment and medium based on ensemble learning model |
| SG10202008469RA (en) * | 2020-09-01 | 2020-10-29 | Ensign Infosecurity Pte Ltd | A deep embedded self-taught learning system and method for detecting suspicious network behaviours |
| CN113349791B (en) * | 2021-05-31 | 2024-07-16 | 平安科技(深圳)有限公司 | Abnormal electrocardiosignal detection method, device, equipment and medium |
-
2018
- 2018-06-29 CN CN201810716540.8A patent/CN110659759A/en active Pending
-
2019
- 2019-06-17 WO PCT/US2019/037408 patent/WO2020005599A1/en not_active Ceased
Non-Patent Citations (4)
| Title |
|---|
| AKITA RYO ET AL: "Deep learning for stock prediction using numerical and textual information", 2016 IEEE/ACIS 15TH INTERNATIONAL CONFERENCE ON COMPUTER AND INFORMATION SCIENCE (ICIS), IEEE, 26 June 2016 (2016-06-26), pages 1 - 6, XP032948521, DOI: 10.1109/ICIS.2016.7550882 * |
| LU JIANG ET AL: "Self-Paced Curriculum Learning", PROCEEDINGS / EIGHTEENTH NATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE (AAAI-2002), FOURTEENTH INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE (IAAI-2002) : [JULY 28 - AUGUST 1, 2002, EDMONTON, ALBERTA, CANADA], 25 January 2015 (2015-01-25), US, pages 2694 - 2700, XP055475953, ISBN: 978-0-262-51129-2 * |
| YOSHIHARA AKIRA ET AL: "Predicting Stock Market Trends by Recurrent Deep Neural Networks", 1 December 2014, INTERNATIONAL CONFERENCE ON COMPUTER ANALYSIS OF IMAGES AND PATTERNS. CAIP 2017: COMPUTER ANALYSIS OF IMAGES AND PATTERNS; [LECTURE NOTES IN COMPUTER SCIENCE; LECT.NOTES COMPUTER], SPRINGER, BERLIN, HEIDELBERG, PAGE(S) 759 - 769, ISBN: 978-3-642-17318-9, XP047311129 * |
| ZINIU HU ET AL: "Listening to Chaotic Whispers : A Deep Learning Framework for News-oriented Stock Trend Prediction", WEB SEARCH AND DATA MINING, 6 December 2017 (2017-12-06), 2 Penn Plaza, Suite 701New YorkNY10121-0701USA, pages 261 - 269, XP055621582, ISBN: 978-1-4503-5581-0, DOI: 10.1145/3159652.3159690 * |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117851823A (en) * | 2024-01-08 | 2024-04-09 | 山东大学 | A method and system for improving the sampling rate of a source measurement unit by interpolation based on time series prediction |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110659759A (en) | 2020-01-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12050887B2 (en) | Information processing method and terminal device | |
| EP3834137B1 (en) | Committed information rate variational autoencoders | |
| US11307864B2 (en) | Data processing apparatus and method | |
| US11307865B2 (en) | Data processing apparatus and method | |
| JP7224447B2 (en) | Encoding method, apparatus, equipment and program | |
| US11887008B2 (en) | Contextual text generation for question answering and text summarization with supervised representation disentanglement and mutual information minimization | |
| CN108629414B (en) | Deep hash learning method and device | |
| CN112085186A (en) | Neural network quantitative parameter determination method and related product | |
| US20180204120A1 (en) | Improved artificial neural network for language modelling and prediction | |
| WO2020005599A1 (en) | Trend prediction based on neural network | |
| US12093659B2 (en) | Text generation with customizable style | |
| US12205001B2 (en) | Random classification model head for improved generalization | |
| US20250077182A1 (en) | Arithmetic apparatus, operating method thereof, and neural network processor | |
| CN112149809A (en) | Model hyper-parameter determination method and device, calculation device and medium | |
| US20240220867A1 (en) | Incorporation of decision trees in a neural network | |
| WO2024054639A1 (en) | Compositional image generation and manipulation | |
| CN114792388A (en) | Image description character generation method and device and computer readable storage medium | |
| CN119003876A (en) | Method, apparatus, device and storage medium for content recommendation | |
| JP7536574B2 (en) | Computing device, computer system, and computing method | |
| US20200302303A1 (en) | Optimization of neural network in equivalent class space | |
| US20240403636A1 (en) | Self-attention based neural networks for processing network inputs from multiple modalities | |
| US20200110635A1 (en) | Data processing apparatus and method | |
| CN109063934B (en) | Artificial intelligence-based combined optimization result obtaining method and device and readable medium | |
| CN115511042B (en) | Training method and device for realizing continuous learning neural network model and electronic equipment | |
| CN121835771A (en) | Text sequence prediction method, system, storage medium and electronic equipment |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19737345 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19737345 Country of ref document: EP Kind code of ref document: A1 |





