WO2020066724A1 - データ処理装置、データ処理方法およびプログラム - Google Patents
データ処理装置、データ処理方法およびプログラム Download PDFInfo
- Publication number
- WO2020066724A1 WO2020066724A1 PCT/JP2019/036262 JP2019036262W WO2020066724A1 WO 2020066724 A1 WO2020066724 A1 WO 2020066724A1 JP 2019036262 W JP2019036262 W JP 2019036262W WO 2020066724 A1 WO2020066724 A1 WO 2020066724A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- vector
- unit
- estimation model
- learning
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H10/00—ICT specially adapted for the handling or processing of patient-related medical or healthcare data
- G16H10/60—ICT specially adapted for the handling or processing of patient-related medical or healthcare data for patient-specific data, e.g. for electronic patient records
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/10—Pre-processing; Data cleansing
- G06F18/15—Statistical pre-processing, e.g. techniques for normalisation or restoring missing data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/17—Function evaluation by approximation methods, e.g. inter- or extrapolation, smoothing, least mean square method
- G06F17/175—Function evaluation by approximation methods, e.g. inter- or extrapolation, smoothing, least mean square method of multidimensional data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/18—Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/213—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
- G06F18/2134—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on separation criteria, e.g. independent component analysis
- G06F18/21342—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on separation criteria, e.g. independent component analysis using statistical independence, i.e. minimising mutual information or maximising non-gaussianity
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/217—Validation; Performance evaluation; Active pattern learning techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0499—Feedforward networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
Definitions
- One embodiment of the present invention relates to a data processing device, a data processing method, and a program for effectively utilizing data including loss.
- IoT Internet of Things
- home appliances such as sphygmomanometers and scales are connected to the network, and an environment is being established in which health data such as blood pressure and weight measured in daily life are collected through the network. is there.
- the health data often recommends periodic measurement, and often includes information indicating the date and time of measurement together with the measured value.
- the health data is likely to be lost due to forgetting to measure or malfunction of the measuring device. This deficiency causes a decrease in accuracy in analyzing health data.
- the data is reduced as one of the problems.
- the analysis is performed ignoring the loss, such as when the size of the entire acquired data is small or when the ratio of the loss to the size of the entire data is large, the effective data may be reduced in a small amount.
- FIG. 4 shows an example of blood pressure measurement data for five days including such data loss.
- blood pressure is set to be measured three times a day, data with no loss is obtained on June 22 and 26, but the second time is measured on the 23rd.
- the third data is missing on the 24th, the third data is missing on the 24th day, and all data is missing on the 25th day.
- it is determined that the data on the day that has been lost even once has been ignored only the data for two days out of the data for five days cannot be used as valid data for analysis.
- the present invention has been made in view of the above circumstances, and an object of the present invention is to provide a data processing device, a data processing method, and a program for effectively utilizing data including loss.
- a first aspect of the present invention is a data processing device, comprising: a data acquisition unit that acquires a series of data including a defect; A statistic calculation unit that calculates a representative value of data and an effective rate representing a ratio of valid data, an output obtained by inputting the representative value and the effective rate to an estimation model, And a learning unit that learns the estimation model so as to minimize an error based on a difference from the value.
- the learning unit connects a predetermined number of representative values and an effective rate corresponding to each of the representative values to the estimation model.
- An input vector composed of elements is input.
- the learning unit is configured such that: X is a vector having the predetermined number of representative values as elements, W is a vector having an effective rate corresponding to each element of X, Y is the input vector, and the input vector is input to the estimation model.
- the statistical data is obtained from the series of data for each aggregation unit.
- the representative value of the data and the effective rate representing the ratio of valid data that are calculated by the amount calculating unit are input to the learned estimation model, and the output from the intermediate layer of the estimation model according to the input is input. Is output as a feature amount of the series of data.
- the statistical data is obtained from the series of data for each aggregation unit.
- the representative value of the data calculated by the amount calculating unit and the effective rate representing the ratio of valid data are input to the learned estimation model, and the output from the estimation model according to the input is
- the apparatus further includes a second estimating unit that outputs loss data as estimated interpolation data.
- the estimation model is learned so as to minimize an error based on a difference between the representative value and an output value obtained by inputting an input value based on the representative value and the effective rate to the estimation model.
- an input vector composed of an element obtained by connecting a predetermined number of representative values and an effective rate corresponding to each representative value is input to the estimation model, Used for model learning.
- a vector X having a predetermined number of representative values as elements, a vector W having an effective rate corresponding to each element of X as an element, and the input vector as an estimation model
- the effective rate is applied to both the input side vector X and the output side vector Y, and the estimation model can be learned using an error that clearly considers the degree of loss.
- the fourth aspect of the present invention when a series of data including a defect to be estimated is obtained, a representative value of data for each aggregation unit calculated from the series of data and valid data exist.
- the effective rate indicating the ratio of the estimation model is input to the learned estimation model, and the output from the intermediate layer of the estimation model corresponding to the input is output as the feature amount of the series of data.
- the fifth aspect of the present invention when a series of data including a defect to be estimated is acquired, a representative value of data for each aggregation unit calculated from the series of data and valid data exist.
- the effective rate indicating the ratio of the estimation model is input to the learned estimation model, and the output from the estimation model according to the input is output as estimation data obtained by interpolating the loss.
- FIG. 1 is a block diagram showing a functional configuration of a data processing device according to an embodiment of the present invention.
- FIG. 2 is a flowchart illustrating an example of a processing procedure of a learning phase performed by the data processing apparatus illustrated in FIG. 1 and details of the processing.
- FIG. 3 is a flowchart illustrating an example of a procedure of an estimation phase performed by the data processing apparatus illustrated in FIG.
- FIG. 4 is a diagram illustrating an example of data including loss.
- FIG. 5 is a diagram illustrating an example of a result of calculating a statistic on a daily basis from data including loss.
- FIG. 6 is a diagram illustrating an example of an estimation model and inputs and outputs to the estimation model.
- FIG. 1 is a block diagram showing a functional configuration of a data processing device according to an embodiment of the present invention.
- FIG. 2 is a flowchart illustrating an example of a processing procedure of a learning phase performed by the data processing apparatus illustrated in FIG. 1 and details of the processing.
- FIG. 7 is a diagram illustrating an example of a result of calculating a statistic in units of aggregation every three days from data including a defect.
- FIG. 8 is a diagram illustrating a first example of input vector generation.
- FIG. 9 is a diagram illustrating a second example of input vector generation.
- FIG. 10 is a diagram illustrating a first example of input vector generation based on a plurality of types of data.
- FIG. 11 is a diagram illustrating a second example of input vector generation based on a plurality of types of data.
- FIG. 1 is a block diagram showing a functional configuration of a data processing device 1 according to one embodiment of the present invention.
- the data processing device 1 is managed by, for example, a medical institution or a health management center, and is configured by, for example, a server computer or a personal computer.
- the data processing apparatus 1 can acquire a series of data including a defect (also referred to as a “data group”) such as health data via the network NW or an input device (not shown).
- the data processing device 1 may be installed alone, but may include a terminal of a medical worker such as a doctor, an electronic medical records (EMR) server installed for each medical institution, and a plurality of medical institutions.
- An electronic health record (EHR) server installed in each area including the service, a cloud server of a service provider, or the like may be provided as one of the extended functions.
- the data processing device 1 may be provided as one of its extended functions in a user terminal or the like owned by the user.
- the data processing device 1 includes an input / output interface unit 10, a control unit 20, and a storage unit 30.
- the input / output interface unit 10 includes, for example, one or more wired or wireless communication interface units, and enables transmission and reception of information with external devices.
- the wired interface for example, a wired LAN is used
- the wireless interface for example, an interface adopting a low-power wireless data communication standard such as a wireless LAN or Bluetooth (registered trademark) is used.
- the input / output interface unit 10 receives data transmitted from a measuring device such as a sphygmomanometer having a communication function, or accesses a database server to read stored data. Then, a process of passing the data to the control unit 20 as an analysis target is performed.
- the input / output interface unit 10 can also perform a process of outputting instruction information input by an input device (not shown) such as a keyboard to the control unit 20.
- the input / output interface unit 10 outputs the learning result or the estimation result output from the control unit 20 to a display device (not shown) such as a liquid crystal display, or transmits the learning result or the estimation result to an external device via the network NW. It can be performed.
- the storage unit 30 uses, as a storage medium, a non-volatile memory that can be written and read at any time, such as a hard disk drive (HDD) or a solid state drive (SSD).
- a storage area necessary for the program a data storage unit 31, a statistic storage unit 32, and a model storage unit 33 are provided in addition to the program storage unit.
- the data storage unit 31 is used to store a data group to be analyzed obtained through the input / output interface unit 10.
- the statistic storage unit 32 is used to store the statistic calculated from the data group.
- the model storage unit 33 is used to store an estimation model for estimating a data group in which a loss has been interpolated from a data group including a loss.
- the storage units 31 to 33 are not indispensable components, and the data processing device 1 may obtain necessary data from a measuring device or a user device as needed.
- the storage units 31 to 33 do not have to be built in the data processing device 1, and may be provided in an external storage medium such as a USB memory or a storage device such as a database server arranged in a cloud. May be obtained.
- the control unit 20 has a hardware processor such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), not shown, and a memory such as a DRAM (Dynamic Random Access Memory) and an SRAM (Static Random Access Memory).
- the processing functions required to implement this embodiment include a data acquisition unit 21, a statistic calculation unit 22, a vector generation unit 23, a learning unit 24, an estimation unit 25, and an output control unit 26. ing. All of these processing functions are realized by causing the processor to execute a program stored in the storage unit 30.
- the control unit 20 may also be implemented in various other forms, including an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (field-programmable gate array).
- the data acquisition unit 21 performs a process of acquiring a data group to be analyzed via the input / output interface unit 10 and storing the data group in the data storage unit 31.
- the statistic calculation unit 22 reads the data stored in the data storage unit 31, calculates a statistic for each predetermined aggregation unit, and stores the calculated result in the statistic storage unit 32.
- the statistic includes a representative value of the data included in each aggregation unit and an effective rate indicating a ratio of valid data included in each aggregation unit.
- the vector generation unit 23 performs a process of reading the statistic stored in the statistic storage unit 32 and generating a vector including a predetermined number of elements.
- the vector generation unit 23 generates a vector X having a predetermined number of representative values as elements and a vector W having an effective rate corresponding to each element of the vector X as an element.
- the vector generation unit 23 outputs the generated vector X and the vector W to the learning unit 24 in the learning phase, and outputs to the estimating unit 25 in the estimation phase.
- the learning unit 24 reads the estimation model stored in the model storage unit 33, inputs the vector X and the vector W received from the vector generation unit 23 to the estimation model, and learns each parameter of the estimation model. Perform the following processing.
- the learning unit 24 inputs a vector obtained by connecting the elements of the vector X and the elements of the vector W to the estimation model, and acquires a vector Y output from the estimation model in response to the input.
- the learning unit 24 learns each parameter of the estimation model so as to minimize an error calculated based on the difference between the vector X and the vector Y, and updates the estimation model stored in the model storage unit 33 as needed. Perform the following processing.
- the estimation unit 25 reads the learned estimation model stored in the model storage unit 33, inputs the vector X and the vector W received from the vector generation unit 23 to the estimation model, and performs a data estimation process. I do.
- the estimating unit 25 inputs a vector obtained by connecting the elements of the vector X and the elements of the vector W to the learned estimation model, and according to the input, the vector Y or the intermediate layer output from the estimation model. Is output to the output control unit 26 as an estimation result.
- the output control unit 26 performs a process of outputting the vector Y or the feature amount Z output from the estimation unit 25. Alternatively, the output control unit 26 can output the parameters related to the learned estimated model stored in the model storage unit 33.
- the data processing apparatus 1 can operate as a learning phase or an estimation phase, for example, by receiving an instruction signal input from an operator through an input device or the like.
- FIG. 2 is a flowchart showing a processing procedure and processing contents of the learning phase by the data processing device 1.
- step S201 the data processing device 1 converts a series of data including a defect into the learning data through the input / output interface unit 10 under the control of the data acquisition unit 21. And stores the obtained data in the data storage unit 31.
- FIG. 4 shows, as an example of acquired and stored data, blood pressure measurement results for a specific user for five days, with a measurement frequency set to three times a day.
- the term “three times a day” may be measured in different time zones, such as immediately after getting up, before lunch, before going to bed, or may be measured three times in the same time zone.
- the blood pressure measurement value may be any measurement value such as systolic blood pressure, diastolic blood pressure, and pulse pressure. It should be noted that the numerical values shown in FIG. 4 are merely examples for explanation, and are not intended to represent a specific health condition.
- the acquired data may include a user ID, a device ID, information indicating a measurement date and time, and the like, along with numerical data indicating a blood pressure measurement value.
- FIG. 4 for convenience, a serial number is assigned to each record for one day, and a description regarding the loss is added.
- the symbol “ ⁇ ” means that valid data does not exist or data is missing.
- three data are measured on June 22 (# 1) and 26 (# 5), and there is no loss.
- the data processing device 1 reads the data stored in the data storage unit 31 under the control of the statistics calculation unit 22 in step S202, and To calculate the statistic.
- the aggregation unit is arbitrarily set by the operator, designer, manager, or the like of the data processing device 1, for example, for each type of data, and is stored in the storage unit 30.
- the statistic calculation unit 22 reads the setting of the aggregation unit stored in the storage unit 30, divides the data read from the data storage unit 31 into each aggregation unit, and calculates the statistic.
- FIG. 5 shows a representative value as a statistic and an effective rate calculated using the data shown in FIG.
- a totaling unit for each day is set, and an average value is set as a representative value.
- the representative value is not limited to this, and arbitrary statistics such as a median, a maximum, a minimum, a mode, a variance, and a standard deviation can be used.
- what kind of statistics should be calculated can be set in advance by the administrator or the like.
- the average value of valid data in the aggregation unit is calculated as the representative value.
- blood pressure measurement data 110, 111, 111 for three times was obtained on June 22 (# 1)
- a representative value "122" 122/1) was calculated as an average value between valid data.
- no measurement data was acquired on June 25 (# 4)
- "NA" indicating that calculation was not possible is shown.
- the result calculated by the statistic calculation unit 22 as described above can be stored in the statistic storage unit 32 as statistic data in association with, for example, an identification number for identifying the aggregation unit or date information.
- the totaling unit is not limited to one day, and any unit can be adopted. For example, it may be set to an arbitrary time width such as several hours, three days, or one week, or may be a unit defined by the number of data including loss without using time information. . Further, the counting units may overlap each other. For example, in association with a specific date, the moving average may be calculated from data of two days before and on the day before the date.
- step S203 the data processing device 1 reads out the statistic data stored in the statistic storage unit 32 under the control of the vector generation unit 23, and performs learning of the estimation model.
- a process of generating two types of vectors (vector X and vector W) to be used is performed.
- the vector generation unit 23 selects a preset number (n) of aggregation units from the read statistic data, extracts a representative value and an effective rate from each of the n aggregation units, and A vector X (x 1 , x 2 ,..., X n ) having a representative value as an element and a vector W (w 1 , w 2 , ..., w n ).
- the number n of elements corresponds to one half of the number of input dimensions of the estimation model to be learned, and the number of input dimensions of the estimation model can be arbitrarily determined by the designer or administrator of the data processing apparatus 1.
- Can be set to The number N of the generated vector pairs (the vector X and the vector W) corresponds to the number of samples of the learning data, and the number N can also be set arbitrarily.
- the vector generation unit 23 sets the first vector pair to, for example, # 1 to # 3 , A representative value is extracted to generate a vector X 1 (110.6667, 122, 121.5), and an effective rate is extracted to generate a vector W 1 (1, 0.333, 0.666). Further, the vector generation unit 23 selects, for example, a totaling unit of # 2 to # 4 as a second vector pair, and calculates a vector X 2 (122, 121.5, 0) and a vector W 2 (0.333, 0.666, 0). Can be generated.
- the representative value “NA” can be replaced with 0 when the vector is generated.
- the counting units selected at the time of vector generation may or may not overlap each other. Instead of setting the number N of vector pairs to be generated, the number of vector pairs corresponding to all selectable combinations may be set from the read statistical data.
- the vector generation unit 23 outputs the vector pair (vector X and vector W) generated as described above to the learning unit 24.
- step S204 under the control of the learning unit 24, the data processing device 1 reads out the learning target estimation model stored in advance in the model storage unit 33, and The vector X and the vector W received from 23 are input to the estimation model and learning is performed.
- the estimation model to be learned can be arbitrarily set by a designer, a manager, or the like.
- a hierarchical neural network is used as the estimation model.
- FIG. 6 shows an example of such a neural network and images of input and output vectors for the example.
- the estimation model shown in FIG. 6 includes an input layer, three intermediate layers, and an output layer, and the number of units is set to 10, 3, 2, 3, and 5, respectively. However, the details of the number of these units are merely set for convenience of explanation, and can be arbitrarily set according to the nature of the data to be analyzed, the purpose of the analysis, the work environment, and the like.
- the number of intermediate layers is not limited to three, and the number of layers other than three can be arbitrarily selected to form the intermediate layer.
- each element of an input vector is input to each node of an input layer, weighted and added together, biased to enter a node of the next layer, and an activation function is applied at the node.
- the weighting factor is A
- the bias is B
- the activation function is f
- the output Q of the intermediate layer (first layer) when P is input to the input layer is generally expressed by the following equation.
- Q f (AP + B) (1)
- a vector obtained by connecting the elements of the vector X and the elements of the vector W is input to the input layer.
- a vector X 110.6667, 122, 121.5, 0, 115.3333
- the input vector 110.6667, # 122, # 121.5, # 0, # 115.3333, # 1, # 0.333, # 0.666, # 0, # 1 obtained by connecting these elements is input to the estimation model.
- Y represents an output vector from the estimation model, and has the same number of elements as the vector X. Therefore, in this embodiment, since the number of elements of the vector X and the vector W are the same, the number of output dimensions of the estimation model is ⁇ of the number of input dimensions. In the example of FIG. 6, the number of units in the intermediate layer is designed to be smaller than that in the input layer and the output layer.
- Z represents the feature amount of the intermediate layer.
- the feature quantity Z is obtained as an output from a node in the hidden layer, and can be expressed based on the above equation (1).
- the suffix 1 or 2 means a parameter that contributes to the output of the first layer or the second layer, respectively.
- the feature amount Z obtained from the trained model in which the number of units in the hidden layer is smaller than that in the input layer is a useful value that represents the essential features of the input data in a smaller dimension. It is known that such information can be useful information.
- equation (4) the vector W of the effective rate is applied to both the input side vector X and the output side vector Y, and the degree of loss in the data is taken into account when learning the estimation model. I understand.
- the learning unit 24 learns the estimation model as a self-encoder (auto encoder) so that the output from the output layer reproduces the input as much as possible.
- the learning unit 24 can learn the estimation model so as to minimize the error L by using a stochastic gradient descent method such as Adam or AdaDelta, for example.
- the learning unit 24 is not limited thereto. Can be used.
- the learning unit 24 performs a process of updating the estimation model stored in the model storage unit 33 in step S205.
- the data processing device 1 outputs each parameter of the learned model stored in the model storage unit 33 through the output control unit 26 under the control of the control unit 20 in response to, for example, input of an instruction signal from an operator. It may be configured as follows.
- the data processing device 1 can perform data estimation based on a newly acquired data group including a defect using the learned model stored in the model storage unit 33. It becomes possible.
- FIG. 3 is a flowchart showing a processing procedure and processing contents of the estimation phase by the data processing device 1. A detailed description of the same processing as in FIG. 2 is omitted.
- step S301 the data processing apparatus 1 performs a series of operations including a defect through the input / output interface unit 10 under the control of the data acquisition unit 21 as in step S201. Is obtained as estimation data, and the obtained data is stored in the data storage unit 31.
- step S302 the data processing device 1 reads out the data stored in the data storage unit 31 under the control of the statistics calculation unit 22 and performs the setting in the same manner as in step S202. A process of calculating a statistic is performed for each of the tabulated units. It is preferable to use the same setting as that used in the learning phase, but it is not necessarily limited thereto. Similarly, as the representative value, it is preferable to use the same representative value used in the learning phase (for example, the average value between valid data in the above example), but it is not necessarily limited to this.
- the statistic calculation unit 22 associates the calculation result with, for example, an identification number for identifying the tabulation unit and date information, and calculates the statistics as statistic data It can be stored in the quantity storage unit 32.
- step S303 the data processing device 1 reads out the statistic data stored in the statistic storage unit 32 under the control of the vector generation unit 23 as in step S203.
- the vector generation unit 23 selects a set number (n) of aggregation units from the read statistic data, extracts a representative value and an effective rate from each of the n aggregation units, and extracts the n representative units.
- a vector X (x 1 , x 2 ,..., X n ) having values as elements, and a vector W (w 1 , w 2 ,...) Having n effective rates corresponding to each element of the vector X. .., w n ).
- the number n of elements may be obtained by storing the value of n used for learning or by multiplying the number of input dimensions of the learned model stored in the model storage unit 33 by 1 /. Can be.
- the vector generation unit 23 outputs the generated vector pair (vector X and vector W) to the estimation unit 25.
- step S304 the data processing device 1 reads the learned estimation model stored in the model storage unit 33 under the control of the estimation unit 25, and The received vector X and vector W are input to the learned estimation model, and a process of obtaining an output vector Y output from the estimation model is performed on the input.
- the output vector Y (110.0, 122.2, 122.4, 0.1, 114.9) is output from the estimation model.
- Each element of the input vector X is replaced by a numerical value in consideration of the effective rate in the vector Y.
- step S305 the data processing device 1 inputs the estimation result by the estimation unit 25 under the control of the output control unit 26, for example, in response to the input of an instruction signal from the operator. Output can be performed via the output interface unit 10.
- the output control unit 26 acquires, for example, an output vector Y output from the estimation model, and outputs the acquired output vector Y to a display device such as a liquid crystal display as a data group obtained by interpolating the loss corresponding to the input data group.
- the data can be transmitted to an external device via the network NW.
- the output control unit 26 can extract the feature amount Z of the intermediate layer corresponding to the input data group and output this.
- the feature amount Z can be considered to represent an essential feature in the input data group with a smaller dimension than the original input data group. Therefore, by using the feature amount Z as an input to any other learning device, it is possible to perform processing with a reduced load as compared with the case where the original input data group is used as it is.
- the learning device is used for a logistic regression, a support vector machine, a classifier such as a random forest, or a regression model using a multiple regression analysis or a regression tree.
- a series of data including a loss is acquired by the data acquisition unit 21, and a statistical amount calculation unit 22 compiles a series of data from the series of data for each predetermined aggregation unit.
- a representative value of the data and an effective rate indicating a ratio of valid data are calculated.
- the loss is not represented by a binary value with / without, but is represented by a continuous value as a ratio.
- the vector generation unit 23 generates a vector X having a representative value extracted from a predetermined number n of aggregation units as an element and a vector W having a corresponding effective rate as an element.
- the learning unit 24 inputs an input vector obtained by connecting the elements of the vector X and the elements of the vector W to the estimation model, and minimizes the error L based on the vector Y output from the estimation model with respect to the input. As a result, learning of the estimation model is performed as an auto encoder.
- the aggregation unit can be effectively utilized without discarding and used for learning. Can be reduced. This is particularly advantageous when the ratio of loss is large relative to the size of the entire data or when the size of the entire data is small.
- learning can be performed on the representative value of each tabulation unit in consideration of the degree of loss for each tabulation unit. As shown in Expression (4), learning is performed so that the contribution of data having a large loss is reduced by W included in the error L, so that the data is effectively used by effectively using the degree of the loss. be able to.
- the vector generation unit 23 uses the vector X having a representative value extracted from a predetermined number n of aggregation units as an element, and the vector W having the effective rate corresponding thereto as an element. Is generated. Then, an input vector obtained by connecting the elements of the vector X and the elements of the vector W is input by the estimating unit 25 to the learned estimation model trained as described above, and output from the estimation model in accordance with the input. The obtained vector Y or the feature amount Z output from the hidden layer is obtained.
- the original Estimation processing can be performed by effectively utilizing the data without discarding it and by considering the degree of the loss.
- neither the learning phase nor the estimation phase requires an overly complicated operation for calculating the statistic and generating the input vector. It is possible for an administrator or the like to perform any setting or correction according to the above.
- the vector generation unit 23 has been described as generating the vector X and the vector W by extracting the representative value and the effective rate calculated for each tabulation by a predetermined number of elements.
- the vector X may be generated from the raw data before calculating the statistics.
- the vector X 1 (110, 111, 111) can be generated by directly extracting the measurement value from the record of # 1.
- the vector W 1 corresponding, for example, the # 1 record with "1" as the effective rate since there is no defect, it is possible to generate a vector W 1 (1, 1, 1).
- the vector X 2 (122, 0, 0) can be generated from the record # 2 in FIG.
- the vector W 2 (0.333, 0.333, 0.333) was obtained using “0.333” as the effective rate.
- the vector W 2 (1, 0, 0) may be generated on the assumption that only the first measurement value is valid.
- FIG. 7 shows an example of a method for calculating a statistic when the totaling unit is three days.
- an average value and an effective rate for three days before and after are calculated as a total unit from measurement data representing body weight measured every day. That is, in FIG. 7, for # 2 linked on June 23, the average value (representative value) “60.5” for the three days from June 22 to 24 is the same as the effective rate for the same three days. (Ratio of valid data) “0.666” is calculated as a statistic.
- the generation of the vector by the vector generation unit 23 is not limited to the above-described embodiment.
- 8 and 9 show an example of extracting five-dimensional data from time-series data for generating a vector.
- the original data is divided every five days and input to the estimation model as shown in FIG.
- data for five days is extracted while being shifted one day at a time to obtain an input vector.
- FIGS. 10 and 11 show an example of input vector generation from two types of data (data A and data B).
- data A is assumed to include data on health such as blood pressure and weight, test values such as blood glucose and urine test values, and answers to questionnaires (questionnaires).
- Sensor data such as sleep time measured by a wearable device, position information measured by GPS or the like, and answers to a questionnaire (questionnaire) are assumed.
- Data A blood pressure measurement value
- step count measurement value data as “Data B”
- the above embodiment is not limited to such health-related data, and various types of data acquired in various fields such as manufacturing, transportation, and agriculture can be used.
- FIG. 10 when there are two types of data, it is possible to configure so as to generate an input vector by connecting data extracted from each of them.
- the first three dimensions are assigned to data A and the second three dimensions are assigned to data B, and data for three days extracted from each of data A and data B is input vector.
- the case where the extraction is performed while being shifted in the same period as the input dimension is described. However, the input may be performed while being shifted by one day as described above with reference to FIG. 9.
- the example in FIG. 10 is applicable even when there are more than two types of data.
- a plurality of data may be assigned to respective input channels and input. This is realized by a general method used when inputting image data to a neural network when one pixel has three pieces of information like an RGB image.
- the time-series data that is recorded particularly every day is described as an example.
- the recording frequency of the data does not need to be one day, and the data recorded at an arbitrary frequency may be used. Can be.
- the above embodiment can be applied to data other than the time-series data.
- temperature data recorded for each observation point or image data may be used.
- data represented by a two-dimensional array such as image data as described in the case where there are a plurality of types of data, it is realized by extracting, connecting, and inputting each line.
- the above-described embodiment can be applied to a total result of a questionnaire or a test.
- a questionnaire it is expected that data will be missing for some questions, or that data will be completely unanswered for certain subjects, for reasons such as not being applicable or not wanting to answer. You.
- learning and estimation can be performed by effectively utilizing data without discarding, while discriminating between partially unanswered and completely unanswered data.
- the data can be digitized by any method, such as analyzing the frequency of appearance of keywords using text mining, and the above embodiment can be applied.
- the functional units 21 to 26 included in the data processing device 1 may be distributed and arranged in a cloud computer, an edge router, or the like, and the learning and estimation may be performed by cooperating with each other. As a result, the processing load on each device can be reduced, and the processing efficiency can be increased.
- the present invention is not limited to the above-described embodiment as it is, and can be embodied by modifying its constituent elements in an implementation stage without departing from the scope of the invention.
- Various inventions can be formed by appropriately combining a plurality of constituent elements disclosed in the above embodiments. For example, some components may be deleted from all the components shown in the embodiment. Further, components of different embodiments may be appropriately combined.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Evolutionary Computation (AREA)
- Mathematical Physics (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Health & Medical Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Biomedical Technology (AREA)
- General Health & Medical Sciences (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Computing Systems (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Databases & Information Systems (AREA)
- Medical Informatics (AREA)
- Probability & Statistics with Applications (AREA)
- Public Health (AREA)
- Algebra (AREA)
- Primary Health Care (AREA)
- Epidemiology (AREA)
- Operations Research (AREA)
- Pathology (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Complex Calculations (AREA)
Abstract
欠損を含むデータ群を有効に活用するデータ処理装置を提供する。欠損を含む一連の学習用データを取得し、当該一連の学習用データから、あらかじめ定められた集計単位ごとに、データの代表値と、有効なデータが存在する割合を表す有効率とを算出し、前記代表値および前記有効率を推定モデルに入力して得られる出力と、前記代表値との差に基づく誤差を最小化するように、前記推定モデルを学習する。また、欠損を含む一連の推定用データを取得し、当該一連の推定用データから、あらかじめ定められた集計単位ごとに、データの代表値と、有効なデータが存在する割合を表す有効率とを算出し、前記代表値および前記有効率を学習済みの推定モデルに入力して、一連の推定用データについて特徴量の取得またはデータ推定を行う。
Description
この発明の一態様は、欠損を含むデータを有効に活用するための、データ処理装置、データ処理方法、およびプログラムに関する。
IoT(Internet of Things)技術の発展に伴い、例えば血圧計や体重計などの家電製品がネットワークに接続され、日常生活で計測した血圧や体重などの健康データがネットワークを通じて収集される環境が整いつつある。健康データは、定期的な計測が推奨されることが多く、また計測値とともに計測日時を表す情報を含むことが多い。ここで、健康データには、計測し忘れや計測機器の不具合などによってデータが欠損しやすいという課題がある。この欠損は、健康データを解析する上で精度の低下等をまねく原因になる。
欠損を考慮したデータ解析として、欠損を表す配列を用いて、欠損の無い部分のみで誤差を最小化することで、欠損の影響を考慮した学習方法が提案されている(例えば、特許文献1参照)。
ところが、欠損を含むデータの解析では、課題の1つとしてデータが削減されてしまうことが考えられる。特に、取得されたデータ全体のサイズが小さい場合や、データ全体のサイズに対して欠損の割合が大きい場合など、欠損を無視して解析すると、有効なデータが少量になってしまうことがある。
例えば、血圧のように1日に複数回計測される健康データでは、1日の計測値のうち一部が欠損する場合がある。図4は、そのようなデータの欠損を含む5日分の血圧計測データの例を示す。図4の例では、1日3回の血圧を計測するように設定されている場合に、6月22日と26日には欠損のないデータが得られているが、23日は2回目と3回目のデータが、24日は3回目のデータが、25日はすべてのデータがそれぞれ欠損している。このようなケースで、例えば1回でも欠損した日のデータを無視すると決めると、5日間のデータのうち2日分のデータしか有効なデータとして解析に使用できなくなってしまう。
もう1つの課題が、欠損の度合いが考慮されないことである。例えば、図4の場合、欠損が1回だけの日から3回すべて欠損している日まで、欠損の程度に差がある。しかし、欠損の有無だけで判断すると、これらの日はすべて欠損ありとして判断されてしまう。集計単位が大きくなるほど、欠損の有無だけでなく欠損の度合いを適切に表現することが重要となり得る。
この発明は上記事情に着目してなされたもので、その目的とするところは、欠損を含むデータを有効に活用するための、データ処理装置、データ処理方法、およびプログラムを提供することにある。
上記課題を解決するために、この発明の第1の態様は、データ処理装置にあって、欠損を含む一連のデータを取得するデータ取得部と、上記一連のデータから、あらかじめ定められた集計単位ごとに、データの代表値と有効なデータが存在する割合を表す有効率とを算出する統計量算出部と、上記代表値および上記有効率を推定モデルに入力して得られる出力と、上記代表値との差に基づく誤差を最小化するように上記推定モデルを学習する学習部と、を具備するようにしたものである。
この発明の第2の態様は、上記第1の態様において上記学習部が、上記推定モデルに対し、あらかじめ定められた個数の代表値と、当該代表値の各々に対応する有効率とを連結した要素からなる入力ベクトルを入力するようにしたものである。
この発明の第3の態様は、上記第2の態様において上記学習部が、
Xを、上記あらかじめ定められた個数の代表値を要素とするベクトル、Wを、Xの各要素に対応する有効率を要素とするベクトル、Yを、上記入力ベクトルを上記推定モデルに入力して得られる出力ベクトルと、それぞれ定義したときに、次式:
L=|W・(Y-X)|2で表される誤差Lを最小化するように上記推定モデルを学習するようにしたものである。
Xを、上記あらかじめ定められた個数の代表値を要素とするベクトル、Wを、Xの各要素に対応する有効率を要素とするベクトル、Yを、上記入力ベクトルを上記推定モデルに入力して得られる出力ベクトルと、それぞれ定義したときに、次式:
L=|W・(Y-X)|2で表される誤差Lを最小化するように上記推定モデルを学習するようにしたものである。
この発明の第4の態様は、上記第1の態様において、上記データ取得部により推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから上記集計単位ごとに上記統計量算出部により算出される、データの代表値と有効なデータが存在する割合を表す有効率とを学習済みの上記推定モデルに入力し、当該入力に応じた上記推定モデルの中間層からの出力を、上記一連のデータの特徴量として出力する、第1の推定部をさらに具備するようにしたものである。
この発明の第5の態様は、上記第1の態様において、上記データ取得部により推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから上記集計単位ごとに上記統計量算出部により算出される、データの代表値と有効なデータが存在する割合を表す有効率とを学習済みの上記推定モデルに入力し、当該入力に応じた上記推定モデルからの出力を、上記欠損を補間した推定データとして出力する、第2の推定部をさらに具備するようにしたものである。
この発明の第1の態様によれば、欠損を含む一連のデータから、あらかじめ定められた集計単位ごとに、データの代表値と、有効なデータが存在する割合を表す有効率とが算出され、代表値と有効率とに基づく入力値を推定モデルに入力して得られる出力値と、前記代表値との差に基づく誤差を最小化するように、推定モデルが学習される。
これにより、取得された一連のデータが欠損を含む場合でも、あらかじめ定められた集計単位ごとに統計量としての代表値および有効率を算出して学習に用いることにより、データを破棄することなく、集計単位ごとの情報としてすべてのデータを有効に活用することができる。また、単に欠損があるかないかだけでなく、集計単位ごとに有効なデータが存在する割合が算出されて学習に用いられるので、欠損の度合いまで考慮に入れた効果的な学習を行うことができる。
この発明の第2の態様によれば、あらかじめ定められた個数の代表値と、各代表値に対応する有効率とを連結した要素からなる入力ベクトルが、推定モデルに対して入力され、当該推定モデルの学習に用いられる。これにより、学習用のデータ群が規則性のない欠損を含む場合でも、複雑なデータ処理を要することなく、各集計単位の代表値と有効率とを確実に対応付けて学習を行うことができる。
この発明の第3の態様によれば、あらかじめ定められた個数の代表値を要素とするベクトルXと、Xの各要素に対応する有効率を要素とするベクトルWと、上記入力ベクトルを推定モデルに入力して得られるベクトルYとから算出される誤差L=|W・(Y-X)|2を最小化するように、推定モデルの学習が行われる。これにより、入力側のベクトルXおよび出力側のベクトルYの両方に有効率が適用され、欠損の度合いを明確に考慮した誤差を用いて、推定モデルの学習を行うことができる。
この発明の第4の態様によれば、推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから算出される集計単位ごとのデータの代表値と有効なデータが存在する割合を表す有効率とが学習済みの推定モデルに入力され、当該入力に応じた推定モデルの中間層からの出力が上記一連のデータの特徴量として出力される。これにより、欠損を含む一連のデータについて、欠損の度合いまでも考慮に入れた特徴量を得ることができ、当該一連のデータの特徴をより的確に把握することができる。
この発明の第5の態様によれば、推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから算出される集計単位ごとのデータの代表値と有効なデータが存在する割合を表す有効率とが学習済みの推定モデルに入力され、当該入力に応じた推定モデルからの出力が、欠損を補間した推定データとして出力される。これにより、欠損を含む一連のデータについて、欠損の度合いまでも考慮に入れた推定結果を得ることができる。
すなわちこの発明の各態様によれば、欠損を含むデータを有効に活用する技術を提供することができる。
以下、図面を参照してこの発明に係わる実施形態を説明する。
[一実施形態]
(構成)
図1は、この発明の一実施形態に係るデータ処理装置1の機能構成を示すブロック図である。
[一実施形態]
(構成)
図1は、この発明の一実施形態に係るデータ処理装置1の機能構成を示すブロック図である。
データ処理装置1は、例えば、医療機関や保健管理センター等によって管理されるもので、例えばサーバコンピュータまたはパーソナルコンピュータにより構成される。データ処理装置1は、ネットワークNWを介して、または図示しない入力デバイスを介して、健康データなど、欠損を含む一連のデータ(「データ群」とも言う)を取得することができる。データ処理装置1は、単独で設置されてもよいが、医師等の医療従事者の端末や、医療機関ごとに設置されている電子医療記録(Electronic Medical Records:EMR)サーバ、複数の医療機関を含む地域ごとに設置される電子健康記録(Electronic Health Records:EHR)サーバ、さらにはサービス事業者のクラウドサーバ等に、その拡張機能の1つとして設けられるものであってもよい。さらには、データ処理装置1は、ユーザが所持するユーザ端末等にその拡張機能の1つとして設けられてもよい。
一実施形態に係るデータ処理装置1は、入出力インタフェースユニット10と、制御ユニット20と、記憶ユニット30とを備える。
入出力インタフェースユニット10は、例えば1つ以上の有線または無線の通信インタフェースユニットを含んでおり、外部機器との間で情報の送受信を可能にする。有線インタフェースとしては、例えば有線LANが使用され、また無線インタフェースとしては、例えば無線LANやBluetooth(登録商標)などの小電力無線データ通信規格を採用したインタフェースが使用される。
例えば、入出力インタフェースユニット10は、制御ユニット20の制御の下、通信機能を備えた血圧計などの計測機器から送信されたデータを受信し、またはデータベースサーバにアクセスして蓄積されたデータを読み出し、そのデータを解析対象として制御ユニット20に渡す処理を行う。入出力インタフェースユニット10はまた、キーボードなどの入力デバイス(図示せず)によって入力された指示情報を制御ユニット20に出力する処理を行うことができる。さらに、入出力インタフェースユニット10は、制御ユニット20から出力された学習結果や推定結果を、液晶ディスプレイなどの表示デバイス(図示せず)に出力したり、ネットワークNWを介して外部機器に送信する処理を行うことができる。
記憶ユニット30は、記憶媒体として、例えばHDD(Hard Disk Drive)またはSSD(Solid State Drive)等の随時書込および読み出しが可能な不揮発性メモリを用いたものであり、この実施形態を実現するために必要な記憶領域として、プログラム記憶部の他に、データ記憶部31と、統計量記憶部32と、モデル記憶部33とを備えている。
データ記憶部31は、入出力インタフェースユニット10を介して取得された、解析対象のデータ群を記憶するために用いられる。
統計量記憶部32は、データ群から算出された統計量を記憶するために用いられる。
モデル記憶部33は、欠損を含むデータ群から欠損を補間したデータ群を推定するための推定モデルを記憶するために用いられる。
ただし、上記記憶部31~33は、必須の構成ではなく、データ処理装置1が計測機器やユーザ機器から必要なデータを随時取得するようにしてもよい。あるいは、上記記憶部31~33は、データ処理装置1に内蔵されたものでなくてもよく、例えば、USBメモリなどの外付け記憶媒体や、クラウドに配置されたデータベースサーバ等の記憶装置に設けられたものであってもよい。
制御ユニット20は、図示しないCPU(Central Processing Unit)やMPU(Micro Processing Unit)等のハードウェアプロセッサと、DRAM(Dynamic Random Access Memory)やSRAM(Static Random Access Memory)等のメモリとを有し、この実施形態を実施するために必要な処理機能として、データ取得部21と、統計量算出部22と、ベクトル生成部23と、学習部24と、推定部25と、出力制御部26とを備えている。これらの処理機能は、いずれも上記記憶ユニット30に格納されたプログラムを上記プロセッサに実行させることにより実現される。制御ユニット20は、また、ASIC(Application Specific Integrated Circuit)やFPGA(field-programmable gate array)などの集積回路を含む、他の多様な形式で実現されてもよい。
データ取得部21は、入出力インタフェースユニット10を介して、解析対象とするデータ群を取得し、データ記憶部31に格納する処理を行う。
統計量算出部22は、データ記憶部31に格納されたデータを読み出し、あらかじめ定められた集計単位ごとに統計量を算出し、算出した結果を統計量記憶部32に格納する処理を行う。一実施形態では、統計量は、各集計単位に含まれるデータの代表値と、各集計単位に含まれる有効なデータの割合を表す有効率とを含む。
ベクトル生成部23は、統計量記憶部32に格納された統計量を読み出し、あらかじめ定められた個数の要素からなるベクトルを生成する処理を行う。一実施形態では、ベクトル生成部23は、あらかじめ定められた個数の代表値を要素とするベクトルXと、ベクトルXの各要素に対応する有効率を要素とするベクトルWとを生成する。ベクトル生成部23は、生成されたベクトルXおよびベクトルWを、学習フェーズにおいては学習部24に出力し、推定フェーズにおいては推定部25に出力する。
学習部24は、学習フェーズにおいて、モデル記憶部33に格納された推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該推定モデルに入力して、推定モデルの各パラメータを学習する処理を行う。一実施形態では、学習部24は、ベクトルXの要素とベクトルWの要素を連結したベクトルを推定モデルに入力し、その入力に応じて当該推定モデルから出力されるベクトルYを取得する。そして、学習部24は、ベクトルXとベクトルYとの差に基づいて算出される誤差を最小化するように推定モデルの各パラメータを学習し、モデル記憶部33に格納された推定モデルを随時更新する処理を行う。
推定部25は、推定フェーズにおいて、モデル記憶部33に格納された学習済みの推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該推定モデルに入力して、データの推定処理を行う。一実施形態では、推定部25は、ベクトルXの要素とベクトルWの要素を連結したベクトルを学習済みの推定モデルに入力し、その入力に応じて当該推定モデルから出力されるベクトルYまたは中間層の特徴量Zを、推定結果として出力制御部26に出力する。
出力制御部26は、推定部25から出力されたベクトルYまたは特徴量Zを出力する処理を行う。あるいは、出力制御部26は、モデル記憶部33に格納された学習済みの推定モデルに関するパラメータを出力することも可能である。
(動作)
次に、以上のように構成されたデータ処理装置1による情報処理動作を説明する。データ処理装置1は、例えば、入力デバイス等を通じて入力されたオペレータからの指示信号を受け付けて、学習フェーズまたは推定フェーズとして動作することができる。
次に、以上のように構成されたデータ処理装置1による情報処理動作を説明する。データ処理装置1は、例えば、入力デバイス等を通じて入力されたオペレータからの指示信号を受け付けて、学習フェーズまたは推定フェーズとして動作することができる。
(1)学習フェーズ
学習フェーズが設定されると、データ処理装置1は、以下のように推定モデルの学習処理を実行する。図2は、データ処理装置1による学習フェーズの処理手順と処理内容を示すフローチャートである。
学習フェーズが設定されると、データ処理装置1は、以下のように推定モデルの学習処理を実行する。図2は、データ処理装置1による学習フェーズの処理手順と処理内容を示すフローチャートである。
(1-1)学習用データの取得
はじめに、データ処理装置1は、ステップS201において、データ取得部21の制御の下、入出力インタフェースユニット10を介して、欠損を含む一連のデータを学習用データとして取得し、取得したデータをデータ記憶部31に格納する。
はじめに、データ処理装置1は、ステップS201において、データ取得部21の制御の下、入出力インタフェースユニット10を介して、欠損を含む一連のデータを学習用データとして取得し、取得したデータをデータ記憶部31に格納する。
図4は、取得され格納されるデータの一例として、1日3回の計測頻度を設定された、特定のユーザの5日分の血圧計測結果を示す。1日3回とは、例えば、起床直後、昼食前、就寝前など、異なる時間帯に計測されるものであってもよいし、同じ時間帯に3回計測が繰り返されるものであってもよい。また、血圧計測値は、収縮期血圧、拡張期血圧、脈圧など、いずれの計測値であってもよい。なお、図4に示した数値は説明のために例示するものにすぎず、特定の健康状態を表すことを意図したものではない。また、取得されるデータは、血圧計測値を表す数値データとともに、ユーザID、装置ID、計測日時を表す情報等を含むこともできる。
なお、図4では、便宜上、1日分のレコードごとに連続番号を付し、欠損に関する説明を付記している。図4において、記号「-」は、有効なデータが存在しない、またはデータが欠損していることを意味する。図4に示すように、6月22日(#1)および26日(#5)には3回分のデータが計測されており欠損はないが、23日(#2)には1回のデータしか計測されておらず、24日(#3)には2回のデータしか計測されておらず、25日(#4)にはまったく計測されていない。
(1-2)統計量の算出
次いで、データ処理装置1は、ステップS202において、統計量算出部22の制御の下、データ記憶部31に格納されたデータを読み出し、あらかじめ設定された集計単位ごとに統計量を算出する処理を行う。集計単位は、データ処理装置1のオペレータ、設計者または管理者等によって、例えばデータの種類ごとに任意に設定され、記憶ユニット30に記憶されているものとする。統計量算出部22は、記憶ユニット30に記憶された集計単位の設定を読み出し、データ記憶部31から読み出したデータを集計単位ごとに分割して、統計量を算出する。
次いで、データ処理装置1は、ステップS202において、統計量算出部22の制御の下、データ記憶部31に格納されたデータを読み出し、あらかじめ設定された集計単位ごとに統計量を算出する処理を行う。集計単位は、データ処理装置1のオペレータ、設計者または管理者等によって、例えばデータの種類ごとに任意に設定され、記憶ユニット30に記憶されているものとする。統計量算出部22は、記憶ユニット30に記憶された集計単位の設定を読み出し、データ記憶部31から読み出したデータを集計単位ごとに分割して、統計量を算出する。
図5は、図4に示したデータを用いて算出された、統計量としての代表値および有効率を示す。ここでは、日ごとの集計単位が設定され、代表値として平均値が設定されている。ただし、代表値はこれだけに限られるものではなく、中央値、最大値、最小値、最頻値、分散や標準偏差など、任意の統計量を用いることができる。集計単位と同様に、どのような種類の統計量を算出すべきかについても、あらかじめ管理者等によって設定しておくことが可能である。
図5に示した例では、代表値として、集計単位内の有効なデータの平均値が算出される。例えば、6月22日(#1)には3回分の血圧計測データ(110,111,111)が得られたので、代表値(平均値)として「110.6667」(=(110+111+111)/3)が算出されている。一方、6月23日(#2)には1回分の血圧計測データ(122)しか得られなかったので、有効なデータ間の平均値として代表値「122」(=122/1)が算出されている。また、6月25日(#4)には計測データが全く取得されなかったので、算出不可を意味する「NA」が示されている。
有効率は、集計単位内に有効なデータが存在する割合を示す。図5に示したように、集計単位が1日で、1日3回の計測頻度が設定されている場合、3回分の計測データが得られれば「1(=3/3)」、2回ならば「0.666(=2/3)」、1回ならば「0.333(=1/3)」、0回ならば「0(=0/3)」として算出される。
以上のようにして統計量算出部22によって算出された結果は、例えば集計単位を識別する識別番号や日付情報に紐づけて、統計量データとして統計量記憶部32に記憶させることができる。
なお、集計単位は、1日単位に限定されるものではなく、任意の単位を採用することができる。例えば、数時間単位、3日単位、1週間単位など、任意の時間幅に設定されてもよいし、時間情報を用いず、欠損を含めたデータの個数によって定義される単位であってもよい。さらに、集計単位は、互いに重複するものであってもよい。例えば、特定の日付に関連付けて、その日付の前日と当日の2日分のデータから移動平均を算出するように設定されてもよい。
(1-3)ベクトルの生成
次に、データ処理装置1は、ステップS203において、ベクトル生成部23の制御の下、統計量記憶部32に格納された統計量データを読み出し、推定モデルの学習に用いるための2種のベクトル(ベクトルXおよびベクトルW)を生成する処理を行う。
次に、データ処理装置1は、ステップS203において、ベクトル生成部23の制御の下、統計量記憶部32に格納された統計量データを読み出し、推定モデルの学習に用いるための2種のベクトル(ベクトルXおよびベクトルW)を生成する処理を行う。
ベクトル生成部23は、読み出した統計量データから、あらかじめ設定された数(n)の集計単位を選択し、それらn個の集計単位の各々から代表値および有効率を抽出して、n個の代表値を要素とするベクトルX(x1, x2,..., xn)と、ベクトルXの各要素に対応するn個の有効率を要素とするベクトルW(w1, w2,..., wn)とを生成する。要素の数nは、後述するように、学習対象である推定モデルの入力次元数の1/2に対応し、推定モデルの入力次元数は、データ処理装置1の設計者や管理者等が任意に設定することができる。生成されるベクトル対(ベクトルXとベクトルW)の数Nは、学習データのサンプル数に対応し、その数Nもまた任意に設定することができる。
例えば、要素の数n=3、ベクトル対の数N=2と設定された場合、図5に示した例では、ベクトル生成部23は、1つ目のベクトル対として、例えば#1~#3の集計単位を選択し、代表値を抽出してベクトルX1(110.6667, 122, 121.5)を生成し、有効率を抽出してベクトルW1(1, 0.333, 0.666)を生成することができる。さらにベクトル生成部23は、2つ目のベクトル対として、例えば#2~#4の集計単位を選択し、ベクトルX2(122, 121.5, 0)およびベクトルW2(0.333, 0.666, 0)を生成することができる。このように、ベクトル生成の際には、代表値「NA」は0で置き換えることができる。またこのように、ベクトル生成の際に選択される集計単位は互いに重複していても重複していなくてもよい。生成すべきベクトル対の数Nを設定せず、読み出された統計量データから選択可能なすべての組合せに対応する個数のベクトル対を生成するように設定してもよい。
ベクトル生成部23は、以上のように生成したベクトル対(ベクトルXとベクトルW)を学習部24に出力する。
(1-4)推定モデルの学習
次に、データ処理装置1は、ステップS204において、学習部24の制御の下、あらかじめモデル記憶部33に格納された学習対象の推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該推定モデルに入力してその学習を行う。学習対象とする推定モデルは、設計者や管理者等によって任意に設定されることができる。
次に、データ処理装置1は、ステップS204において、学習部24の制御の下、あらかじめモデル記憶部33に格納された学習対象の推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該推定モデルに入力してその学習を行う。学習対象とする推定モデルは、設計者や管理者等によって任意に設定されることができる。
一実施形態では、推定モデルとして階層型ニューラルネットワークが使用される。図6は、そのようなニューラルネットワークの一例と、それに対する入力および出力ベクトルのイメージを示す。図6に示した推定モデルは、入力層と、3層の中間層と、出力層とから構成され、ユニット数はそれぞれ順に10、3、2、3、5と設定されている。ただし、これらのユニット数の詳細は、説明のために便宜的に設定したものにすぎず、解析対象とするデータの性質や解析の目的、作業環境等に応じて任意に設定することができる。また、中間層については3層に限定されるものではなく、3層以外の層数を任意に選択して中間層を構成することができる。
ニューラルネットワークでは、一般に、入力層の各ノードに入力ベクトルの各要素が入力され、それぞれ重みづけされて足し合わされ、バイアスを付加されて次の層のノードに入り、当該ノードで活性化関数を適用後に出力される。したがって、重み係数をA、バイアスをB、活性化関数をfとすると、入力層にPが入力されたときの中間層(第1層)の出力Qは、一般に、次式で表される。
Q=f(AP+B) (1)
Q=f(AP+B) (1)
この実施形態では、入力層には、ベクトルXの要素とベクトルWの要素とを連結したベクトルが入力される。図6に示した例では、図5のデータから要素数n=5としてベクトルX(110.6667, 122, 121.5, 0, 115.3333)、およびベクトルW(1, 0.333, 0.666, 0, 1)が生成され、これらの要素を連結した入力ベクトル(110.6667, 122, 121.5, 0, 115.3333, 1, 0.333, 0.666, 0, 1)が推定モデルに入力される。
図6において、Yは、推定モデルからの出力ベクトルを表し、ベクトルXと同じ要素数を有する。したがって、この実施形態では、ベクトルXとベクトルWの要素数が同一であることから、推定モデルの出力次元数は、入力次元数の1/2となっている。図6の例ではまた、入力層および出力層に比べて中間層のユニット数が小さくなるように設計されている。
図6において、Zは、中間層の特徴量を表す。特徴量Zは、中間層のノードからの出力として得られ、上式(1)に基づいて表すことができる。例えば、図6の例で、中間層(第1層)の特徴量Z1は、
Z1=f1(A1P+B1) (2)
で表され、中間層(第2層)の特徴量Z2は、
Z2=f2(A2(f1(A1P+B1))+B2) (3)
で表される。なお、添え字1または2は、それぞれ第1層または第2層の出力に寄与するパラメータであることを意味する。
Z1=f1(A1P+B1) (2)
で表され、中間層(第2層)の特徴量Z2は、
Z2=f2(A2(f1(A1P+B1))+B2) (3)
で表される。なお、添え字1または2は、それぞれ第1層または第2層の出力に寄与するパラメータであることを意味する。
特徴量は、一般に、入力されたデータにどのような特徴があるかを表す。図6に示したように、入力層よりも中間層のユニット数の方が少ない学習済みモデルから得られる特徴量Zは、入力されたデータの本質的な特徴をより少ない次元で表した、有益な情報となり得ることが知られている。
学習部24は、このような推定モデルに対して、上記のようにベクトルXの要素とベクトルWの要素を連結した入力ベクトルを入力し、その入力に対して推定モデルから出力される出力ベクトルYを取得する。そして、学習部24は、生成されたすべてのベクトル対(ベクトルXとベクトルW)について、次式(4)を用いて算出される誤差Lを最小化するように、推定モデルのパラメータ(重み係数やバイアスなど)を学習する。
L=|W・(Y-X)|2 (4)
L=|W・(Y-X)|2 (4)
式(4)において、入力側のベクトルXおよび出力側のベクトルYの両方に有効率のベクトルWが適用されており、推定モデルを学習する際にデータ中の欠損の度合いが考慮されていることがわかる。
このように、学習部24では、出力層からの出力ができるだけ入力を再現したものとなるように、推定モデルが自己符号化器(オートエンコーダ)として学習される。ここで、学習部24は、例えばAdamやAdaDeltaなどの確率的勾配降下法を用いて、上記誤差Lを最小化するように推定モデルを学習することができるが、これに限るものではなく、他の任意の手法を用いることができる。
(1-5)モデルの更新
誤差Lを最小化するように推定モデルのパラメータが決定されたら、学習部24は、ステップS205において、モデル記憶部33に格納された推定モデルを更新する処理を行う。データ処理装置1は、例えばオペレータからの指示信号の入力に応答して、モデル記憶部33に格納された学習済みモデルの各パラメータを、制御ユニット20の制御の下、出力制御部26を通じて出力するように構成してもよい。
誤差Lを最小化するように推定モデルのパラメータが決定されたら、学習部24は、ステップS205において、モデル記憶部33に格納された推定モデルを更新する処理を行う。データ処理装置1は、例えばオペレータからの指示信号の入力に応答して、モデル記憶部33に格納された学習済みモデルの各パラメータを、制御ユニット20の制御の下、出力制御部26を通じて出力するように構成してもよい。
上記学習フェーズが終了すると、データ処理装置1は、モデル記憶部33に格納された学習済みモデルを用いて、新たに取得された欠損を含むデータ群をもとに、データの推定を行うことが可能となる。
(2)推定フェーズ
推定フェーズが設定されると、データ処理装置1は、学習済みモデルを用いて以下のようにデータの推定処理を実行することができる。図3は、データ処理装置1による推定フェーズの処理手順と処理内容を示すフローチャートである。なお、図2と同様の処理については詳細な説明は省略する。
推定フェーズが設定されると、データ処理装置1は、学習済みモデルを用いて以下のようにデータの推定処理を実行することができる。図3は、データ処理装置1による推定フェーズの処理手順と処理内容を示すフローチャートである。なお、図2と同様の処理については詳細な説明は省略する。
(2-1)推定用データの取得
はじめに、データ処理装置1は、ステップS301において、ステップS201と同様に、データ取得部21の制御の下、入出力インタフェースユニット10を介して、欠損を含む一連のデータを推定用データとして取得し、取得したデータをデータ記憶部31に格納する。
はじめに、データ処理装置1は、ステップS301において、ステップS201と同様に、データ取得部21の制御の下、入出力インタフェースユニット10を介して、欠損を含む一連のデータを推定用データとして取得し、取得したデータをデータ記憶部31に格納する。
(2-2)統計量の算出
次いで、データ処理装置1は、ステップS302において、ステップS202と同様に、統計量算出部22の制御の下、データ記憶部31に格納されたデータを読み出し、設定された集計単位ごとに統計量を算出する処理を行う。集計単位は、学習フェーズで用いたのと同じ設定を用いることが好ましいが、必ずしもそれに限定されるわけではない。同様に、代表値は、学習フェーズで用いたのと同じ代表値(例えば上記の例では有効なデータ間の平均値)を用いることが好ましいが、必ずしもそれに限定されるわけではない。集計単位ごとに統計量として代表値および有効率が算出されたら、統計量算出部22は、その算出結果を、例えば集計単位を識別する識別番号や日付情報に紐づけて、統計量データとして統計量記憶部32に記憶させることができる。
次いで、データ処理装置1は、ステップS302において、ステップS202と同様に、統計量算出部22の制御の下、データ記憶部31に格納されたデータを読み出し、設定された集計単位ごとに統計量を算出する処理を行う。集計単位は、学習フェーズで用いたのと同じ設定を用いることが好ましいが、必ずしもそれに限定されるわけではない。同様に、代表値は、学習フェーズで用いたのと同じ代表値(例えば上記の例では有効なデータ間の平均値)を用いることが好ましいが、必ずしもそれに限定されるわけではない。集計単位ごとに統計量として代表値および有効率が算出されたら、統計量算出部22は、その算出結果を、例えば集計単位を識別する識別番号や日付情報に紐づけて、統計量データとして統計量記憶部32に記憶させることができる。
(2-3)ベクトルの生成
次に、データ処理装置1は、ステップS303において、ステップS203と同様に、ベクトル生成部23の制御の下、統計量記憶部32に格納された統計量データを読み出し、推定を行うための2種のベクトル(ベクトルXおよびベクトルW)を生成する処理を行う。
次に、データ処理装置1は、ステップS303において、ステップS203と同様に、ベクトル生成部23の制御の下、統計量記憶部32に格納された統計量データを読み出し、推定を行うための2種のベクトル(ベクトルXおよびベクトルW)を生成する処理を行う。
ベクトル生成部23は、読み出した統計量データから、設定された数(n)の集計単位を選択し、それらn個の集計単位の各々から代表値および有効率を抽出して、n個の代表値を要素とするベクトルX(x1, x2,..., xn)と、ベクトルXの各要素に対応するn個の有効率を要素とするベクトルW(w1, w2,..., wn)とを生成する。要素の数nは、例えば、学習に用いたnの値を記憶しておくか、またはモデル記憶部33に格納された学習済みモデルの入力次元数に1/2を乗じた値として取得することができる。
ベクトル生成部23は、生成したベクトル対(ベクトルXとベクトルW)を推定部25に出力する。
(2-4)データの推定
次に、データ処理装置1は、ステップS304において、推定部25の制御の下、モデル記憶部33に格納された学習済みの推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該学習済みの推定モデルに入力して、その入力に対して推定モデルから出力される出力ベクトルYを取得する処理を行う。学習フェーズで説明したのと同様に、図6に示した出力ベクトルYは、次式で表される。
Y=f4(A4(f3(A3(f2(A2(f1(A1P+B1))+B2))+B3))+B4) (5)
次に、データ処理装置1は、ステップS304において、推定部25の制御の下、モデル記憶部33に格納された学習済みの推定モデルを読み出し、ベクトル生成部23から受け取ったベクトルXおよびベクトルWを当該学習済みの推定モデルに入力して、その入力に対して推定モデルから出力される出力ベクトルYを取得する処理を行う。学習フェーズで説明したのと同様に、図6に示した出力ベクトルYは、次式で表される。
Y=f4(A4(f3(A3(f2(A2(f1(A1P+B1))+B2))+B3))+B4) (5)
図6に示した例では、推定モデルから出力ベクトルY(110.0, 122.2, 122.4, 0.1, 114.9)が出力される。入力されたベクトルXの各要素が、ベクトルYでは有効率を考慮した数値に置き換わっており、特に、ベクトルX中のx4=0(欠損)がベクトルYではy4=0.1に置き換わっている。
(2-5)推定結果の出力
データ処理装置1は、ステップS305において、例えばオペレータからの指示信号の入力に応答して、出力制御部26の制御の下、推定部25による推定結果を、入出力インタフェースユニット10を介して出力することができる。出力制御部26は、例えば、推定モデルから出力された出力ベクトルYを取得し、これを、入力データ群に対応する欠損を補間されたデータ群として、液晶ディスプレイなどの表示デバイスに出力したり、ネットワークNWを介して外部機器に送信することができる。
データ処理装置1は、ステップS305において、例えばオペレータからの指示信号の入力に応答して、出力制御部26の制御の下、推定部25による推定結果を、入出力インタフェースユニット10を介して出力することができる。出力制御部26は、例えば、推定モデルから出力された出力ベクトルYを取得し、これを、入力データ群に対応する欠損を補間されたデータ群として、液晶ディスプレイなどの表示デバイスに出力したり、ネットワークNWを介して外部機器に送信することができる。
あるいは、出力制御部26は、入力データ群に対応する中間層の特徴量Zを抽出し、これを出力することもできる。特徴量Zは、上述のように、入力データ群について、元の入力データ群よりも少ない次元で本質的な特徴を表したものと考えることができる。したがって、特徴量Zを任意の別の学習器の入力として用いることにより、元の入力データ群をそのまま用いる場合に比べて負荷を軽減した処理を行うことができる。そのような任意の別の学習器として、例えば、ロジスティック回帰やサポートベクターマシン、ランダムフォレストのような分類器や、重回帰分析や回帰木などを用いた回帰モデルへの活用が想定される。
(効果)
以上詳述したように、この発明の一実施形態では、データ取得部21によって、欠損を含む一連のデータが取得され、統計量算出部22によって、この一連のデータから所定の集計単位ごとに統計量としてデータの代表値と有効なデータが存在する割合を表す有効率とが算出される。この有効率の算出の際、上記実施形態では、欠損をあり/なしの2値で表現するのではなく、割合としての連続値で表現するようにしている。
以上詳述したように、この発明の一実施形態では、データ取得部21によって、欠損を含む一連のデータが取得され、統計量算出部22によって、この一連のデータから所定の集計単位ごとに統計量としてデータの代表値と有効なデータが存在する割合を表す有効率とが算出される。この有効率の算出の際、上記実施形態では、欠損をあり/なしの2値で表現するのではなく、割合としての連続値で表現するようにしている。
そして、学習フェーズにおいては、ベクトル生成部23によって、所定の個数nの集計単位から抽出される代表値を要素とするベクトルXと、それに対応する有効率を要素とするベクトルWとが生成される。次いで、学習部24によって、ベクトルXの要素とベクトルWの要素を連結した入力ベクトルが推定モデルに対して入力され、その入力に対して推定モデルから出力されるベクトルYに基づく誤差Lを最小化するように、オートエンコーダとして推定モデルの学習が行われる。
これにより、推定モデルの学習に際して、集計単位内の一部のデータまたはすべてのデータが欠損している場合でも、その集計単位を破棄することなく有効に活用して学習に用いることができ、データの削減を抑えることができる。これは、欠損の割合がデータ全体のサイズに対して大きい場合や、データ全体のサイズが小さい場合に特に有利である。
さらに、上記実施形態によれば、集計単位ごとの代表値に対し、集計単位ごとの欠損の度合いを考慮して学習を行うことができる。式(4)に示したように、誤差Lに含まれるWによって、欠損の大きいデータの寄与が小さくなるように学習されるので、欠損の度合いまでも効果的に用いてデータを有効に活用することができる。
推定フェーズにおいても、学習フェーズと同様に、ベクトル生成部23によって、所定の個数nの集計単位から抽出される代表値を要素とするベクトルXと、それに対応する有効率を要素とするベクトルWとが生成される。そして、推定部25によって、ベクトルXの要素とベクトルWの要素を連結した入力ベクトルが、上記のように学習された学習済みの推定モデルに対して入力され、その入力に応じて推定モデルから出力されるベクトルYまたは中間層から出力される特徴量Zが取得される。
したがって、欠損を含むデータ群をもとに、学習済みの推定モデルを用いてデータを推定するときにも、または学習済みの推定モデルの中間層から特徴量を取得するときにも、もとのデータを破棄することなく有効に活用して、またその欠損の度合いまでも考慮して、推定処理を行うことができる。
さらに、上記実施形態によれば、学習フェーズおよび推定フェーズのいずれについても、統計量の算出や入力ベクトル生成のために過度に複雑な操作を要求するものではないので、データの性質や分析の目的に応じて管理者等が任意の設定や修正を行って実施することが可能である。
[他の実施形態]
なお、この発明は上記実施形態に限定されるものではない。
なお、この発明は上記実施形態に限定されるものではない。
例えば、図5および図6に関して、ベクトル生成部23が、集計単位ごとに算出された代表値および有効率を所定の要素数だけ抽出してベクトルXおよびベクトルWを生成するものとして説明したが、統計量を算出する前の生データからベクトルXを生成するようにしてもよい。
例えば図4の例では、#1のレコードから計測値をそのまま抽出してベクトルX1(110, 111, 111)を生成することもできる。この場合、対応するベクトルW1として、例えば#1のレコードには欠損がないので有効率として「1」を用いて、ベクトルW1(1, 1, 1)を生成することができる。また同様に、図4の#2のレコードからベクトルX2(122, 0, 0)を生成することができる。この場合、対応するベクトルW2として、#2のレコードでは1回目の計測値しか得られなかったので、有効率として「0.333」を用いて、ベクトルW2(0.333, 0.333, 0.333)を生成することができる。あるいは、1回目の計測値だけが有効であったとしてベクトルW2(1, 0, 0)を生成するようにしてもよい。
また、統計量算出部22が用いる集計単位は、上記実施形態に限定されるものではなく、任意の集計単位を設定することができる。図7は、集計単位を3日としたときの統計量の算出方法の一例を示す。図7では、日ごとに計測された体重を表す計測データから、集計単位として前後3日間の平均値および有効率が算出されている。すなわち、図7において、6月23日に紐づけられた#2については、6月22日~24日の3日間の平均値(代表値)「60.5」と、同じ3日間の有効率(有効データが存在する割合)「0.666」とが統計量として算出されている。同様に、6月27日に紐づけられた#6については、6月26日~28日の3日間に計測データが全く取得されなかったので、代表値として「NA(算出不可)」と、有効率「0」とが算出されている。なお、上述のように、「NA」はベクトル生成時に「0」に置き換えることができる。
さらに、ベクトル生成部23によるベクトルの生成も、上記で説明した実施形態に限定されるものではない。図8および図9は、ベクトル生成のための時系列データからの5次元のデータ抽出の例を示す。図8の例では、元のデータを5日間ごとに分割して、図6に示したような推定モデルに入力するようにしている。図9の例では、5日間のデータを1日ずつずらしながら抽出して入力ベクトルとするようにしている。同様に、2日ずつ、3日ずつ、または4日ずつずらして抽出することも可能であり、他の抽出方法を採用して上記実施形態に適用することも可能である。
またさらに、複数の種類のデータが存在する場合にも、上記実施形態を適用することができる。図10および図11は、2種類のデータ(データAおよびデータB)からの入力ベクトル生成の例を示す。ここでは、「データA」として、血圧値や体重などの健康に関するデータや、血糖値や尿検査値などの検査値、問診(アンケート)の回答などが想定され、「データB」として、歩数や睡眠時間などウェアラブルデバイスで計測されるようなセンサデータや、GPSなどで計測される位置情報、問診(アンケート)の回答などが想定される。例えば、「データA」として血圧計測値データ、「データB」として歩数計測値データを収集し、両者を同時に考慮して解析することにより、被検者の健康管理や病気の予防などに役立てようとする場合が考えられる。ただし、上記実施形態は、このような健康関連データに限るものではなく、製造業、運輸業、農業など、多種多様な分野において取得される多種多様なデータを用いることができる。
図10に示すように、2種類のデータが存在する場合、それぞれから抽出したデータを連結して入力ベクトルを生成するように構成することができる。図10の例では、6次元の入力に対して、前半の3次元をデータA、後半の3次元をデータBに割り当てて、データAおよびデータBそれぞれから抽出した3日間分のデータを入力ベクトルとしている。図10の例では、入力次元と同じ期間でずらしながら抽出した場合を記載したが、図9に関して上述したように1日ずつずらしながら入力してもよい。2種類を超える種類のデータが存在する場合にも、図10の例を適用可能である。
あるいは、図11に示すように、複数のデータをそれぞれ入力のチャネルに割り当てて入力してもよい。これは、RGB画像のように1つのピクセルが3つの情報を持っているときに、画像データをニューラルネットワークに入力する際などに使用される一般的な手法で実現される。
以上の実施形態では、特に1日ごとに記録されるような時系列データを例に記載したが、データの記録頻度は1日である必要はなく、任意の頻度で記録されたデータを用いることができる。
さらに、上述したように時系列データ以外のデータに対して上記実施形態を適用することも可能である。例えば、観測地点ごとに記録された気温データのようなものでもよいし、画像データなどでもよい。画像データのように2次元の配列で表現されるデータの場合は、複数の種類のデータが存在する事例について述べたように、行ごとに抽出して連結して入力することで実現される。
また、アンケートや試験などの集計結果に対して上記実施形態を適用することも可能である。例えば、アンケートの場合、該当なしまたは回答したくないなどの理由により、一部の質問に対してデータが欠損したり、特定の被検者に関して完全に無回答のデータが得られることが予想される。このような場合にも、上記実施形態によれば、一部無回答と完全無回答とを区別して考慮しつつ、データを破棄することなく有効に活用して学習や推定を行うことができる。なお、アンケートの自由回答のようにデータが言語情報を含む場合、テキストマイニングを用いてキーワードの出現頻度を解析するなど、任意の方法でデータを数値化し、上記実施形態を適用することができる。
またさらに、データ処理装置1が備える各機能部の必ずしもすべてを単一の装置に設ける必要はない。例えば、データ処理装置1が備える機能部21~26を、クラウドコンピュータやエッジルータ等に分散配置し、これらの装置が互いに連携することにより学習および推定を行うようにしてもよい。これにより、各装置の処理負荷を軽減し、処理効率を高めることができる。
その他、統計量の算出やデータの格納形式等についても、この発明の要旨を逸脱しない範囲で種々変形して実施可能である。
要するにこの発明は、上記実施形態そのままに限定されるものではなく、実施段階ではその要旨を逸脱しない範囲で構成要素を変形して具体化できる。また、上記実施形態に開示されている複数の構成要素の適宜な組み合せにより種々の発明を形成できる。例えば、実施形態に示される全構成要素から幾つかの構成要素を削除してもよい。さらに、異なる実施形態に亘る構成要素を適宜組み合せてもよい。
1…データ処理装置
10…入出力インタフェースユニット
20…制御ユニット
21…データ取得部
22…統計量算出部
23…ベクトル生成部
24…学習部
25…推定部
26…出力制御部
30…記憶ユニット
31…データ記憶部
32…統計量記憶部
33…モデル記憶部
10…入出力インタフェースユニット
20…制御ユニット
21…データ取得部
22…統計量算出部
23…ベクトル生成部
24…学習部
25…推定部
26…出力制御部
30…記憶ユニット
31…データ記憶部
32…統計量記憶部
33…モデル記憶部
Claims (8)
- 欠損を含む一連のデータを取得する、データ取得部と、
前記一連のデータから、あらかじめ定められた集計単位ごとに、データの代表値と、有効なデータが存在する割合を表す有効率とを算出する、統計量算出部と、
前記代表値および前記有効率を推定モデルに入力して得られる出力と、前記代表値との差に基づく誤差を最小化するように前記推定モデルを学習する、学習部と、
を具備するデータ処理装置。 - 前記学習部は、前記推定モデルに対し、あらかじめ定められた個数の代表値と、当該代表値の各々に対応する有効率とを連結した要素からなる入力ベクトルを入力する、請求項1に記載のデータ処理装置。
- 前記学習部は、
Xを、前記あらかじめ定められた個数の代表値を要素とするベクトル、Wを、Xの各要素に対応する有効率を要素とするベクトル、Yを、前記入力ベクトルを前記推定モデルに入力して得られる出力ベクトルと、それぞれ定義したときに、
次式で表される誤差Lを最小化するように前記推定モデルを学習する、
L=|W・(Y-X)|2
請求項2に記載のデータ処理装置。 - 前記データ取得部により推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから前記集計単位ごとに前記統計量算出部により算出されるデータの代表値と有効なデータが存在する割合を表す有効率とを学習済みの前記推定モデルに入力し、当該入力に応じた前記推定モデルの中間層からの出力を、前記一連のデータの特徴量として出力する、第1の推定部をさらに具備する、請求項1に記載のデータ処理装置。
- 前記データ取得部により推定対象となる欠損を含む一連のデータが取得された場合に、当該一連のデータから前記集計単位ごとに前記統計量算出部により算出されるデータの代表値と有効なデータが存在する割合を表す有効率とを学習済みの前記推定モデルに入力し、当該入力に応じた前記推定モデルからの出力を、前記欠損を補間した推定データとして出力する、第2の推定部をさらに具備する、請求項1に記載のデータ処理装置。
- データ処理装置が実行する、データ処理方法であって、
欠損を含む一連のデータを取得する過程と、
前記一連のデータから、あらかじめ定められた集計単位ごとに、データの代表値と、有効なデータが存在する割合を表す有効率とを算出する過程と、
前記代表値および前記有効率を推定モデルに入力して得られる出力と、前記代表値との差に基づく誤差を最小化するように前記推定モデルを学習する過程と、
を具備するデータ処理方法。 - 前記学習する過程は、
Xを、あらかじめ定められた個数の代表値を要素とするベクトル、Wを、Xの各要素に対応する有効率を要素とするベクトル、Yを、Xの各要素とWの各要素とを連結した要素からなる入力ベクトルを前記推定モデルに入力して得られる出力ベクトルと、それぞれ定義したときに、
次式で表される誤差Lを最小化するように前記推定モデルを学習する、
L=|W・(Y-X)|2
請求項6に記載のデータ処理方法。 - 請求項1乃至5のいずれか一項に記載のデータ処理装置の各部による処理をプロセッサに実行させるプログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/276,767 US20220027686A1 (en) | 2018-09-28 | 2019-09-17 | Data processing apparatus, data processing method, and program |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018-183608 | 2018-09-28 | ||
| JP2018183608A JP7056493B2 (ja) | 2018-09-28 | 2018-09-28 | データ処理装置、データ処理方法およびプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020066724A1 true WO2020066724A1 (ja) | 2020-04-02 |
Family
ID=69952686
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/036262 Ceased WO2020066724A1 (ja) | 2018-09-28 | 2019-09-17 | データ処理装置、データ処理方法およびプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20220027686A1 (ja) |
| JP (1) | JP7056493B2 (ja) |
| WO (1) | WO2020066724A1 (ja) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2023056190A (ja) * | 2021-10-07 | 2023-04-19 | トヨタ自動車株式会社 | 欠損値推定方法、機械学習方法および欠損値推定装置 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022153142A (ja) * | 2021-03-29 | 2022-10-12 | ソニーグループ株式会社 | 情報処理システム、生体試料処理装置及びプログラム |
| US12315290B2 (en) | 2021-07-05 | 2025-05-27 | Nec Corporation | Information processing system, information processing method, biometric matching system, biometric matching method, and storage medium |
| WO2024242197A1 (ja) * | 2023-05-25 | 2024-11-28 | 株式会社Preferred Networks | 情報処理装置、情報処理プログラム及び情報処理方法 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006163521A (ja) * | 2004-12-02 | 2006-06-22 | Research Organization Of Information & Systems | 時系列データ分析装置および時系列データ分析プログラム |
| WO2018047655A1 (ja) * | 2016-09-06 | 2018-03-15 | 日本電信電話株式会社 | 時系列データ特徴量抽出装置、時系列データ特徴量抽出方法及び時系列データ特徴量抽出プログラム |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2010100701A1 (ja) | 2009-03-06 | 2010-09-10 | 株式会社 東芝 | 学習装置、識別装置及びその方法 |
| WO2016164680A2 (en) * | 2015-04-09 | 2016-10-13 | Equifax, Inc. | Automated model development process |
| WO2018005489A1 (en) * | 2016-06-27 | 2018-01-04 | Purepredictive, Inc. | Data quality detection and compensation for machine learning |
| US10592368B2 (en) * | 2017-10-26 | 2020-03-17 | International Business Machines Corporation | Missing values imputation of sequential data |
| US12175345B2 (en) * | 2018-03-06 | 2024-12-24 | Tazi AI Systems, Inc. | Online machine learning system that continuously learns from data and human input |
| WO2019240787A1 (en) * | 2018-06-13 | 2019-12-19 | Nokia Technologies Oy | Generalized virtual pim measurement for enhanced accuracy |
-
2018
- 2018-09-28 JP JP2018183608A patent/JP7056493B2/ja active Active
-
2019
- 2019-09-17 WO PCT/JP2019/036262 patent/WO2020066724A1/ja not_active Ceased
- 2019-09-17 US US17/276,767 patent/US20220027686A1/en not_active Abandoned
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006163521A (ja) * | 2004-12-02 | 2006-06-22 | Research Organization Of Information & Systems | 時系列データ分析装置および時系列データ分析プログラム |
| WO2018047655A1 (ja) * | 2016-09-06 | 2018-03-15 | 日本電信電話株式会社 | 時系列データ特徴量抽出装置、時系列データ特徴量抽出方法及び時系列データ特徴量抽出プログラム |
Non-Patent Citations (1)
| Title |
|---|
| RIMOLDINI, LORENZO, WEIGHTED STATISTICAL PARAMETERS FOR IRREGULARLY SAMPLED TIME SERIES, vol. 3, 6 September 2013 (2013-09-06), pages 1 - 18, XP055700155, Retrieved from the Internet <URL:https://arxiv.org/pdf/1304.6616.pdf> [retrieved on 20191128] * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2023056190A (ja) * | 2021-10-07 | 2023-04-19 | トヨタ自動車株式会社 | 欠損値推定方法、機械学習方法および欠損値推定装置 |
| JP7636312B2 (ja) | 2021-10-07 | 2025-02-26 | トヨタ自動車株式会社 | 欠損値推定方法、機械学習方法および欠損値推定装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20220027686A1 (en) | 2022-01-27 |
| JP2020052886A (ja) | 2020-04-02 |
| JP7056493B2 (ja) | 2022-04-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Thorsen-Meyer et al. | Dynamic and explainable machine learning prediction of mortality in patients in the intensive care unit: a retrospective study of high-frequency data in electronic patient records | |
| Peel et al. | Statistical inference links data and theory in network science | |
| Rizopoulos et al. | Combining dynamic predictions from joint models for longitudinal and time-to-event data using Bayesian model averaging | |
| Mavrogiorgou et al. | Analyzing data and data sources towards a unified approach for ensuring end-to-end data and data sources quality in healthcare 4.0 | |
| WO2020066724A1 (ja) | データ処理装置、データ処理方法およびプログラム | |
| Hsieh et al. | Rarefaction and extrapolation: making fair comparison of abundance-sensitive phylogenetic diversity among multiple assemblages | |
| Valsaraj et al. | Development and validation of echocardiography-based machine-learning models to predict mortality | |
| US12087444B2 (en) | Population-level gaussian processes for clinical time series forecasting | |
| US20210397951A1 (en) | Data processing apparatus, data processing method, and program | |
| US20250127493A1 (en) | Surfacing insights into left and right ventricular dysfunction through deep learning | |
| Kim et al. | The partial derivative framework for substantive regression effects. | |
| Schinkel et al. | Detecting changes in the performance of a clinical machine learning tool over time | |
| Rimella et al. | Inference on extended-spectrum beta-lactamase Escherichia coli and Klebsiella pneumoniae data through SMC2 | |
| Levy et al. | Development and validation of self-monitoring auto-updating prognostic models of survival for hospitalized COVID-19 patients | |
| Quevedo et al. | Online monitoring of nonlinear profiles using a Gaussian process model with heteroscedasticity | |
| Mayya et al. | Empirical study of feature selection methods in regression for large-scale healthcare data: a case study on estimating dental expenditures | |
| US20170147776A1 (en) | Continuous monitoring of event trajectories system and related method | |
| KR102536184B1 (ko) | 건강검진 결과지 생성 시스템 | |
| AU2020367789A1 (en) | Method for enhancing patient compliance with a medical therapy plan and mobile device therefor | |
| Carroll et al. | Temporally dependent accelerated failure time model for capturing the impact of events that alter survival in disease mapping | |
| Thorpe et al. | Sensing behaviour in healthcare design | |
| Wadhai | Adaptive real time data mining methodology for wireless body area network based healthcare applications | |
| Pirracchio et al. | Recalibrating our prediction models in the ICU: time to move from the abacus to the computer | |
| Ieva et al. | A semiparametric bivariate probit model for joint modeling of outcomes in STEMI patients | |
| Chakraborty et al. | A method to estimate intra-cluster correlation for clustered categorical data |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19864011 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19864011 Country of ref document: EP Kind code of ref document: A1 |