WO2020215671A1 - 数据智能分析方法、装置、计算机设备及存储介质 - Google Patents

数据智能分析方法、装置、计算机设备及存储介质 Download PDF

Info

Publication number
WO2020215671A1
WO2020215671A1 PCT/CN2019/116942 CN2019116942W WO2020215671A1 WO 2020215671 A1 WO2020215671 A1 WO 2020215671A1 CN 2019116942 W CN2019116942 W CN 2019116942W WO 2020215671 A1 WO2020215671 A1 WO 2020215671A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
sample data
processed
public opinion
target
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/116942
Other languages
English (en)
French (fr)
Inventor
陈娴娴
阮晓雯
徐亮
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Priority to JP2021506707A priority Critical patent/JP7165809B2/ja
Priority to SG11202008324YA priority patent/SG11202008324YA/en
Publication of WO2020215671A1 publication Critical patent/WO2020215671A1/zh
Priority to US17/168,925 priority patent/US20210158973A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H40/00ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices
    • G16H40/60ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices
    • G16H40/67ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices for remote operation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/21Design, administration or maintenance of databases
    • G06F16/215Improving data quality; Data cleansing, e.g. de-duplication, removing invalid entries or correcting typographical errors
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/24Querying
    • G06F16/245Query processing
    • G06F16/2458Special types of queries, e.g. statistical queries, fuzzy queries or distributed queries
    • G06F16/2465Query processing support for facilitating data mining operations in structured databases
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/951Indexing; Web crawling techniques
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques
    • G06F18/243Classification techniques relating to the number of classes
    • G06F18/24323Tree-organised classifiers
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/20Ensemble learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/01Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/80ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for detecting, monitoring or modelling epidemics or pandemics, e.g. flu
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02ATECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
    • Y02A90/00Technologies having an indirect contribution to adaptation to climate change
    • Y02A90/10Information and communication technologies [ICT] supporting adaptation to climate change, e.g. for weather forecasting or climate simulation

Definitions

  • This application relates to the field of data prediction technology, and in particular to a data intelligent analysis method, device, computer equipment and storage medium.
  • the embodiments of the present application provide a data intelligent analysis method, device, computer equipment, and storage medium to solve the current problem of low model prediction accuracy when data prediction is performed on lagging data.
  • An intelligent data analysis method including:
  • the hit entry corresponds to a public opinion factor
  • the public opinion index carries a time label
  • the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain a target prediction model; the improved multi-granularity cascaded random forest algorithm includes a pooling layer for retaining data features.
  • a data intelligent analysis device including:
  • the public opinion data acquisition module is used to crawl the public opinion data obtained by the third-party information platform using a crawler tool according to preset keywords.
  • the hit term determining module is used to determine at least one hit term based on public opinion data; the hit term corresponds to a public opinion factor.
  • the public opinion index acquisition module is used to acquire medical data within historical unit time and the public opinion index corresponding to the hit entry; the public opinion index carries a time label.
  • the first portrait data acquisition module is configured to use the public opinion factor and the public opinion index carrying the time tag as the first portrait data.
  • the original sample data obtaining module is used to obtain original sample data based on the first portrait data and the medical data.
  • the sample data acquisition module to be processed is used to perform data cleaning on the original sample data to obtain sample data to be processed;
  • the lagging sample data acquisition module is used to perform lagging processing on the sample data to be processed to obtain lagging sample data
  • the target sample data acquisition module is used to perform feature expansion processing on the lagging sample data to acquire target sample data
  • the target prediction model acquisition module is used to train the target sample data by an improved multi-granularity cascaded random forest algorithm to obtain a target prediction model;
  • the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pool
  • the transformation layer is used to retain data characteristics.
  • a computer device including a memory, a processor, and computer readable instructions stored in the memory and capable of running on the processor, and the processor implements the above data intelligent analysis method when the processor executes the computer readable instructions A step of.
  • a readable storage medium stores computer readable instructions, and when the computer readable instructions are executed by a processor, the steps of the above intelligent data analysis method are realized.
  • FIG. 1 is a schematic diagram of an application environment of a data intelligent analysis method in an embodiment of the present application
  • Figure 2 is a flowchart of a data intelligent analysis method in an embodiment of the present application
  • FIG. 3 is a specific flowchart of step S60 in FIG. 2;
  • FIG. 4 is a specific flowchart of step S80 in FIG. 2;
  • FIG. 5 is a flowchart of a data intelligent analysis method in an embodiment of the present application.
  • FIG. 6 is a specific flowchart of step S90 in FIG. 2;
  • FIG. 7 is a specific flowchart of step S92 in FIG. 6;
  • FIG. 8 is a schematic diagram of an intelligent data analysis device in an embodiment of the present application.
  • Fig. 9 is a schematic diagram of a computer device in an embodiment of the present application.
  • the data intelligent analysis method provided in the embodiments of this application can be applied to this method can be applied to a data intelligent analysis tool, which can train different samples according to the sample data corresponding to different topics (such as chickenpox, flu, etc.)
  • the prediction model especially for the sample data with lag, can effectively guarantee the accuracy of the model prediction.
  • the data intelligent analysis method can be applied in the application environment as shown in Fig. 1, in which the computer equipment communicates with the server through the network.
  • Computer equipment can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and portable wearable devices.
  • the server can be implemented as an independent server.
  • a method for intelligent data analysis is provided.
  • the method is applied to the server in FIG. 1 as an example for description, and includes the following steps:
  • S10 Use a crawler tool to crawl public opinion data obtained by a third-party information platform according to preset keywords.
  • the preset keywords are preset keywords related to communicable diseases, such as chickenpox, redness, pruritic herpes and blisters.
  • Public opinion data refers to the text data publicly released by different users in a third-party information platform to reflect the occurrence of social events. Specifically, with the rapid development of the information age at any time, users are more inclined to use various information platforms to query the required information, such as querying whether they have a disease based on their own symptoms, etc.
  • a crawler tool is used to crawl public opinion data containing preset keywords in a third-party information platform (such as Baidu, Weibo, or WeChat) .
  • a third-party information platform such as Baidu, Weibo, or WeChat.
  • S20 Determine at least one hit term based on the public opinion data, and the hit term corresponds to a public opinion factor.
  • the daily public opinion factors of 20 years in different regions are selected as another part of the portrait data.
  • the public opinion factors include, but are not limited to, chickenpox, redness, pruritic herpes and blisters.
  • the public opinion data includes at least one original entry (such as a Baidu entry). Specifically, the expert judges whether it is related to varicella according to the information contained in each crawled original entry, so as to determine at least one entry that is actually related to varicella as a hit entry. Then, according to the determined hit entry.
  • Each hit entry corresponds to a public opinion factor.
  • the public opinion factor refers to at least one factor related to a preset keyword contained in the hit entry, such as chickenpox, redness, pruritic herpes, and blisters.
  • S30 Obtain the medical data in historical unit time and the public opinion index corresponding to the hit entry.
  • the public opinion index carries a time label.
  • medical data refers to the historical unit time of sentinel hospitals in different regions provided by the CDC, such as the historical number of patients (ie, tag data) in a unit time of 20 years.
  • the unit time is the time label, and the unit time can be selected by the user and is not limited here.
  • the unit time may be one day, one week, one month, one quarter, or one year, etc., which will not be listed here.
  • the unit time is one week as an example.
  • the public opinion index and medical data corresponding to the hit entry within the unit time are obtained.
  • Each public opinion index carries a time label, which refers to the time label of the hit entry. release time.
  • S40 Use the public opinion factor and the public opinion index carrying the time label as the first portrait data.
  • the first portrait data refers to the public opinion factor and the public opinion index carrying the time label as the feature data for model training.
  • the time interval can be one week, one month, one quarter, or one year.
  • the processing of sample data will be different.
  • public opinion factors such as chickenpox, redness and herpes
  • the public opinion index of the Nth week can be used as row labels to establish partial portrait data.
  • the public opinion index of the Nth week includes, but is not limited to, the average public opinion index of the Nth week (that is, the average public opinion index of 7 days a week), the largest public opinion index of the Nth week, and the smallest public opinion index of the Nth week.
  • the first portrait data is used as the feature data for model training
  • the medical data is used as the label data for model training to obtain the original sample data.
  • S60 Perform data cleaning on the original sample data to obtain sample data to be processed.
  • the original sample data may include missing values or abnormal values
  • S70 Perform lag processing on the sample data to be processed to obtain lag sample data.
  • lag processing is a feature engineering method, by expanding the sample data set, that is, increasing the feature portrait to collect more information.
  • the corresponding sample data has lag, such as disease outbreaks or economic-related data.
  • the prediction theme is to predict chickenpox, and there is a hysteresis in the outbreak of chickenpox. For example, if the temperature suddenly rises this week and the climate is humid, it may not cause the outbreak of chickenpox this week, but it will usher in the next week. During the outbreak period, it is necessary to perform lag processing on the sample data to be processed to ensure the accuracy of subsequent model predictions.
  • the sample data to be processed is subjected to n lag processing (n generally takes 1 to 3). If n is 1, the sample data to be processed is subjected to lag processing, that is, the data of the original first week is regarded as the data of the second week. The data of the second week is used as the data of the third week, and so on, to get the lagging sample data.
  • n is set to 2
  • the sample data to be processed is subjected to lag processing, that is, the data of the original first week is regarded as the data of the third week, and the data of the second week is The data is used as the data of the fourth week, and so on, to obtain the lagging data, and integrate the lagging data obtained each time to obtain the lagging sample data to achieve the purpose of expanding the sample data set
  • the concat function is used to combine the lag sample data obtained by multiple lag processing with the sample data to be processed into a data frame (DataFrame), that is, the lag sample data.
  • DataFrame data frame
  • the concat function is a function used to connect two or more arrays.
  • the data frame is a two-dimensional data structure, that is, the data is arranged in a table of rows and columns.
  • S80 Perform feature expansion processing on the lagging sample data to obtain target sample data.
  • feature expansion processing is performed on the lagging sample data to obtain target sample data, so as to achieve the purpose of further expanding the sample data set.
  • the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain a target prediction model.
  • the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pooling layer is used to retain data features.
  • the improved multi-granularity cascaded random forest algorithm is an algorithm that introduces the idea of pooling in the convolutional neural network into the multi-granularity cascaded random forest algorithm.
  • the multi-granularity cascaded random forest algorithm is a decision tree integration method, which stacks multiple layers of random forests in a cascaded manner to obtain better feature representation and learning performance. The algorithm does not need to adjust the hyperparameters to achieve good results Performance.
  • each layer in the multi-granularity cascaded random forest is composed of multiple random forests.
  • the feature information of the input feature vector is learned through random forest, and then input to the next layer after processing.
  • a variety of different types of random forests are selected for each layer, for example, two random forest structures are selected for each layer, namely completely-random tree forests and random forests. .
  • a crawler tool is used to crawl public opinion data obtained by a third-party information platform, so as to determine at least one hit entry that is truly related to the predicted topic based on the public opinion data to ensure subsequent acquisitions The validity and accuracy of the public opinion factor. Then obtain the public opinion index and medical data corresponding to the hit entry per unit time. Finally, the public opinion factor and the public opinion index carrying the time label are used as the original sample data, so that the model analyzes the public opinion data in a unit time of 20 years in history. Then, by performing data cleaning on the original sample data, the sample data to be processed is obtained to ensure the quality of the sample data to be processed.
  • the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain the target prediction model to obtain better feature representation and learning performance, and the algorithm can achieve good performance without excessive adjustment of hyperparameters. Ensure the accuracy of model predictions.
  • the improved multi-granularity cascaded random forest algorithm also includes a pooling layer to fully retain data features and further improve the accuracy of model prediction.
  • the data intelligent analysis method before step S10, further includes:
  • this embodiment can select different portrait data according to different predicted themes.
  • the prediction of chickenpox is taken as an example for illustration. Since the climate conditions and the chickenpox virus are very closely related, the history of different regions is selected.
  • the daily meteorological factors of the year are used as part of the profile data.
  • the meteorological factors include, but are not limited to, day and night temperature, day and night pressure, day and night precipitation, humidity, light intensity, and wind in different regions.
  • the second portrait data refers to the feature data trained with meteorological factors and corresponding meteorological data as the model.
  • the method of establishing portrait data for meteorological factors is consistent with step S40, that is, meteorological factors can be used as column labels, and the meteorological conditions of the Nth week can be used as row labels to establish the second portrait data.
  • the meteorological conditions of the Nth week include but are not limited to the average meteorological conditions of the Nth week (such as average precipitation), the maximum meteorological conditions of the Nth week (such as maximum precipitation), and the minimum meteorological conditions of the Nth week (such as minimum Precipitation).
  • step S50 that is, based on the first portrait data and medical data, obtaining the original sample data includes:
  • S51 Use the first portrait data, the second portrait data and the medical data as original sample data.
  • the meteorological situation is combined with the idea of mass dissemination of public opinion data to effectively predict the period of disease outbreak and improve the accuracy of model prediction.
  • step S60 data cleaning is performed on the original sample data to obtain sample data to be processed, which specifically includes the following steps:
  • S61 Perform missing value filling on the original sample data to obtain the first sample data.
  • missing value filling methods include, but are not limited to, mean filling, mode filling, median filling, expectation maximization methods, multiple filling, and k-means clustering methods.
  • the portrait data where the missing value is located is clustered, and the missing value is filled with the mean value of the clustered cluster.
  • S62 Perform abnormal value detection on the first sample data to obtain at least one abnormal value, and mark the abnormal value as empty.
  • S63 Fill the outliers marked as empty with missing values to obtain sample data to be processed.
  • outlier detection includes, but is not limited to, the use of statistical variable analysis (such as box-plot analysis, average, maximum and minimum analysis, and the 3 ⁇ rule), distance-based methods, density-based outlier detection, and density-based outlier detection.
  • Group point detection and isolation forest Isolation Forest
  • an outlier is defined as a value whose deviation from the average value in a set of measured values exceeds 3 times the standard deviation, because in the normal distribution Under the assumption of, the probability of occurrence of values beyond 3 ⁇ from the average value is less than 0.003), that is, data exceeding ⁇ +3 ⁇ and data not exceeding ⁇ -3 ⁇ are regarded as abnormal values.
  • the outliers are deleted and marked as null values, and then the outliers marked as null values are filled with missing values again to obtain the sample data to be processed.
  • the sample data to be processed is obtained by filling the outliers marked as null with missing values, so as to avoid directly removing the sample data corresponding to the outliers, causing the sample data to lack this part of the feature and affecting the accuracy of model prediction The problem of rate.
  • the original sample data is filled with missing values to obtain the first sample data, and then abnormal value detection is performed on the first sample data to obtain at least one abnormal value. And the missing values are processed to achieve the purpose of data cleaning and ensure the quality of sample data. Then, the obtained outliers are marked as empty, so that the outliers marked as empty are filled with missing values again to obtain the sample to be processed, and the original sample data is filled with missing values twice to ensure the quality and standard of the sample data To improve the accuracy of model predictions
  • step S80 that is, performing feature expansion processing on lagging sample data to obtain target sample data, specifically includes the following steps:
  • S81 Perform feature expansion on the lagging sample data to obtain a feature value corresponding to at least one statistical indicator.
  • S82 Splicing the eigenvalues with the lagging sample data to obtain target sample data.
  • the statistical indicators include but are not limited to the maximum, minimum, mean and standard deviation corresponding to each row of data.
  • Each statistical indicator is added as a new column to the lagging sample data to expand the data set and increase the collection of feature images. Multiple feature information improves the accuracy of model prediction.
  • the lagging sample data is a matrix. The eigenvalues and the lagging sample data are spliced together to obtain the target sample data, that is, N columns are added to the sample matrix. Value, minimum and average), the maximum, minimum, and average of the data corresponding to each row are the characteristic values.
  • the feature value corresponding to at least one statistical indicator is obtained by feature expansion of the lagging sample data, and the feature value is spliced with the lagging sample data to obtain target sample data to expand the data set and increase the collection of feature images.
  • Feature information improves the accuracy of model prediction.
  • the data intelligent analysis method further includes the following steps:
  • S111 Perform variance analysis on the target sample data, and remove data whose variance is less than a preset variance threshold to obtain second sample data.
  • S112 Perform singular value decomposition on the second sample data to update the target sample data.
  • the analysis of variance refers to the analysis based on the variance of the data column to remove the sequence with too small variance (that is, less than the preset variance threshold) to obtain the second sample data.
  • the size of the variance describes the amount of information of a variable, and the sequence with too small variance is considered to contain less information, so all data columns with small variance are removed to achieve the effect of data dimensionality reduction and reduce the amount of data processing , Improve the efficiency of subsequent model training.
  • the singular value decomposition of the second sample data is also required to remove redundant data, to achieve the purpose of data compression, and to ensure the quality of the target sample data.
  • the improved multi-granularity cascaded random forest algorithm includes a multi-particle scanning algorithm and a cascaded random forest algorithm, and the multi-particle scanning algorithm corresponds to at least one sliding window, as shown in FIG. 6, in step S90, specifically including the following steps :
  • S91 Use a multi-particle scanning algorithm to perform multi-particle scanning on the target sample data according to at least one sliding window to obtain at least one intermediate data.
  • multi-particle scanning refers to scanning the target sample data using a sliding window to obtain at least one intermediate data.
  • sliding windows of different dimensions can be set.
  • the sliding window can be an i*j window.
  • the target table sample data row label is the i-th week
  • the sliding window window_size can be 2 (every 2 weeks), 4 (every month), 12 (every quarter), etc.
  • the sliding window can scan at least one feature portrait, that is, every row, every two rows, and every j row can be scanned to maximize the search for the internal relationship between the feature and the tag set, and the feature and the feature.
  • S92 Based on the pooling layer, perform pooling processing on at least one intermediate data to obtain data to be trained.
  • At least one intermediate data is pooled through the pooling layer to obtain the data to be trained, so as to achieve the purpose of reducing the dimensionality of the data, reduce the amount of calculation, and improve the efficiency of model training.
  • multi-granularity cascaded neural network ensemble Random Forest algorithm based on the idea, the i-th complete-random tree forest predicted label cforest i and column labels random forest column rforest i predicted as a target sample data continuously added
  • the portrait column is expanded with further features, and finally the following feature portrait is obtained [orgf 1 ,orgf 2 ,K,orgf n ,cforest 1 ,rforest 1 ,K,cforest k ,rforest k ].
  • orgf is the target sample data.
  • the obtained data to be trained into the cascade forest for training.
  • sliding windows of three dimensions are used.
  • the sliding window of the first dimension is used to scan to obtain a feature vector, and then the original feature vector is input into the complete-random tree forest and random forest to obtain two predicted sequence (i.e. cforest i and rforest i), then the two predicted sequence spliced to give a first feature vector, the feature vector input into the original first hierarchical linking forest training, to obtain a first predicted sequence.
  • the obtained first prediction sequence and the first feature vector are spliced together to obtain the second feature vector, which is used as the input data of the second layer of cascaded forest;
  • the third feature vector obtained by the sliding window of the dimension (same method as the first feature vector) is spliced as the input data of the third-level cascaded forest;
  • the third prediction sequence obtained by the third-level cascade forest training is then combined with the third
  • the fourth feature vector obtained by the sliding window of the dimension is spliced and used as the input of the next layer, and the above process is repeated until convergence, and the target prediction model is obtained.
  • multi-particle scanning is performed on the target sample data according to at least one sliding window to obtain at least one intermediate data to maximize the search for features and tag sets, and between features and features. Internal relevance. Then, by combining the pooling layer, at least one intermediate data is pooled to obtain the data to be trained, so as to combine machine learning and neural network ideas to obtain more intuitive and unobtainable information to enrich the model and further improve the model Forecast accuracy rate.
  • step S92 that is, based on the pooling layer, pooling is performed on at least one intermediate data to obtain the data to be trained, which specifically includes the following steps:
  • S921 Select two adjacent intermediate data as a group of to-be-processed data groups to obtain at least one group of to-be-processed data groups corresponding to the intermediate data.
  • S922 Perform an averaging operation on each data group to be processed to obtain a first data sequence.
  • S923 Perform a minimum value operation on each group of to-be-processed data groups to obtain a second data sequence.
  • the second data column includes the minimum value of the two intermediate data of each group of to-be-processed data groups.
  • S924 Perform a maximum value operation on each group of to-be-processed data groups to obtain a third data sequence.
  • the third data column includes the maximum value of the two intermediate data of each group of to-be-processed data groups.
  • S925 Join the first data sequence, the second data sequence, and the third data sequence to obtain data to be trained.
  • model prediction requires more linear or nonlinear methods to spatially warp data, so as to obtain more information that is not intuitively available to enrich the model. Therefore, in this embodiment, The three pooling methods pool at least one intermediate data, and then integrate the results obtained by pooling in each method to obtain the data to be trained to obtain more intuitive and unobtainable information to enrich the model, and Fully retain data characteristics. Assuming that one column of portrait data in the intermediate data is Feature: f 1 , f 2 , f 3 , f 4 , f 5 , K f n , the following three pooling methods are used to pool at least one intermediate data.
  • Feature_new_1 (f 1 +f 2 )/2,(f 2 +f 3 )/2,K,(f n-1 +f n )/2
  • Feature_new_2 max(f 1 ,f 2 ),max(f 2 ,f 3 ),K,max(f n-1 ,f n )
  • Feature_new_3 min(f 1 ,f 2 ),min(f 2 ,f 3 ),K,min(f n-1 ,f n )
  • At least one intermediate data is pooled by using three pooling methods, and then the results obtained by pooling in each method are integrated to obtain the data to be trained, so as to fully retain the data characteristics and ensure the sample data Quality, improve the accuracy of model prediction.
  • a data intelligent analysis device corresponds one-to-one with the data intelligent analysis method in the foregoing embodiment.
  • the data intelligent analysis device includes a public opinion data acquisition module 10, a hit entry determination module 20, a public opinion index acquisition module 30, a first portrait data acquisition module 40, an original sample data acquisition module 50, and sample data to be processed
  • the acquisition module 60, the lagging sample data acquisition module 70, the target sample data acquisition module 80, and the target prediction model acquisition module 90 is as follows:
  • the public opinion data acquisition module 10 is used for crawling public opinion data obtained by a third-party information platform using a crawler tool according to preset keywords.
  • the hit term determining module 20 is configured to determine at least one hit term based on public opinion data; the hit term corresponds to a public opinion factor.
  • the public opinion index acquisition module 30 is used to acquire the medical data in historical unit time and the public opinion index corresponding to the hit entry; the public opinion index carries a time label.
  • the first portrait data acquisition module 40 is configured to use the public opinion factor and the public opinion index carrying the time tag as the first portrait data.
  • the original sample data obtaining module 50 is used to obtain original sample data based on the first portrait data and medical data.
  • the sample data acquisition module 60 to be processed is used to perform data cleaning on the original sample data to obtain sample data to be processed.
  • the lagging sample data acquisition module 70 is configured to perform lag processing on the sample data to be processed to obtain lagging sample data.
  • the target sample data acquisition module 80 is configured to perform feature expansion processing on the lagging sample data to acquire target sample data.
  • the target prediction model acquisition module 90 is used to train the target sample data with the improved multi-granularity cascaded random forest algorithm to obtain the target prediction model;
  • the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pooling layer is used for retention Data characteristics.
  • the sample data acquisition module to be processed includes a first sample data acquisition unit, an abnormal value acquisition unit, and a sample data acquisition unit to be processed.
  • the first sample data acquisition unit is used to fill in missing values on the original sample data to obtain the first sample data.
  • the abnormal value obtaining unit is used to perform abnormal value detection on the first sample data to obtain at least one abnormal value, and mark the abnormal value as empty.
  • the to-be-processed sample data acquisition unit is used to fill in the missing value of the outliers marked as empty to obtain the to-be-processed sample data.
  • the target sample data acquisition module includes a feature value acquisition unit and a target sample data acquisition unit.
  • the feature value obtaining unit is used to perform feature expansion on the lagging sample data to obtain a feature value corresponding to at least one statistical indicator.
  • the target sample data acquisition unit is used to splice the feature value and the lagging sample data to acquire the target sample data.
  • the data intelligent analysis device includes a second sample data acquisition unit and a target sample data update unit.
  • the second sample data acquisition unit is configured to perform variance analysis on the target sample data, remove data whose variance is less than a preset variance threshold, and obtain second sample data.
  • the target sample data update unit is used to perform singular value decomposition on the second sample data to update the target sample data.
  • the improved multi-granularity cascaded random forest algorithm includes a multi-particle scanning algorithm and a cascaded random forest algorithm.
  • the multi-particle scanning algorithm corresponds to at least one sliding window;
  • the target prediction model acquisition module includes a target prediction model, a data acquisition unit to be trained, and a target Predictive model acquisition unit.
  • the intermediate data acquisition unit is configured to use a multi-particle scanning algorithm to perform multi-particle scanning on the target sample data according to at least one sliding window to obtain at least one intermediate data.
  • the to-be-trained data acquisition unit is configured to perform pooling processing on at least one intermediate data based on the pooling layer to obtain the to-be-trained data.
  • the target prediction model acquisition unit is used to train the training data by using the cascaded random forest algorithm to obtain the target prediction model.
  • the data acquisition unit to be trained includes a data group acquisition subunit to be processed, a first data sequence acquisition subunit, a second data sequence acquisition subunit, a third data sequence acquisition subunit, and a data sequence acquisition subunit to be trained.
  • the to-be-processed data group acquiring subunit is used to select two adjacent intermediate data as a group of to-be-processed data groups to obtain at least one group of to-be-processed data groups corresponding to the intermediate data.
  • the first data sequence obtaining subunit is used to perform an averaging operation on each group of to-be-processed data groups to obtain the first data sequence.
  • the second data sequence obtaining subunit is used to perform minimum value operation on each group of to-be-processed data groups to obtain a second data sequence.
  • the second data column includes the minimum value of the two intermediate data of each group of to-be-processed data groups.
  • the third data sequence obtaining subunit is used to perform a maximum value operation on each group of to-be-processed data groups to obtain a third data sequence.
  • the third data column includes the maximum value of the two intermediate data of each group of to-be-processed data groups.
  • the to-be-trained data acquisition subunit is used to splice the first data sequence, the second data sequence and the third data sequence to obtain the to-be-trained data.
  • Each module in the above-mentioned data intelligent analysis device can be implemented in whole or in part by software, hardware and a combination thereof.
  • the foregoing modules may be embedded in the form of hardware or independent of the processor in the computer device, or may be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the foregoing modules.
  • a computer device is provided.
  • the computer device may be a server, and its internal structure diagram may be as shown in FIG. 10.
  • the computer equipment includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide calculation and control capabilities.
  • the memory of the computer device includes a readable storage medium and an internal memory.
  • the readable storage medium stores an operating system, computer readable instructions, and a database.
  • the internal memory provides an environment for the operation of the operating system and computer readable instructions in the readable storage medium.
  • the database of the computer equipment is used to store data generated or obtained during the execution of the intelligent data analysis method, such as target sample data.
  • the network interface of the computer device is used to communicate with an external terminal through a network connection.
  • the computer readable instructions are executed by the processor to realize a data intelligent analysis method.
  • a computer device including a memory, a processor, and computer-readable instructions stored in the memory and capable of running on the processor.
  • the steps of the data intelligent analysis method are, for example, steps S10-S90 shown in FIG. 2 or the steps shown in FIG. 3 to FIG. 7.
  • the processor implements the functions of the modules/units in this embodiment of the data intelligent analysis device when the processor executes the computer-readable instructions, such as the functions of the modules/units shown in FIG. 8. To avoid repetition, details are not described herein again.
  • one or more readable storage media storing computer readable instructions are provided.
  • the computer readable storage medium stores computer readable instructions, wherein the computer readable instructions are controlled by one or When multiple processors are executed, the one or more processors are executed to implement the steps of the data intelligent analysis method in the foregoing embodiment, for example, steps S10-S90 shown in FIG. 2 or shown in FIGS. 3 to 7 To avoid repetition, I won’t repeat them here.
  • the computer-readable instruction is executed by the processor, the function of each module/unit in the embodiment of the above-mentioned data intelligent analysis device is realized, for example, the function of each module/unit shown in FIG. Repeat it again.
  • the readable storage medium in this embodiment includes a nonvolatile readable storage medium and a volatile readable storage medium.
  • Non-volatile memory may include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • ROM read only memory
  • PROM programmable ROM
  • EPROM electrically programmable ROM
  • EEPROM electrically erasable programmable ROM
  • Volatile memory may include random access memory (RAM) or external cache memory.
  • RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain Channel (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Theoretical Computer Science (AREA)
  • Medical Informatics (AREA)
  • Public Health (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • Biomedical Technology (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Primary Health Care (AREA)
  • Epidemiology (AREA)
  • Pathology (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computational Linguistics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Evolutionary Biology (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Business, Economics & Management (AREA)
  • General Business, Economics & Management (AREA)
  • Fuzzy Systems (AREA)
  • Probability & Statistics with Applications (AREA)
  • Quality & Reliability (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

一种数据智能分析方法、装置、计算机设备及存储介质,该数据智能分析方法包括:将获取到的舆情因子和携带时间标签的舆情指数作为第一画像数据(S40);基于所述第一画像数据和医疗数据,获取原始样本数据(S50);对所述原始样本数据进行数据清洗,得到待处理样本数据(S60);对所述待处理样本数据进行滞后处理,得到滞后样本数据(S70);对所述滞后样本数据进行特征扩充处理,获取目标样本数据(S80);采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征(S90),该数据智能分析方法可有效提高模型预测准确率和模型训练效率。

Description

数据智能分析方法、装置、计算机设备及存储介质
本申请以2019年8月19日提交的申请号为201910763137.5,名称为“数据智能分析方法、装置、计算机设备及存储介质”的中国发明专利申请为基础,并要求其优先权。
技术领域
本申请涉及数据预测技术领域,尤其涉及一种数据智能分析方法、装置、计算机设备及存储介质。
背景技术
随着信息时代的飞速发展,数据预测技术也在不断发展。目前各大科研机构针对医疗数据进行预测时,由于部分医疗数据具有滞后性,导致模型预测准确率较低,例如对于具有一定潜伏期的传染性疾病(如水痘),在满足疾病爆发的条件(如气温、湿度等)时,可能会在下一时段爆发,这就导致模型预测准确率较低,使公民不能及时预防,无法对疾病爆发的严重程度起到控制作用。
发明内容
本申请实施例提供一种数据智能分析方法、装置、计算机设备及存储介质,以解决目前对滞后性的数据进行数据预测时,模型预测准确率较低的问题。
一种数据智能分析方法,包括:
按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;
基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;
获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;
将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;
基于所述第一画像数据和所述医疗数据,获取原始样本数据;
对所述原始样本数据进行数据清洗,得到待处理样本数据;
对所述待处理样本数据进行滞后处理,得到滞后样本数据;
对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
一种数据智能分析装置,包括:
舆情数据获取模块,用于按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据。
命中词条确定模块,用于基于舆情数据,确定至少一个命中词条;命中词条对应一舆情因子。
舆情指数获取模块,用于获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签。
第一画像数据获取模块,用于将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据。
原始样本数据获取模块,用于基于所述第一画像数据和所述医疗数据,获取原始样本数据。
待处理样本数据获取模块,用于对所述原始样本数据进行数据清洗,得到待处理样本数据;
滞后样本数据获取模块,用于对所述待处理样本数据进行滞后处理,得到滞后样本数据;
目标样本数据获取模块,用于对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
目标预测模型获取模块,用于采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
一种计算机设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现上述数据智能分析方法的步骤。
一种可读存储介质,所述可读存储介质存储有计算机可读指令,所述计算机可读指令被处理器执行时实现上述数据智能分析方法的步骤。
本申请的一个或多个实施例的细节在下面的附图和描述中提出,本申请的其他特征和优点将从说明书、附图以及权利要求变得明显。
附图说明
为了更清楚地说明本申请实施例的技术方案,下面将对本申请实施例的描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1是本申请一实施例中数据智能分析方法的一应用环境示意图;
图2是本申请一实施例中数据智能分析方法的一流程图;
图3是图2中步骤S60的一具体流程图;
图4是图2中步骤S80的一具体流程图;
图5是本申请一实施例中数据智能分析方法的一流程图;
图6是图2中步骤S90的一具体流程图;
图7是图6中步骤S92的一具体流程图;
图8是本申请一实施例中数据智能分析装置的一示意图;
图9是本申请一实施例中计算机设备的一示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本申请实施例提供的数据智能分析方法可应用在本方法可应用在一种数据智能分析工具中,该数据智能分析工具可根据不同的主题(例如水痘、流感等)对应的样本数据训练不同的预测模型,尤其针对具有滞后性的样本数据,可有效保证模型预测的准确率。该数据智能分析方法可应用在如图1的应用环境中,其中,计算机设备通过网络与服务器进行通信。计算机设备可以但不限于各种个人计算机、笔记本电脑、智能手机、平板电脑和便携式可穿戴设备。服务器可以用独立的服务器来实现。
在一实施例中,如图2所示,提供一种数据智能分析方法,以该方法应用在图1中的服务器为例进行说明,包括如下步骤:
S10:按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据。
其中,预设关键词是预先设置的涉及传播性疾病的一些关键词,如水痘、红肿、瘙痒性疱疹和水疱疹等。舆情数据是指第三方信息平台中不同用户所公开发布的文字数据,用于反映社会事件的发生。具体地,随时信息时代的飞速发展,用户更倾向于采用各种信息平台查询所需的信息,例如根据自身症状查询是否患有疾病等,当某一传播性疾病爆发(如水痘)时,必然会有更大的搜索量或关注度,故本实施例中还按照预设关键词,采用爬虫工具爬取第三方信息平台(如百度、微博或微信)中包含预设关键词的舆情数据。需要说明的是,本实施例中的涉及传播性疾病的一些预设关键词可预先设定部分默认的关键词,再取该默认的关键词对应的近义词,以得到更多的关键词进行爬取,获取更多相关的信息,为后续模型训练提供充足的数据集。
S20:基于舆情数据,确定至少一个命中词条,命中词条对应一舆情因子。
具体地,随时信息时代的飞速发展,用户更倾向于采用各种信息平台查询所需的信息,例如根据自身症状查询是否患有疾病等,当某一传播性疾病爆发(如水痘)时,必然会有更大的搜索量或关注度,故本实施例中选择不同地区历史20年每天的舆情因子作为另一部分画像数据。该舆情因子包括但不限于水痘、红肿、瘙痒性疱疹和水疱疹等。
其中,舆情数据包括至少一个原始词条(如百度词条)。具体地,通过专家根据爬取到的每一原始词条中所包含的信息判断是否与水痘有关,以确定至少一个与水痘真实相关的词条作为命中词条。然后,再根据确定的命中词条。每一命中词条对应一舆情因子。该舆情因子是指命中词条中包含的至少一个与预设关键词相关的因子,如水痘、红肿、瘙痒性疱疹和水疱疹。
S30:获取历史单位时间内的医疗数据和命中词条对应的舆情指数舆情指数携带时间标签。
其中,医疗数据是指疾控中心提供的,不同地区哨点医院历史单位时间,如历史20年的单位时间内的历史发病人数(即标签数据)。可以理解地,该单位时间即为时间标签,该单位时间可由用户自定义选定,此处不做限定。本实施例中,该单位时间可为一天、一周、一个月、一个季度或者一年等等,在此不一一列举。
本实施例中,以单位时间为一周为例进行说明,具体地,获取单位时间内命中词条对应的舆情指数和医疗数据,每一舆情指数携带时间标签,该时间标签即指命中词条的发布时间。
S40:将舆情因子和携带时间标签的舆情指数作为第一画像数据。
其中,第一画像数据即指将舆情因子和携带时间标签的舆情指数作为模型训练的特征数据。具体地,当需要预测未来一个时间区间内某疾病是否爆发,该时间区间可为一周、一个月、一个季度或者一年,根据预测的时间区间不同,在样本数据的处理上会有所不同,以时间区间为一周进行举例说明,可以舆情因子(如水痘、红肿和疱疹)为列标签,以第N周的舆情指数为行标签,建立部分画像数据。其中,第N周舆情指数包括但不限于第N周平均舆情指数(即一周7天的舆情指数取平均)、第N周最大舆情指数以及第N周最小舆情指数。
需要说明的是,如下表格为本实施例中根据舆情因子建立的画像数据示意图。可以理解地,该示意图仅做示例,在此不做限定。
Figure PCTCN2019116942-appb-000001
Figure PCTCN2019116942-appb-000002
S50:基于第一画像数据和医疗数据,获取原始样本数据
具体地,将第一画像数据作为模型训练的特征数据,将医疗数据作为模型训练的标签数据,以获取原始样本数据。
S60:对原始样本数据进行数据清洗,得到待处理样本数据。
具体地,由于原始样本数据中可能包括缺失值或异常值,为进一步的保证后续模型预测的准确率,需要对原始样本数据进行数据清洗,以保证待处理样本数据的质量。
S70:对待处理样本数据进行滞后处理,得到滞后样本数据。
其中,滞后处理是一种特征工程方法,通过扩充样本数据集,即增大特征画像以收集更多信息的方法。从业务逻辑层面理解就是,延迟特征的效果。具体地,由于部分模型预测的主题不同,其对应的样本数据存在滞后性,如疾病的爆发或者与经济相关的数据。本实施例中,假设预测主题为预测水痘,而水痘的爆发存在滞后性,例如本周的气温突然升高且气候潮湿,可能这一周并不会带来水痘的爆发,但是下一周会迎来爆发期,故需要对待处理样本数据进行滞后处理,以保证后续模型预测的准确率。具体地,对待处理样本数据进行n次滞后处理(n一般取1~3),假设n取1,则对待处理样本数据进行滞后处理,即将原第一周的数据作为第二周的数据,第二周的数据作为第三周的数据,以此类推,以得到滞后样本数据。若n取2,由于是在第一次时候得到的样本数据的基础上进行再次之后,故对待处理样本数据进行滞后处理,即将原第一周的数据作为第三周的数据,第二周的数据作为第四周的数据,以此类推,得到滞后数据,将每次得到的滞后数据集成,以得到滞后样本数据,实现扩充样本数据集的目的
最后,采用concat函数将多次滞后处理得到的滞后样本数据与待处理样本数据合并为一个数据帧(DataFrame)即滞后样本数据。其中,concat函数是用于连接两个或多个数组的函数。数据帧是二维数据结构,即数据以行和列的表格方式排列。
S80:对滞后样本数据进行特征扩充处理,获取目标样本数据。
具体地,为了扩充样本数据集,进一步提高模型预测的准确率,本实施例中会对滞后样本数据进行特征扩充处理,得到目标样本数据,以达到进一步扩充样本数据集的目的。
S90:采用改进多粒度级联随机森林算法对目标样本数据进行训练,得到目标预测模型,改进多粒度级联随机森林算法包括一池化层,池化层用于保留数据特征。
其中,改进多粒度级联随机森林算法是在多粒度级联随机森林算法中引入卷积神经网络中池化思想的算法。多粒度级联随机森林算法是一种决策树集成方法,通过级联的方式堆叠多层随机森林,以获得更好的特征表示和学习性能,该算法无需过多调节超参数,即可达到良好的性能。
其中,多粒度级联随机森林(Gcforest)中每一层都由多个随机森林组成。通过随机森林学习输入特征向量的特征信息,经过处理后输入到下一层。为了增强模型的泛化能力,每一层选取多种不同类型的随机森林,例如每一层选取两种随机森林结构,分别为completely-random tree forests(完全随机森林)和random forests(随机森林)。
本实施例中,先按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据,以便基于舆情数据,确定至少一个与预测主题真实相关的命中词条,以保证后续获取的舆情因子的有效性和准确性。再获取单位时间内命中词条对应的舆情指数和医疗数据。最后,将舆情因子和携带时间标签的舆情指数作为原始样本数据,以使模型通过对历史20年,单位时间内的舆情数据进行分析。然后,通过对原始样本数据进行数据清洗,得到待处理样本数据,以保证待处理样本数据的质量。然后,对对待处理样本数据进行滞后处理, 得到滞后样本数据,以扩充样本数据集。并且,针对具有滞后性数据,可实现延迟特征的效果,保证模型预测的准确率。接着,对滞后样本数据进行特征扩充处理,获取目标样本数据,以达到进一步扩充样本数据集的目的,提高模型预测的准确率。最后,采用改进多粒度级联随机森林算法对目标样本数据进行训练,得到目标预测模型,以获得更好的特征表示和学习性能,且算法无需过多调节超参数,即可达到良好的性能,保证模型预测的准确率。并且,改进多粒度级联随机森林算法还包括一池化层,以充分保留数据特征,进一步提高模型预测的准确率。
在一实施例中,步骤S10之前中,该数据智能分析方法还包括:
S101:获取气象因子和对应的气象数据。
可以理解地,本实施例可根据预测主题的不同选择不同的画像数据,本实施例中以预测水痘为例进行说明,由于气候情况和水痘病毒存在非常紧密的相关性,故选择不同地区历史20年每天的气象因子作为一部分画像数据。该气象因子包括但不限于不同地区的昼夜气温、昼夜气压、昼夜降水量、湿度、光照强度和风力等。
S102:将气象因子和对应的气象数据作为第二画像数据;
其中,第二画像数据即指将气象因子和对应的气象数据作为模型训练的特征数据。具体地,针对气象因子建立画像数据的方式与步骤S40一致,即可以气象因子为列标签,以第N周的气象情况为行标签,建立第二画像数据。其中,第N周的气象情况包括但不限于第N周的平均气象情况(如平均降水量)、第N周的最大气象情况(如最大降水量)以及第N周的最小气象情况(如最小降水量)。
相应地,步骤S50中,即基于第一画像数据和医疗数据,获取原始样本数据,包括:
S51:将第一画像数据、第二画像数据和医疗数据作为原始样本数据。
本实施例中,通过气象情况结合舆情数据大量传播的思想,以有效预测疾病爆发时段,提高模型预测的准确率。
在一实施例中,如图3所示,步骤S60中,即对原始样本数据进行数据清洗,得到待处理样本数据,具体包括如下步骤:
S61:对原始样本数据进行缺失值填充,得到第一样本数据。
其中,缺失值填充方法包括但不限于均值填充、众数填充、中位数填充、期望值最大化方法、多重填补以及k-means聚类方法等。具体地,以k-means聚类方法进行填充为例,将缺失值所在的画像数据进行聚类,并将缺失值以所聚类类簇的均值进行填充。
S62:对第一样本数据进行异常值检测,得到至少一个异常值,将异常值标记为空。
S63:对标记为空的异常值进行缺失值填充,得到待处理样本数据。
具体地,异常值检测包括但不限于采用统计变量分析(如箱型图分析、平均值、最大最小值分析以及3σ法则)、基于距离的方法、基于密度的离群点检测、基于密度的离群点检测和孤立森林(Isolation Forest)等。本实施例中,以3σ法则为例,若数据服从正太分布,在3σ原则下,异常值被定义为一组测定值中与平均值的偏差超3倍标准差的值,因为在正态分布的假设下,距离平均值3σ之外的值出现的概率小于0.003),即超过μ+3σ的数据以及不超过μ-3σ的数据作为异常值。
具体地,由于异常值对应的样本数据不一定是不必要的,若直接将该异常值对应的样本数据删除,会导致样本数据中的特征缺失,影响样本数据的质量,进而影响模型预测的准确率,故本实施例中会将异常值删除并标记为空值,再对标记为空值的异常值再次进行缺失值填充,得到待处理样本数据。本实施例中,通过对标记为空值的异常值进行缺失值填充,得到待处理样本数据,以避免直接将异常值对应的样本数据去除,导致样本数据缺少该部分特征,影响模型预测的准确率的问题。
本实施例中,通过对原始样本数据进行缺失值填充,得到第一样本数据,再对对第一 样本数据进行异常值检测,得到至少一个异常值,以通过对样本数据中的异常值和缺失值进行处理,达到数据清洗的目的,保证样本数据的质量。然后,将得到的异常值标记为空,以便对标记为空的异常值再次进行缺失值填充,得到待处理样本,以通过对原始样本数据进行两次缺失值填充,保证样本数据的质量和规范性,提升模型预测的准确率
在一实施例中,如图4所示,步骤S80中,即对滞后样本数据进行特征扩充处理,获取目标样本数据,具体包括如下步骤:
S81:对滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值。
S82:将特征值与滞后样本数据进行拼接,获取目标样本数据。
其中,统计指标包括但不限于每一行数据对应的最大值、最小值、均值和标准差,将每一统计指标作为新的列加入滞后样本数据中,以扩充数据集并增大特征画像收集更多特征信息,提高模型预测的准确率。可以理解地,该滞后样本数据为一矩阵,将特征值与滞后样本数据进行拼接,获取目标样本数据,即在样本矩阵中增加N个列,N为统计指标(如每一行对应的数据的最大值、最小值和均值)的个数,每一行对应的数据的最大值、最小值和均值即为特征值。
本实施例中,通过对滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值,将特征值与滞后样本数据进行拼接,获取目标样本数据以扩充数据集并增大特征画像收集更多特征信息,提高模型预测的准确率。
在一实施例中,如图5所示,步骤S80之后,该数据智能分析方法还包括如下步骤:
S111:对目标样本数据进行方差分析,去除方差小于预设方差阈值的数据,得到第二样本数据。
S112:对第二样本数据进行奇异值分解,以更新目标样本数据。
具体地,由于数据量有时过犹不及,在数据分析应用中大量的数据反而会产生更坏的性能。故需要对目标样本数据进行筛选,以去除冗余数据,达到减少数据列数的同时保证丢失的数据信息尽可能少。
其中,方差分析是指根据数据列的方差进行分析,以去除方差过于小(即小于预设方差阈值)的序列,得到第二样本数据。具体地,方差的大小描述的是一个变量的信息量,方差过于小的序列则认为其包含的信息量少,故去除所有方差小的数据列,以达到数据降维的效果,降低数据处理量,提高后续模型训练效率。
具体地,在目标样本数据中包含多个特征,但某些特征对于模型的预测精度的影响并不大,或者可认为相关性过大的特征可同等替换,故可将冗余变量去除,以达到数据降维的目的,节约模型训练时间。具体地,采用方差分析时,是将方差小于预设方差阈值的数据列去除,故方差分析的准确性取决于预设方差阈值,因此,为进一步去除冗余数据,且能保证丢失的数据信息尽可能少,本实施例中还需对第二样本数据进行奇异值分解,以除冗余数据,实现数据压缩的目的,保证目标样本数据的质量。
本实施例中,通过对目标样本数据进行方差分析,去除方差小于预设方差阈值的数据,得到第二样本数据,以去除冗余数据,达到减少数据列数的同时保证丢失的数据信息尽可能少,节约模型训练时间。然后,对第二样本数据进行奇异值分解,更新目标样本数据,以进一步去除冗余数据,保证目标样本数据的质量。
在一实施例中,改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,多粒子扫描算法对应至少一个滑动窗口,如图6所示,步骤S90中,具体包括如下步骤:
S91:采用多粒子扫描算法,按照至少一个滑动窗口,对目标样本数据进行多粒子扫描,得到至少一个中间数据。
其中,多粒子扫描是指采用滑动窗口对目标样本数据进行扫描,得到至少一个中间数据。本实施例中,可设置不同维度的滑动窗口,可以理解地,该滑动窗口可为i*j的窗口。 例如目标表样本数据行标签为第i周,则滑动窗口window_size可取2(每2周)、4(每个月)、12(每个季度)等。需要说明的是,该滑动窗口可扫描至少一个特征画像,即可扫描每一列、每两列、每j列,以极大化地搜寻特征与标签集、特征与特征之间的内在关联性。
S92:基于池化层,对至少一个中间数据进行池化处理,得到待训练数据。
本实施例中,通过池化层对至少一个中间数据进行池化处理,得到待训练数据,以达到对数据进行降维的目的,减小计算量,提高模型训练效率。
S93:采用级联随机森林算法对待训练数据进行训练,获取目标预测模型。
具体地,多粒度级联随机森林算法基于神经网络集成的思想,将第i次complete-random tree forest预测得到的标签列cforest i和random forest预测得到的标签列rforest i作为不断加入目标样本数据的画像列,以进一步特征扩充,最终得到如下的特征画像[orgf 1,orgf 2,K,orgf n,cforest 1,rforest 1,K,cforest k,rforest k]。其中,orgf是目标样本数据。最后,将该特征画像输入到最后m个(m一般取3~5,一般数量级取3,千万数量级取3~4,超过千万数量级取4~5)random forest中进行预测,取最终的Max值作为最终的预测概率值。
具体地,将得到的待训练数据输入到级联森林中进行训练。例如,本实施例中采用三种维度的滑动窗口,首先使用第一维度的滑动窗口进行扫描得到一特征向量,再将该原始特征向量输入到complete-random tree forest和random forest中,分别得到两个预测序列(即cforest i和rforest i),再将这两个预测序列拼接,得到第一特征向量,将原始特征向量输入到第一层级联森林中进行训练,得到第一预测序列。然后将得到的第一预测序列与第一特征向量进行拼接,得到第二特征向量,作为第二层的级联森林的输入数据;第二层级联森林训练得到的第二预测序列再与第二维度的滑动窗得到的第三特征向量(与第一特征向量的获取方法相同)进行拼接,作为第三层级联森林的输入数据;第三层级联森林训练得到的第三预测序列再与第三维度的滑动窗得到的第四特征向量进行拼接,作为下一层的输入,不断重复上述过程,直至收敛,得到目标预测模型。
本实施例中,通过采用多粒子扫描算法,按照至少一个滑动窗口,对目标样本数据进行多粒子扫描,得到至少一个中间数据,以极大化地搜寻特征与标签集、特征与特征之间的内在关联性。然后,通过结合池化层,对至少一个中间数据进行池化处理,得到待训练数据,以将机器学习和神经网络思想相结合,获取更多直观无法获取的信息,来丰富模型,进一步提高模型预测准确率。
在一实施例中,如图7所示,步骤S92中,即基于池化层,对至少一个中间数据进行池化处理,得到待训练数据,具体包括如下步骤:
S921:选取相邻的两个中间数据作为一组待处理数据组,以得到中间数据对应的至少一组待处理数据组。
S922:对每组待处理数据组进行取平均运算,得到第一数据序列。
S923:对每组待处理数据组进行最小值运算,得到第二数据序列,第二数据列中包括每组待处理数据组的两个中间数据中的最小值。
S924:对每组待处理数据组进行最大值运算,得到第三数据序列,第三数据列中包括每组待处理数据组的两个中间数据中的最大值。
S925:将第一数据序列、第二数据序列和第三数据序列进行拼接,得到待训练数据。
具体地,从业务逻辑层面上来说,模型预测需要更多线性、或者非线性的方法来对数 据进行空间扭曲,从而获取更多直观无法获取的信息,来丰富模型,故本实施例中,采用三种池化方式对至少一个中间数据进行池化,再对每种方式进行池化所得到的结果进行集成,得到待训练数据,以获取更多直观无法获取的信息,来丰富模型,并可充分保留数据特征。假设中间是中间数据中某一列画像数据为Feature:f 1,f 2,f 3,f 4,f 5,K f n,则采用如下三种池化方式对至少一个中间数据进行池化。
Feature_new_1:(f 1+f 2)/2,(f 2+f 3)/2,K,(f n-1+f n)/2
Feature_new_2:max(f 1,f 2),max(f 2,f 3),K,max(f n-1,f n)
Feature_new_3:min(f 1,f 2),min(f 2,f 3),K,min(f n-1,f n)
本实施例中,通过采用三种池化方式对至少一个中间数据进行池化,再对每种方式进行池化所得到的结果进行集成,得到待训练数据,以充分保留数据特征,保证样本数据质量,提高模型预测准确率。
应理解,上述实施例中各步骤的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本申请实施例的实施过程构成任何限定。
在一实施例中,提供一种数据智能分析装置,该数据智能分析装置与上述实施例中数据智能分析方法一一对应。如图8所示,该数据智能分析装置包括舆情数据获取模块10、命中词条确定模块20、舆情指数获取模块30、第一画像数据获取模块40、原始样本数据获取模块50、待处理样本数据获取模块60、滞后样本数据获取模块70、目标样本数据获取模块80、和目标预测模型获取模块90。各功能模块详细说明如下:
舆情数据获取模块10,用于按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据。
命中词条确定模块20,用于基于舆情数据,确定至少一个命中词条;命中词条对应一舆情因子。
舆情指数获取模块30,用于获取历史单位时间内的医疗数据和命中词条对应的舆情指数;舆情指数携带时间标签。
第一画像数据获取模块40,用于将舆情因子和携带时间标签的舆情指数作为第一画像数据。
原始样本数据获取模块50,用于基于第一画像数据和医疗数据,获取原始样本数据。
待处理样本数据获取模块60,用于对原始样本数据进行数据清洗,得到待处理样本数据。
滞后样本数据获取模块70,用于对待处理样本数据进行滞后处理,得到滞后样本数据。
目标样本数据获取模块80,用于对滞后样本数据进行特征扩充处理,获取目标样本数据。
目标预测模型获取模块90,用于采用改进多粒度级联随机森林算法对目标样本数据进行训练,得到目标预测模型;改进多粒度级联随机森林算法包括一池化层,池化层用于保留数据特征。
具体地,待处理样本数据获取模块包括第一样本数据获取单元、异常值获取单元和待处理样本数据获取单元。
第一样本数据获取单元,用于对原始样本数据进行缺失值填充,得到第一样本数据。
异常值获取单元,用于对第一样本数据进行异常值检测,得到至少一个异常值,将异常值标记为空。
待处理样本数据获取单元,用于对标记为空的异常值进行缺失值填充,得到待处理样本数据。
具体地,目标样本数据获取模块包括特征值获取单元和目标样本数据获取单元。
特征值获取单元,用于对滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值。
目标样本数据获取单元,用于将特征值与滞后样本数据进行拼接,获取目标样本数据。
具体地,该数据智能分析装置包括第二样本数据获取单元和目标样本数据更新单元。
第二样本数据获取单元,用于对目标样本数据进行方差分析,去除方差小于预设方差阈值的数据,得到第二样本数据。
目标样本数据更新单元,用于对第二样本数据进行奇异值分解,以更新目标样本数据。
具体地,改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,多粒子扫描算法对应至少一个滑动窗口;目标预测模型获取模块包括目标预测模型、待训练数据获取单元和目标预测模型获取单元。
中间数据获取单元,用于采用多粒子扫描算法,按照至少一个滑动窗口,对目标样本数据进行多粒子扫描,得到至少一个中间数据。
待训练数据获取单元,用于基于池化层,对至少一个中间数据进行池化处理,得到待训练数据。
目标预测模型获取单元,用于采用级联随机森林算法对待训练数据进行训练,获取目标预测模型。
具体地,待训练数据获取单元包括待处理数据组获取子单元、第一数据序列获取子单元、第二数据序列获取子单元、第三数据序列获取子单元、和待训练数据获取子单元。
待处理数据组获取子单元,用于选取相邻的两个中间数据作为一组待处理数据组,以得到中间数据对应的至少一组待处理数据组。
第一数据序列获取子单元,用于对每组待处理数据组进行取平均运算,得到第一数据序列。
第二数据序列获取子单元,用于对每组待处理数据组进行最小值运算,得到第二数据序列,第二数据列中包括每组待处理数据组的两个中间数据中的最小值。
第三数据序列获取子单元,用于对每组待处理数据组进行最大值运算,得到第三数据序列,第三数据列中包括每组待处理数据组的两个中间数据中的最大值。
待训练数据获取子单元,用于将第一数据序列、第二数据序列和第三数据序列进行拼接,得到待训练数据。
关于数据智能分析装置的具体限定可以参见上文中对于数据智能分析方法的限定,在此不再赘述。上述数据智能分析装置中的各个模块可全部或部分通过软件、硬件及其组合来实现。上述各模块可以硬件形式内嵌于或独立于计算机设备中的处理器中,也可以以软件形式存储于计算机设备中的存储器中,以便于处理器调用执行以上各个模块对应的操作。
在一个实施例中,提供了一种计算机设备,该计算机设备可以是服务器,其内部结构图可以如图10所示。该计算机设备包括通过系统总线连接的处理器、存储器、网络接口和数据库。其中,该计算机设备的处理器用于提供计算和控制能力。该计算机设备的存储器包括可读存储介质、内存储器。该可读存储介质存储有操作系统、计算机可读指令和数据库。该内存储器为可读存储介质中的操作系统和计算机可读指令的运行提供环境。该计算机设备的数据库用于存储执行数据智能分析方法过程中生成或获取的数据,如目标样本数据。该计算机设备的网络接口用于与外部的终端通过网络连接通信。该计算机可读指令被处理器执行时以实现一种数据智能分析方法。
在一个实施例中,提供了一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机可读指令,处理器执行计算机可读指令时实现上述实施例中的数据智能分析方法的步骤,例如图2所示的步骤S10-S90,或者图3至图7中所示的步骤。 或者,处理器执行计算机可读指令时实现数据智能分析装置这一实施例中的各模块/单元的功能,例如图8所示的各模块/单元的功能,为避免重复,这里不再赘述。
在一实施例中,提供一个或多个存储有计算机可读指令的可读存储介质,所述计算机可读存储介质存储有计算机可读指令,其特征在于,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行时实现上述实施例中数据智能分析方法的步骤,例如图2所示的步骤S10-S90,或者图3至图7中所示的步骤,为避免重复,这里不再赘述。或者,该计算机可读指令被处理器执行时实现上述数据智能分析装置这一实施例中的各模块/单元的功能,例如图8所示的各模块/单元的功能,为避免重复,这里不再赘述。本实施例中的可读存储介质包括非易失性可读存储介质和易失性可读存储介质。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,所述的计算机可读指令可存储于一非易失性计算机可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
所属领域的技术人员可以清楚地了解到,为了描述的方便和简洁,仅以上述各功能单元、模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能单元、模块完成,即将所述装置的内部结构划分成不同的功能单元或模块,以完成以上描述的全部或者部分功能。
以上所述实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围,均应包含在本申请的保护范围之内。

Claims (20)

  1. 一种数据智能分析方法,其特征在于,包括:
    按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;
    基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;
    获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;
    将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;
    基于所述第一画像数据和所述医疗数据,获取原始样本数据;
    对所述原始样本数据进行数据清洗,得到待处理样本数据;
    对所述待处理样本数据进行滞后处理,得到滞后样本数据;
    对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
    采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
  2. 如权利要求1所述数据智能分析方法,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:
    获取气象因子和对应的气象数据;
    将所述气象因子和对应的气象数据作为第二画像数据;
    所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:
    将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
  3. 如权利要求1所述数据智能分析方法,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;
    对所述原始样本数据进行缺失值填充,得到第一样本数据;
    对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;
    对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
  4. 如权利要求1所述数据智能分析方法,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:
    对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;
    将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
  5. 如权利要求1或4所述数据智能分析方法,其特征在于,在所述获取目标样本数据之后,所述数据智能分析方法包括:
    对所述目标样本数据进行方差分析,去除方差小于预设方差阈值的数据,得到第二样本数据;
    对所述第二样本数据进行奇异值分解,以更新所述目标样本数据。
  6. 如权利要求1所述数据智能分析方法,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;
    所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:
    采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;
    基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;
    采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
  7. 如权利要求6所述数据智能分析方法,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:
    选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;
    对每组所述待处理数据组进行取平均运算,得到第一数据序列;
    对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;
    对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;
    将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
  8. 一种数据智能分析装置,其特征在于,包括:
    舆情数据获取模块,用于按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;
    命中词条确定模块,基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;
    舆情指数获取模块,用于获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;
    第一画像数据获取模块,用于将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;
    原始样本数据获取模块,用于基于所述第一画像数据和所述医疗数据,获取原始样本数据;
    待处理样本数据获取模块,用于对所述原始样本数据进行数据清洗,得到待处理样本数据;
    滞后样本数据获取模块,用于对所述待处理样本数据进行滞后处理,得到滞后样本数据;
    目标样本数据获取模块,用于对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
    目标预测模型获取模块,用于采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
  9. 一种计算机设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,其特征在于,所述处理器执行所述计算机可读指令时实现如下步骤:
    按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;
    基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;
    获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;
    将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;
    基于所述第一画像数据和所述医疗数据,获取原始样本数据;
    对所述原始样本数据进行数据清洗,得到待处理样本数据;
    对所述待处理样本数据进行滞后处理,得到滞后样本数据;
    对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
    采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
  10. 如权利要求9所述的计算机设备,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:
    获取气象因子和对应的气象数据;
    将所述气象因子和对应的气象数据作为第二画像数据;
    所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:
    将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
  11. 如权利要求9所述的计算机设备,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;
    对所述原始样本数据进行缺失值填充,得到第一样本数据;
    对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;
    对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
  12. 如权利要求9所述的计算机设备,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:
    对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;
    将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
  13. 如权利要求9所述的计算机设备,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;
    所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:
    采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;
    基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;
    采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
  14. 如权利要求13所述的计算机设备,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:
    选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;
    对每组所述待处理数据组进行取平均运算,得到第一数据序列;
    对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;
    对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;
    将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
  15. 一个或多个存储有计算机可读指令的可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行如下步骤:
    按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;
    基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;
    获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;
    将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;
    基于所述第一画像数据和所述医疗数据,获取原始样本数据;
    对所述原始样本数据进行数据清洗,得到待处理样本数据;
    对所述待处理样本数据进行滞后处理,得到滞后样本数据;
    对所述滞后样本数据进行特征扩充处理,获取目标样本数据;
    采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模 型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
  16. 如权利要求15所述的可读存储介质,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:
    获取气象因子和对应的气象数据;
    将所述气象因子和对应的气象数据作为第二画像数据;
    所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:
    将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
  17. 如权利要求15所述的可读存储介质,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;
    对所述原始样本数据进行缺失值填充,得到第一样本数据;
    对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;
    对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
  18. 如权利要求15所述的可读存储介质,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:
    对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;
    将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
  19. 如权利要求15所述的可读存储介质,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;
    所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:
    采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;
    基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;
    采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
  20. 如权利要求19所述的可读存储介质,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:
    选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;
    对每组所述待处理数据组进行取平均运算,得到第一数据序列;
    对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;
    对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;
    将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
PCT/CN2019/116942 2019-08-19 2019-11-11 数据智能分析方法、装置、计算机设备及存储介质 Ceased WO2020215671A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
JP2021506707A JP7165809B2 (ja) 2019-08-19 2019-11-11 データのインテリジェント分析方法、装置、コンピュータ機器及び記憶媒体
SG11202008324YA SG11202008324YA (en) 2019-08-19 2019-11-11 Intelligent data analysis method and apparatus, computer device, and storage medium
US17/168,925 US20210158973A1 (en) 2019-08-19 2021-02-05 Intelligent data analysis method and device, computer device, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910763137.5 2019-08-19
CN201910763137.5A CN110675959B (zh) 2019-08-19 2019-08-19 数据智能分析方法、装置、计算机设备及存储介质

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/168,925 Continuation US20210158973A1 (en) 2019-08-19 2021-02-05 Intelligent data analysis method and device, computer device, and storage medium

Publications (1)

Publication Number Publication Date
WO2020215671A1 true WO2020215671A1 (zh) 2020-10-29

Family

ID=69075500

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/116942 Ceased WO2020215671A1 (zh) 2019-08-19 2019-11-11 数据智能分析方法、装置、计算机设备及存储介质

Country Status (5)

Country Link
US (1) US20210158973A1 (zh)
JP (1) JP7165809B2 (zh)
CN (1) CN110675959B (zh)
SG (1) SG11202008324YA (zh)
WO (1) WO2020215671A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112395401A (zh) * 2020-11-17 2021-02-23 中国平安人寿保险股份有限公司 自适应负样本对采样方法、装置、电子设备及存储介质
CN113159181A (zh) * 2021-04-23 2021-07-23 湖南大学 基于改进的深度森林的工业控制网络异常检测方法和系统
CN113268921A (zh) * 2021-05-13 2021-08-17 西安交通大学 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质
CN116266211A (zh) * 2021-12-17 2023-06-20 中国移动通信有限公司研究院 数据关联性分析方法、装置、电子设备及可读存储介质

Families Citing this family (20)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111738286A (zh) * 2020-03-17 2020-10-02 北京京东乾石科技有限公司 故障判定及其模型训练方法、装置、设备及存储介质
CN111581877B (zh) * 2020-03-25 2025-01-21 中国平安人寿保险股份有限公司 样本模型训练方法、样本生成方法、装置、设备及介质
CN111986763B (zh) * 2020-09-03 2024-05-14 深圳平安智慧医健科技有限公司 疾病数据分析方法、装置、电子设备及存储介质
CN112134862B (zh) * 2020-09-11 2023-09-08 国网电力科学研究院有限公司 基于机器学习的粗细粒度混合网络异常检测方法及装置
CN112434208B (zh) * 2020-12-03 2024-05-07 百果园技术(新加坡)有限公司 一种孤立森林的训练及其网络爬虫的识别方法与相关装置
CN112579587B (zh) * 2020-12-29 2024-07-02 纽扣互联(北京)科技有限公司 数据清洗方法及装置、设备和存储介质
CN112862179A (zh) * 2021-02-03 2021-05-28 国网山西省电力公司吕梁供电公司 一种用能行为的预测方法、装置及计算机设备
CN113672366B (zh) * 2021-08-05 2025-04-04 Oppo广东移动通信有限公司 应用程序管理方法、装置、电子设备及存储介质
CN114358422B (zh) * 2022-01-04 2024-12-27 中国工商银行股份有限公司 研发进度的异常预测方法及装置、存储介质和电子设备
CN114547970B (zh) * 2022-01-25 2024-02-20 中国长江三峡集团有限公司 一种水电厂顶盖排水系统异常智能诊断方法
CN114581252B (zh) * 2022-03-03 2024-04-05 平安科技(深圳)有限公司 目标案件的预测方法和装置、电子设备、存储介质
CN115146696A (zh) * 2022-04-18 2022-10-04 华北电力大学 一种基于多任务学习的综合能源系统运行状态监测方法
CN115129769B (zh) * 2022-06-01 2024-08-20 苏州大学 一种居民出行调查扩样方法、装置及存储介质
CN115312193B (zh) * 2022-08-05 2025-05-06 天津医科大学总医院 医学潜在相关指标风险监控系统、方法、终端及存储介质
CN115547508B (zh) * 2022-11-29 2023-03-21 联仁健康医疗大数据科技股份有限公司 数据校正方法、装置、电子设备及存储介质
CN117251474A (zh) * 2022-12-19 2023-12-19 中国南方电网有限责任公司 电力系统碳流计算的数据库构建方法、装置、设备和介质
KR102653187B1 (ko) * 2023-02-23 2024-04-01 주식회사 쇼퍼하우스 웹크롤링 기반 학습용 데이터 전처리 전자 장치 및 그 방법
CN117009813B (zh) * 2023-07-05 2025-10-31 上海幻电信息科技有限公司 样本拼接训练方法及装置
CN117786560B (zh) * 2024-02-28 2024-05-07 通用电梯股份有限公司 一种基于多粒度级联森林的电梯故障分类方法及电子设备
CN119272178B (zh) * 2024-12-10 2025-04-29 上海融和元储能源有限公司 数据处理方法以及装置

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017186048A1 (zh) * 2016-04-27 2017-11-02 第四范式(北京)技术有限公司 展示预测模型的方法、装置及调整预测模型的方法、装置
CN108417274A (zh) * 2018-03-06 2018-08-17 东南大学 流行病预测方法、系统及设备
CN108647249A (zh) * 2018-04-18 2018-10-12 平安科技(深圳)有限公司 舆情数据预测方法、装置、终端及存储介质
CN109241987A (zh) * 2018-06-29 2019-01-18 南京邮电大学 基于加权的深度森林的机器学习方法
CN109656918A (zh) * 2019-01-04 2019-04-19 平安科技(深圳)有限公司 流行病发病指数的预测方法、装置、设备及可读存储介质

Family Cites Families (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7547283B2 (en) * 2000-11-28 2009-06-16 Physiosonics, Inc. Methods for determining intracranial pressure non-invasively
WO2006087854A1 (ja) * 2004-11-25 2006-08-24 Sharp Kabushiki Kaisha 情報分類装置、情報分類方法、情報分類プログラム、情報分類システム
US9746985B1 (en) * 2008-02-25 2017-08-29 Georgetown University System and method for detecting, collecting, analyzing, and communicating event-related information
US8255346B2 (en) * 2009-11-11 2012-08-28 International Business Machines Corporation Methods and systems for variable group selection and temporal causal modeling
ES2388413B1 (es) * 2010-07-01 2013-08-22 Telefónica, S.A. Método para la clasificación de videos.
CN105608200A (zh) * 2015-12-28 2016-05-25 湖南蚁坊软件有限公司 一种网络舆论趋势预测分析方法
KR20180052489A (ko) * 2016-11-10 2018-05-18 주식회사 레드아이스 사용자 경험분석 및 환경요인에 기초한 크로스보더 전자상거래 상품 추천 방법
JP6736530B2 (ja) * 2017-09-13 2020-08-05 ヤフー株式会社 予測装置、予測方法、及び予測プログラム
CN107918772B (zh) * 2017-12-10 2021-04-30 北京工业大学 基于压缩感知理论和gcForest的目标跟踪方法
CN108389631A (zh) * 2018-02-07 2018-08-10 平安科技(深圳)有限公司 水痘发病预警方法、服务器及计算机可读存储介质
CN108288502A (zh) * 2018-04-11 2018-07-17 平安科技(深圳)有限公司 疾病预测方法及装置、计算机装置及可读存储介质
CN108648829A (zh) * 2018-04-11 2018-10-12 平安科技(深圳)有限公司 疾病预测方法及装置、计算机装置及可读存储介质
CN108921702A (zh) * 2018-06-04 2018-11-30 北京至信普林科技有限公司 基于大数据的园区招商方法及装置

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017186048A1 (zh) * 2016-04-27 2017-11-02 第四范式(北京)技术有限公司 展示预测模型的方法、装置及调整预测模型的方法、装置
CN108417274A (zh) * 2018-03-06 2018-08-17 东南大学 流行病预测方法、系统及设备
CN108647249A (zh) * 2018-04-18 2018-10-12 平安科技(深圳)有限公司 舆情数据预测方法、装置、终端及存储介质
CN109241987A (zh) * 2018-06-29 2019-01-18 南京邮电大学 基于加权的深度森林的机器学习方法
CN109656918A (zh) * 2019-01-04 2019-04-19 平安科技(深圳)有限公司 流行病发病指数的预测方法、装置、设备及可读存储介质

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112395401A (zh) * 2020-11-17 2021-02-23 中国平安人寿保险股份有限公司 自适应负样本对采样方法、装置、电子设备及存储介质
CN112395401B (zh) * 2020-11-17 2024-06-04 中国平安人寿保险股份有限公司 自适应负样本对采样方法、装置、电子设备及存储介质
CN113159181A (zh) * 2021-04-23 2021-07-23 湖南大学 基于改进的深度森林的工业控制网络异常检测方法和系统
CN113159181B (zh) * 2021-04-23 2022-06-10 湖南大学 基于改进的深度森林的工业控制系统异常检测方法和系统
CN113268921A (zh) * 2021-05-13 2021-08-17 西安交通大学 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质
CN113268921B (zh) * 2021-05-13 2022-12-09 西安交通大学 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质
CN116266211A (zh) * 2021-12-17 2023-06-20 中国移动通信有限公司研究院 数据关联性分析方法、装置、电子设备及可读存储介质

Also Published As

Publication number Publication date
JP2021532501A (ja) 2021-11-25
CN110675959A (zh) 2020-01-10
SG11202008324YA (en) 2020-11-27
JP7165809B2 (ja) 2022-11-04
US20210158973A1 (en) 2021-05-27
CN110675959B (zh) 2023-07-07

Similar Documents

Publication Publication Date Title
WO2020215671A1 (zh) 数据智能分析方法、装置、计算机设备及存储介质
CN112365171B (zh) 基于知识图谱的风险预测方法、装置、设备及存储介质
US9442929B2 (en) Determining documents that match a query
TWI740891B (zh) 利用訓練資料訓練模型的方法和訓練系統
US20200175314A1 (en) Predictive data analytics with automatic feature extraction
WO2022217713A1 (zh) 症候群监测预警方法、装置、计算机设备及存储介质
WO2021012790A1 (zh) 页面数据生成方法、装置、计算机设备及存储介质
US11921681B2 (en) Machine learning techniques for predictive structural analysis
WO2021164171A1 (zh) 知识库中数据处理方法、装置、计算机设备和存储介质
US20230394352A1 (en) Efficient multilabel classification by chaining ordered classifiers and optimizing on uncorrelated labels
WO2021114613A1 (zh) 基于人工智能的故障节点识别方法、装置、设备和介质
CN114547257A (zh) 类案匹配方法、装置、计算机设备及存储介质
WO2021218037A1 (zh) 目标检测方法、装置、计算机设备和存储介质
EP4318265A1 (en) Insight mining using machine learning
CN118734247B (zh) 基于moe架构的智能城市中枢数据融合计算模型训练方法、预警方法及设备
CN112699668A (zh) 一种化学信息抽取模型的训练方法、抽取方法、装置、设备及存储介质
CN115827797A (zh) 一种基于大数据的环境数据分析整合方法及系统
WO2019080419A1 (zh) 标准知识库的构建方法、电子装置及存储介质
CN116434973A (zh) 基于人工智能的传染病预警方法、装置、设备及介质
US12536430B2 (en) Machine learning techniques for efficient data pattern recognition across structured data objects
WO2023040145A1 (zh) 基于人工智能的文本分类方法、装置、电子设备及介质
US11907191B2 (en) Content based log retrieval by using embedding feature extraction
WO2024199121A1 (zh) 网络模型的结构化稀疏、图像分类方法及相关装置
HK40017556B (zh) 数据智能分析方法、装置、计算机设备及存储介质
HK40017556A (zh) 数据智能分析方法、装置、计算机设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19926035

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021506707

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19926035

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 19926035

Country of ref document: EP

Kind code of ref document: A1