WO2020215671A1 - 数据智能分析方法、装置、计算机设备及存储介质 - Google Patents
数据智能分析方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2020215671A1 WO2020215671A1 PCT/CN2019/116942 CN2019116942W WO2020215671A1 WO 2020215671 A1 WO2020215671 A1 WO 2020215671A1 CN 2019116942 W CN2019116942 W CN 2019116942W WO 2020215671 A1 WO2020215671 A1 WO 2020215671A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- sample data
- processed
- public opinion
- target
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H40/00—ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices
- G16H40/60—ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices
- G16H40/67—ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices for remote operation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/20—Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
- G06F16/21—Design, administration or maintenance of databases
- G06F16/215—Improving data quality; Data cleansing, e.g. de-duplication, removing invalid entries or correcting typographical errors
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/20—Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
- G06F16/24—Querying
- G06F16/245—Query processing
- G06F16/2458—Special types of queries, e.g. statistical queries, fuzzy queries or distributed queries
- G06F16/2465—Query processing support for facilitating data mining operations in structured databases
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/951—Indexing; Web crawling techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/243—Classification techniques relating to the number of classes
- G06F18/24323—Tree-organised classifiers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
- G06N20/20—Ensemble learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/01—Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/70—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/80—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for detecting, monitoring or modelling epidemics or pandemics, e.g. flu
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02A—TECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
- Y02A90/00—Technologies having an indirect contribution to adaptation to climate change
- Y02A90/10—Information and communication technologies [ICT] supporting adaptation to climate change, e.g. for weather forecasting or climate simulation
Definitions
- This application relates to the field of data prediction technology, and in particular to a data intelligent analysis method, device, computer equipment and storage medium.
- the embodiments of the present application provide a data intelligent analysis method, device, computer equipment, and storage medium to solve the current problem of low model prediction accuracy when data prediction is performed on lagging data.
- An intelligent data analysis method including:
- the hit entry corresponds to a public opinion factor
- the public opinion index carries a time label
- the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain a target prediction model; the improved multi-granularity cascaded random forest algorithm includes a pooling layer for retaining data features.
- a data intelligent analysis device including:
- the public opinion data acquisition module is used to crawl the public opinion data obtained by the third-party information platform using a crawler tool according to preset keywords.
- the hit term determining module is used to determine at least one hit term based on public opinion data; the hit term corresponds to a public opinion factor.
- the public opinion index acquisition module is used to acquire medical data within historical unit time and the public opinion index corresponding to the hit entry; the public opinion index carries a time label.
- the first portrait data acquisition module is configured to use the public opinion factor and the public opinion index carrying the time tag as the first portrait data.
- the original sample data obtaining module is used to obtain original sample data based on the first portrait data and the medical data.
- the sample data acquisition module to be processed is used to perform data cleaning on the original sample data to obtain sample data to be processed;
- the lagging sample data acquisition module is used to perform lagging processing on the sample data to be processed to obtain lagging sample data
- the target sample data acquisition module is used to perform feature expansion processing on the lagging sample data to acquire target sample data
- the target prediction model acquisition module is used to train the target sample data by an improved multi-granularity cascaded random forest algorithm to obtain a target prediction model;
- the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pool
- the transformation layer is used to retain data characteristics.
- a computer device including a memory, a processor, and computer readable instructions stored in the memory and capable of running on the processor, and the processor implements the above data intelligent analysis method when the processor executes the computer readable instructions A step of.
- a readable storage medium stores computer readable instructions, and when the computer readable instructions are executed by a processor, the steps of the above intelligent data analysis method are realized.
- FIG. 1 is a schematic diagram of an application environment of a data intelligent analysis method in an embodiment of the present application
- Figure 2 is a flowchart of a data intelligent analysis method in an embodiment of the present application
- FIG. 3 is a specific flowchart of step S60 in FIG. 2;
- FIG. 4 is a specific flowchart of step S80 in FIG. 2;
- FIG. 5 is a flowchart of a data intelligent analysis method in an embodiment of the present application.
- FIG. 6 is a specific flowchart of step S90 in FIG. 2;
- FIG. 7 is a specific flowchart of step S92 in FIG. 6;
- FIG. 8 is a schematic diagram of an intelligent data analysis device in an embodiment of the present application.
- Fig. 9 is a schematic diagram of a computer device in an embodiment of the present application.
- the data intelligent analysis method provided in the embodiments of this application can be applied to this method can be applied to a data intelligent analysis tool, which can train different samples according to the sample data corresponding to different topics (such as chickenpox, flu, etc.)
- the prediction model especially for the sample data with lag, can effectively guarantee the accuracy of the model prediction.
- the data intelligent analysis method can be applied in the application environment as shown in Fig. 1, in which the computer equipment communicates with the server through the network.
- Computer equipment can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and portable wearable devices.
- the server can be implemented as an independent server.
- a method for intelligent data analysis is provided.
- the method is applied to the server in FIG. 1 as an example for description, and includes the following steps:
- S10 Use a crawler tool to crawl public opinion data obtained by a third-party information platform according to preset keywords.
- the preset keywords are preset keywords related to communicable diseases, such as chickenpox, redness, pruritic herpes and blisters.
- Public opinion data refers to the text data publicly released by different users in a third-party information platform to reflect the occurrence of social events. Specifically, with the rapid development of the information age at any time, users are more inclined to use various information platforms to query the required information, such as querying whether they have a disease based on their own symptoms, etc.
- a crawler tool is used to crawl public opinion data containing preset keywords in a third-party information platform (such as Baidu, Weibo, or WeChat) .
- a third-party information platform such as Baidu, Weibo, or WeChat.
- S20 Determine at least one hit term based on the public opinion data, and the hit term corresponds to a public opinion factor.
- the daily public opinion factors of 20 years in different regions are selected as another part of the portrait data.
- the public opinion factors include, but are not limited to, chickenpox, redness, pruritic herpes and blisters.
- the public opinion data includes at least one original entry (such as a Baidu entry). Specifically, the expert judges whether it is related to varicella according to the information contained in each crawled original entry, so as to determine at least one entry that is actually related to varicella as a hit entry. Then, according to the determined hit entry.
- Each hit entry corresponds to a public opinion factor.
- the public opinion factor refers to at least one factor related to a preset keyword contained in the hit entry, such as chickenpox, redness, pruritic herpes, and blisters.
- S30 Obtain the medical data in historical unit time and the public opinion index corresponding to the hit entry.
- the public opinion index carries a time label.
- medical data refers to the historical unit time of sentinel hospitals in different regions provided by the CDC, such as the historical number of patients (ie, tag data) in a unit time of 20 years.
- the unit time is the time label, and the unit time can be selected by the user and is not limited here.
- the unit time may be one day, one week, one month, one quarter, or one year, etc., which will not be listed here.
- the unit time is one week as an example.
- the public opinion index and medical data corresponding to the hit entry within the unit time are obtained.
- Each public opinion index carries a time label, which refers to the time label of the hit entry. release time.
- S40 Use the public opinion factor and the public opinion index carrying the time label as the first portrait data.
- the first portrait data refers to the public opinion factor and the public opinion index carrying the time label as the feature data for model training.
- the time interval can be one week, one month, one quarter, or one year.
- the processing of sample data will be different.
- public opinion factors such as chickenpox, redness and herpes
- the public opinion index of the Nth week can be used as row labels to establish partial portrait data.
- the public opinion index of the Nth week includes, but is not limited to, the average public opinion index of the Nth week (that is, the average public opinion index of 7 days a week), the largest public opinion index of the Nth week, and the smallest public opinion index of the Nth week.
- the first portrait data is used as the feature data for model training
- the medical data is used as the label data for model training to obtain the original sample data.
- S60 Perform data cleaning on the original sample data to obtain sample data to be processed.
- the original sample data may include missing values or abnormal values
- S70 Perform lag processing on the sample data to be processed to obtain lag sample data.
- lag processing is a feature engineering method, by expanding the sample data set, that is, increasing the feature portrait to collect more information.
- the corresponding sample data has lag, such as disease outbreaks or economic-related data.
- the prediction theme is to predict chickenpox, and there is a hysteresis in the outbreak of chickenpox. For example, if the temperature suddenly rises this week and the climate is humid, it may not cause the outbreak of chickenpox this week, but it will usher in the next week. During the outbreak period, it is necessary to perform lag processing on the sample data to be processed to ensure the accuracy of subsequent model predictions.
- the sample data to be processed is subjected to n lag processing (n generally takes 1 to 3). If n is 1, the sample data to be processed is subjected to lag processing, that is, the data of the original first week is regarded as the data of the second week. The data of the second week is used as the data of the third week, and so on, to get the lagging sample data.
- n is set to 2
- the sample data to be processed is subjected to lag processing, that is, the data of the original first week is regarded as the data of the third week, and the data of the second week is The data is used as the data of the fourth week, and so on, to obtain the lagging data, and integrate the lagging data obtained each time to obtain the lagging sample data to achieve the purpose of expanding the sample data set
- the concat function is used to combine the lag sample data obtained by multiple lag processing with the sample data to be processed into a data frame (DataFrame), that is, the lag sample data.
- DataFrame data frame
- the concat function is a function used to connect two or more arrays.
- the data frame is a two-dimensional data structure, that is, the data is arranged in a table of rows and columns.
- S80 Perform feature expansion processing on the lagging sample data to obtain target sample data.
- feature expansion processing is performed on the lagging sample data to obtain target sample data, so as to achieve the purpose of further expanding the sample data set.
- the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain a target prediction model.
- the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pooling layer is used to retain data features.
- the improved multi-granularity cascaded random forest algorithm is an algorithm that introduces the idea of pooling in the convolutional neural network into the multi-granularity cascaded random forest algorithm.
- the multi-granularity cascaded random forest algorithm is a decision tree integration method, which stacks multiple layers of random forests in a cascaded manner to obtain better feature representation and learning performance. The algorithm does not need to adjust the hyperparameters to achieve good results Performance.
- each layer in the multi-granularity cascaded random forest is composed of multiple random forests.
- the feature information of the input feature vector is learned through random forest, and then input to the next layer after processing.
- a variety of different types of random forests are selected for each layer, for example, two random forest structures are selected for each layer, namely completely-random tree forests and random forests. .
- a crawler tool is used to crawl public opinion data obtained by a third-party information platform, so as to determine at least one hit entry that is truly related to the predicted topic based on the public opinion data to ensure subsequent acquisitions The validity and accuracy of the public opinion factor. Then obtain the public opinion index and medical data corresponding to the hit entry per unit time. Finally, the public opinion factor and the public opinion index carrying the time label are used as the original sample data, so that the model analyzes the public opinion data in a unit time of 20 years in history. Then, by performing data cleaning on the original sample data, the sample data to be processed is obtained to ensure the quality of the sample data to be processed.
- the improved multi-granularity cascaded random forest algorithm is used to train the target sample data to obtain the target prediction model to obtain better feature representation and learning performance, and the algorithm can achieve good performance without excessive adjustment of hyperparameters. Ensure the accuracy of model predictions.
- the improved multi-granularity cascaded random forest algorithm also includes a pooling layer to fully retain data features and further improve the accuracy of model prediction.
- the data intelligent analysis method before step S10, further includes:
- this embodiment can select different portrait data according to different predicted themes.
- the prediction of chickenpox is taken as an example for illustration. Since the climate conditions and the chickenpox virus are very closely related, the history of different regions is selected.
- the daily meteorological factors of the year are used as part of the profile data.
- the meteorological factors include, but are not limited to, day and night temperature, day and night pressure, day and night precipitation, humidity, light intensity, and wind in different regions.
- the second portrait data refers to the feature data trained with meteorological factors and corresponding meteorological data as the model.
- the method of establishing portrait data for meteorological factors is consistent with step S40, that is, meteorological factors can be used as column labels, and the meteorological conditions of the Nth week can be used as row labels to establish the second portrait data.
- the meteorological conditions of the Nth week include but are not limited to the average meteorological conditions of the Nth week (such as average precipitation), the maximum meteorological conditions of the Nth week (such as maximum precipitation), and the minimum meteorological conditions of the Nth week (such as minimum Precipitation).
- step S50 that is, based on the first portrait data and medical data, obtaining the original sample data includes:
- S51 Use the first portrait data, the second portrait data and the medical data as original sample data.
- the meteorological situation is combined with the idea of mass dissemination of public opinion data to effectively predict the period of disease outbreak and improve the accuracy of model prediction.
- step S60 data cleaning is performed on the original sample data to obtain sample data to be processed, which specifically includes the following steps:
- S61 Perform missing value filling on the original sample data to obtain the first sample data.
- missing value filling methods include, but are not limited to, mean filling, mode filling, median filling, expectation maximization methods, multiple filling, and k-means clustering methods.
- the portrait data where the missing value is located is clustered, and the missing value is filled with the mean value of the clustered cluster.
- S62 Perform abnormal value detection on the first sample data to obtain at least one abnormal value, and mark the abnormal value as empty.
- S63 Fill the outliers marked as empty with missing values to obtain sample data to be processed.
- outlier detection includes, but is not limited to, the use of statistical variable analysis (such as box-plot analysis, average, maximum and minimum analysis, and the 3 ⁇ rule), distance-based methods, density-based outlier detection, and density-based outlier detection.
- Group point detection and isolation forest Isolation Forest
- an outlier is defined as a value whose deviation from the average value in a set of measured values exceeds 3 times the standard deviation, because in the normal distribution Under the assumption of, the probability of occurrence of values beyond 3 ⁇ from the average value is less than 0.003), that is, data exceeding ⁇ +3 ⁇ and data not exceeding ⁇ -3 ⁇ are regarded as abnormal values.
- the outliers are deleted and marked as null values, and then the outliers marked as null values are filled with missing values again to obtain the sample data to be processed.
- the sample data to be processed is obtained by filling the outliers marked as null with missing values, so as to avoid directly removing the sample data corresponding to the outliers, causing the sample data to lack this part of the feature and affecting the accuracy of model prediction The problem of rate.
- the original sample data is filled with missing values to obtain the first sample data, and then abnormal value detection is performed on the first sample data to obtain at least one abnormal value. And the missing values are processed to achieve the purpose of data cleaning and ensure the quality of sample data. Then, the obtained outliers are marked as empty, so that the outliers marked as empty are filled with missing values again to obtain the sample to be processed, and the original sample data is filled with missing values twice to ensure the quality and standard of the sample data To improve the accuracy of model predictions
- step S80 that is, performing feature expansion processing on lagging sample data to obtain target sample data, specifically includes the following steps:
- S81 Perform feature expansion on the lagging sample data to obtain a feature value corresponding to at least one statistical indicator.
- S82 Splicing the eigenvalues with the lagging sample data to obtain target sample data.
- the statistical indicators include but are not limited to the maximum, minimum, mean and standard deviation corresponding to each row of data.
- Each statistical indicator is added as a new column to the lagging sample data to expand the data set and increase the collection of feature images. Multiple feature information improves the accuracy of model prediction.
- the lagging sample data is a matrix. The eigenvalues and the lagging sample data are spliced together to obtain the target sample data, that is, N columns are added to the sample matrix. Value, minimum and average), the maximum, minimum, and average of the data corresponding to each row are the characteristic values.
- the feature value corresponding to at least one statistical indicator is obtained by feature expansion of the lagging sample data, and the feature value is spliced with the lagging sample data to obtain target sample data to expand the data set and increase the collection of feature images.
- Feature information improves the accuracy of model prediction.
- the data intelligent analysis method further includes the following steps:
- S111 Perform variance analysis on the target sample data, and remove data whose variance is less than a preset variance threshold to obtain second sample data.
- S112 Perform singular value decomposition on the second sample data to update the target sample data.
- the analysis of variance refers to the analysis based on the variance of the data column to remove the sequence with too small variance (that is, less than the preset variance threshold) to obtain the second sample data.
- the size of the variance describes the amount of information of a variable, and the sequence with too small variance is considered to contain less information, so all data columns with small variance are removed to achieve the effect of data dimensionality reduction and reduce the amount of data processing , Improve the efficiency of subsequent model training.
- the singular value decomposition of the second sample data is also required to remove redundant data, to achieve the purpose of data compression, and to ensure the quality of the target sample data.
- the improved multi-granularity cascaded random forest algorithm includes a multi-particle scanning algorithm and a cascaded random forest algorithm, and the multi-particle scanning algorithm corresponds to at least one sliding window, as shown in FIG. 6, in step S90, specifically including the following steps :
- S91 Use a multi-particle scanning algorithm to perform multi-particle scanning on the target sample data according to at least one sliding window to obtain at least one intermediate data.
- multi-particle scanning refers to scanning the target sample data using a sliding window to obtain at least one intermediate data.
- sliding windows of different dimensions can be set.
- the sliding window can be an i*j window.
- the target table sample data row label is the i-th week
- the sliding window window_size can be 2 (every 2 weeks), 4 (every month), 12 (every quarter), etc.
- the sliding window can scan at least one feature portrait, that is, every row, every two rows, and every j row can be scanned to maximize the search for the internal relationship between the feature and the tag set, and the feature and the feature.
- S92 Based on the pooling layer, perform pooling processing on at least one intermediate data to obtain data to be trained.
- At least one intermediate data is pooled through the pooling layer to obtain the data to be trained, so as to achieve the purpose of reducing the dimensionality of the data, reduce the amount of calculation, and improve the efficiency of model training.
- multi-granularity cascaded neural network ensemble Random Forest algorithm based on the idea, the i-th complete-random tree forest predicted label cforest i and column labels random forest column rforest i predicted as a target sample data continuously added
- the portrait column is expanded with further features, and finally the following feature portrait is obtained [orgf 1 ,orgf 2 ,K,orgf n ,cforest 1 ,rforest 1 ,K,cforest k ,rforest k ].
- orgf is the target sample data.
- the obtained data to be trained into the cascade forest for training.
- sliding windows of three dimensions are used.
- the sliding window of the first dimension is used to scan to obtain a feature vector, and then the original feature vector is input into the complete-random tree forest and random forest to obtain two predicted sequence (i.e. cforest i and rforest i), then the two predicted sequence spliced to give a first feature vector, the feature vector input into the original first hierarchical linking forest training, to obtain a first predicted sequence.
- the obtained first prediction sequence and the first feature vector are spliced together to obtain the second feature vector, which is used as the input data of the second layer of cascaded forest;
- the third feature vector obtained by the sliding window of the dimension (same method as the first feature vector) is spliced as the input data of the third-level cascaded forest;
- the third prediction sequence obtained by the third-level cascade forest training is then combined with the third
- the fourth feature vector obtained by the sliding window of the dimension is spliced and used as the input of the next layer, and the above process is repeated until convergence, and the target prediction model is obtained.
- multi-particle scanning is performed on the target sample data according to at least one sliding window to obtain at least one intermediate data to maximize the search for features and tag sets, and between features and features. Internal relevance. Then, by combining the pooling layer, at least one intermediate data is pooled to obtain the data to be trained, so as to combine machine learning and neural network ideas to obtain more intuitive and unobtainable information to enrich the model and further improve the model Forecast accuracy rate.
- step S92 that is, based on the pooling layer, pooling is performed on at least one intermediate data to obtain the data to be trained, which specifically includes the following steps:
- S921 Select two adjacent intermediate data as a group of to-be-processed data groups to obtain at least one group of to-be-processed data groups corresponding to the intermediate data.
- S922 Perform an averaging operation on each data group to be processed to obtain a first data sequence.
- S923 Perform a minimum value operation on each group of to-be-processed data groups to obtain a second data sequence.
- the second data column includes the minimum value of the two intermediate data of each group of to-be-processed data groups.
- S924 Perform a maximum value operation on each group of to-be-processed data groups to obtain a third data sequence.
- the third data column includes the maximum value of the two intermediate data of each group of to-be-processed data groups.
- S925 Join the first data sequence, the second data sequence, and the third data sequence to obtain data to be trained.
- model prediction requires more linear or nonlinear methods to spatially warp data, so as to obtain more information that is not intuitively available to enrich the model. Therefore, in this embodiment, The three pooling methods pool at least one intermediate data, and then integrate the results obtained by pooling in each method to obtain the data to be trained to obtain more intuitive and unobtainable information to enrich the model, and Fully retain data characteristics. Assuming that one column of portrait data in the intermediate data is Feature: f 1 , f 2 , f 3 , f 4 , f 5 , K f n , the following three pooling methods are used to pool at least one intermediate data.
- Feature_new_1 (f 1 +f 2 )/2,(f 2 +f 3 )/2,K,(f n-1 +f n )/2
- Feature_new_2 max(f 1 ,f 2 ),max(f 2 ,f 3 ),K,max(f n-1 ,f n )
- Feature_new_3 min(f 1 ,f 2 ),min(f 2 ,f 3 ),K,min(f n-1 ,f n )
- At least one intermediate data is pooled by using three pooling methods, and then the results obtained by pooling in each method are integrated to obtain the data to be trained, so as to fully retain the data characteristics and ensure the sample data Quality, improve the accuracy of model prediction.
- a data intelligent analysis device corresponds one-to-one with the data intelligent analysis method in the foregoing embodiment.
- the data intelligent analysis device includes a public opinion data acquisition module 10, a hit entry determination module 20, a public opinion index acquisition module 30, a first portrait data acquisition module 40, an original sample data acquisition module 50, and sample data to be processed
- the acquisition module 60, the lagging sample data acquisition module 70, the target sample data acquisition module 80, and the target prediction model acquisition module 90 is as follows:
- the public opinion data acquisition module 10 is used for crawling public opinion data obtained by a third-party information platform using a crawler tool according to preset keywords.
- the hit term determining module 20 is configured to determine at least one hit term based on public opinion data; the hit term corresponds to a public opinion factor.
- the public opinion index acquisition module 30 is used to acquire the medical data in historical unit time and the public opinion index corresponding to the hit entry; the public opinion index carries a time label.
- the first portrait data acquisition module 40 is configured to use the public opinion factor and the public opinion index carrying the time tag as the first portrait data.
- the original sample data obtaining module 50 is used to obtain original sample data based on the first portrait data and medical data.
- the sample data acquisition module 60 to be processed is used to perform data cleaning on the original sample data to obtain sample data to be processed.
- the lagging sample data acquisition module 70 is configured to perform lag processing on the sample data to be processed to obtain lagging sample data.
- the target sample data acquisition module 80 is configured to perform feature expansion processing on the lagging sample data to acquire target sample data.
- the target prediction model acquisition module 90 is used to train the target sample data with the improved multi-granularity cascaded random forest algorithm to obtain the target prediction model;
- the improved multi-granularity cascaded random forest algorithm includes a pooling layer, and the pooling layer is used for retention Data characteristics.
- the sample data acquisition module to be processed includes a first sample data acquisition unit, an abnormal value acquisition unit, and a sample data acquisition unit to be processed.
- the first sample data acquisition unit is used to fill in missing values on the original sample data to obtain the first sample data.
- the abnormal value obtaining unit is used to perform abnormal value detection on the first sample data to obtain at least one abnormal value, and mark the abnormal value as empty.
- the to-be-processed sample data acquisition unit is used to fill in the missing value of the outliers marked as empty to obtain the to-be-processed sample data.
- the target sample data acquisition module includes a feature value acquisition unit and a target sample data acquisition unit.
- the feature value obtaining unit is used to perform feature expansion on the lagging sample data to obtain a feature value corresponding to at least one statistical indicator.
- the target sample data acquisition unit is used to splice the feature value and the lagging sample data to acquire the target sample data.
- the data intelligent analysis device includes a second sample data acquisition unit and a target sample data update unit.
- the second sample data acquisition unit is configured to perform variance analysis on the target sample data, remove data whose variance is less than a preset variance threshold, and obtain second sample data.
- the target sample data update unit is used to perform singular value decomposition on the second sample data to update the target sample data.
- the improved multi-granularity cascaded random forest algorithm includes a multi-particle scanning algorithm and a cascaded random forest algorithm.
- the multi-particle scanning algorithm corresponds to at least one sliding window;
- the target prediction model acquisition module includes a target prediction model, a data acquisition unit to be trained, and a target Predictive model acquisition unit.
- the intermediate data acquisition unit is configured to use a multi-particle scanning algorithm to perform multi-particle scanning on the target sample data according to at least one sliding window to obtain at least one intermediate data.
- the to-be-trained data acquisition unit is configured to perform pooling processing on at least one intermediate data based on the pooling layer to obtain the to-be-trained data.
- the target prediction model acquisition unit is used to train the training data by using the cascaded random forest algorithm to obtain the target prediction model.
- the data acquisition unit to be trained includes a data group acquisition subunit to be processed, a first data sequence acquisition subunit, a second data sequence acquisition subunit, a third data sequence acquisition subunit, and a data sequence acquisition subunit to be trained.
- the to-be-processed data group acquiring subunit is used to select two adjacent intermediate data as a group of to-be-processed data groups to obtain at least one group of to-be-processed data groups corresponding to the intermediate data.
- the first data sequence obtaining subunit is used to perform an averaging operation on each group of to-be-processed data groups to obtain the first data sequence.
- the second data sequence obtaining subunit is used to perform minimum value operation on each group of to-be-processed data groups to obtain a second data sequence.
- the second data column includes the minimum value of the two intermediate data of each group of to-be-processed data groups.
- the third data sequence obtaining subunit is used to perform a maximum value operation on each group of to-be-processed data groups to obtain a third data sequence.
- the third data column includes the maximum value of the two intermediate data of each group of to-be-processed data groups.
- the to-be-trained data acquisition subunit is used to splice the first data sequence, the second data sequence and the third data sequence to obtain the to-be-trained data.
- Each module in the above-mentioned data intelligent analysis device can be implemented in whole or in part by software, hardware and a combination thereof.
- the foregoing modules may be embedded in the form of hardware or independent of the processor in the computer device, or may be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the foregoing modules.
- a computer device is provided.
- the computer device may be a server, and its internal structure diagram may be as shown in FIG. 10.
- the computer equipment includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide calculation and control capabilities.
- the memory of the computer device includes a readable storage medium and an internal memory.
- the readable storage medium stores an operating system, computer readable instructions, and a database.
- the internal memory provides an environment for the operation of the operating system and computer readable instructions in the readable storage medium.
- the database of the computer equipment is used to store data generated or obtained during the execution of the intelligent data analysis method, such as target sample data.
- the network interface of the computer device is used to communicate with an external terminal through a network connection.
- the computer readable instructions are executed by the processor to realize a data intelligent analysis method.
- a computer device including a memory, a processor, and computer-readable instructions stored in the memory and capable of running on the processor.
- the steps of the data intelligent analysis method are, for example, steps S10-S90 shown in FIG. 2 or the steps shown in FIG. 3 to FIG. 7.
- the processor implements the functions of the modules/units in this embodiment of the data intelligent analysis device when the processor executes the computer-readable instructions, such as the functions of the modules/units shown in FIG. 8. To avoid repetition, details are not described herein again.
- one or more readable storage media storing computer readable instructions are provided.
- the computer readable storage medium stores computer readable instructions, wherein the computer readable instructions are controlled by one or When multiple processors are executed, the one or more processors are executed to implement the steps of the data intelligent analysis method in the foregoing embodiment, for example, steps S10-S90 shown in FIG. 2 or shown in FIGS. 3 to 7 To avoid repetition, I won’t repeat them here.
- the computer-readable instruction is executed by the processor, the function of each module/unit in the embodiment of the above-mentioned data intelligent analysis device is realized, for example, the function of each module/unit shown in FIG. Repeat it again.
- the readable storage medium in this embodiment includes a nonvolatile readable storage medium and a volatile readable storage medium.
- Non-volatile memory may include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
- ROM read only memory
- PROM programmable ROM
- EPROM electrically programmable ROM
- EEPROM electrically erasable programmable ROM
- Volatile memory may include random access memory (RAM) or external cache memory.
- RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain Channel (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- Medical Informatics (AREA)
- Public Health (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- Biomedical Technology (AREA)
- General Physics & Mathematics (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Software Systems (AREA)
- Primary Health Care (AREA)
- Epidemiology (AREA)
- Pathology (AREA)
- Mathematical Physics (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computational Linguistics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Evolutionary Biology (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Business, Economics & Management (AREA)
- General Business, Economics & Management (AREA)
- Fuzzy Systems (AREA)
- Probability & Statistics with Applications (AREA)
- Quality & Reliability (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
Description
Claims (20)
- 一种数据智能分析方法,其特征在于,包括:按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;基于所述第一画像数据和所述医疗数据,获取原始样本数据;对所述原始样本数据进行数据清洗,得到待处理样本数据;对所述待处理样本数据进行滞后处理,得到滞后样本数据;对所述滞后样本数据进行特征扩充处理,获取目标样本数据;采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
- 如权利要求1所述数据智能分析方法,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:获取气象因子和对应的气象数据;将所述气象因子和对应的气象数据作为第二画像数据;所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
- 如权利要求1所述数据智能分析方法,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;对所述原始样本数据进行缺失值填充,得到第一样本数据;对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
- 如权利要求1所述数据智能分析方法,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
- 如权利要求1或4所述数据智能分析方法,其特征在于,在所述获取目标样本数据之后,所述数据智能分析方法包括:对所述目标样本数据进行方差分析,去除方差小于预设方差阈值的数据,得到第二样本数据;对所述第二样本数据进行奇异值分解,以更新所述目标样本数据。
- 如权利要求1所述数据智能分析方法,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
- 如权利要求6所述数据智能分析方法,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;对每组所述待处理数据组进行取平均运算,得到第一数据序列;对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
- 一种数据智能分析装置,其特征在于,包括:舆情数据获取模块,用于按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;命中词条确定模块,基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;舆情指数获取模块,用于获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;第一画像数据获取模块,用于将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;原始样本数据获取模块,用于基于所述第一画像数据和所述医疗数据,获取原始样本数据;待处理样本数据获取模块,用于对所述原始样本数据进行数据清洗,得到待处理样本数据;滞后样本数据获取模块,用于对所述待处理样本数据进行滞后处理,得到滞后样本数据;目标样本数据获取模块,用于对所述滞后样本数据进行特征扩充处理,获取目标样本数据;目标预测模型获取模块,用于采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
- 一种计算机设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,其特征在于,所述处理器执行所述计算机可读指令时实现如下步骤:按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;基于所述第一画像数据和所述医疗数据,获取原始样本数据;对所述原始样本数据进行数据清洗,得到待处理样本数据;对所述待处理样本数据进行滞后处理,得到滞后样本数据;对所述滞后样本数据进行特征扩充处理,获取目标样本数据;采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
- 如权利要求9所述的计算机设备,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:获取气象因子和对应的气象数据;将所述气象因子和对应的气象数据作为第二画像数据;所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
- 如权利要求9所述的计算机设备,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;对所述原始样本数据进行缺失值填充,得到第一样本数据;对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
- 如权利要求9所述的计算机设备,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
- 如权利要求9所述的计算机设备,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
- 如权利要求13所述的计算机设备,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;对每组所述待处理数据组进行取平均运算,得到第一数据序列;对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
- 一个或多个存储有计算机可读指令的可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行如下步骤:按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据;基于所述舆情数据,确定至少一个命中词条;所述命中词条对应一舆情因子;获取历史单位时间内的医疗数据和所述命中词条对应的舆情指数;所述舆情指数携带时间标签;将所述舆情因子和所述携带时间标签的舆情指数作为第一画像数据;基于所述第一画像数据和所述医疗数据,获取原始样本数据;对所述原始样本数据进行数据清洗,得到待处理样本数据;对所述待处理样本数据进行滞后处理,得到滞后样本数据;对所述滞后样本数据进行特征扩充处理,获取目标样本数据;采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模 型;所述改进多粒度级联随机森林算法包括一池化层,所述池化层用于保留数据特征。
- 如权利要求15所述的可读存储介质,其特征在于,所述按照预设关键词,采用爬虫工具爬取第三方信息平台所得到的舆情数据之前,所述数据智能分析方法还包括:获取气象因子和对应的气象数据;将所述气象因子和对应的气象数据作为第二画像数据;所述基于所述第一画像数据和所述医疗数据,获取原始样本数据,包括:将所述第一画像数据、所述第二画像数据和所述医疗数据作为原始样本数据。
- 如权利要求15所述的可读存储介质,其特征在于,所述对所述原始样本数据进行数据清洗,得到待处理样本数据,包括;对所述原始样本数据进行缺失值填充,得到第一样本数据;对所述第一样本数据进行异常值检测,得到至少一个异常值,将所述异常值标记为空;对所述标记为空的异常值进行缺失值填充,得到所述待处理样本数据。
- 如权利要求15所述的可读存储介质,其特征在于,所述对所述滞后样本数据进行特征扩充处理,获取目标样本数据,包括:对所述滞后样本数据进行特征扩充,得到至少一个统计指标对应的特征值;将所述特征值与所述滞后样本数据进行拼接,获取所述目标样本数据。
- 如权利要求15所述的可读存储介质,其特征在于,所述改进多粒度级联随机森林算法包括多粒子扫描算法和级联随机森林算法,所述多粒子扫描算法对应至少一个滑动窗口;所述采用改进多粒度级联随机森林算法对所述目标样本数据进行训练,得到目标预测模型,包括:采用所述多粒子扫描算法,按照至少一个所述滑动窗口,对所述目标样本数据进行多粒子扫描,得到至少一个中间数据;基于所述池化层,对至少一个所述中间数据进行池化处理,得到待训练数据;采用级联随机森林算法对所述待训练数据进行训练,获取目标预测模型。
- 如权利要求19所述的可读存储介质,其特征在于,所述对至少一个所述中间数据进行池化处理,得到待训练数据,包括:选取相邻的两个中间数据作为一组待处理数据组,以得到所述中间数据对应的至少一组所述待处理数据组;对每组所述待处理数据组进行取平均运算,得到第一数据序列;对每组所述待处理数据组进行最小值运算,得到第二数据序列,所述第二数据列中包括每组所述待处理数据组的两个所述中间数据中的最小值;对每组所述待处理数据组进行最大值运算,得到第三数据序列,所述第三数据列中包括每组所述待处理数据组的两个所述中间数据中的最大值;将所述第一数据序列、所述第二数据序列和所述第三数据序列进行拼接,得到所述待训练数据。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2021506707A JP7165809B2 (ja) | 2019-08-19 | 2019-11-11 | データのインテリジェント分析方法、装置、コンピュータ機器及び記憶媒体 |
| SG11202008324YA SG11202008324YA (en) | 2019-08-19 | 2019-11-11 | Intelligent data analysis method and apparatus, computer device, and storage medium |
| US17/168,925 US20210158973A1 (en) | 2019-08-19 | 2021-02-05 | Intelligent data analysis method and device, computer device, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910763137.5 | 2019-08-19 | ||
| CN201910763137.5A CN110675959B (zh) | 2019-08-19 | 2019-08-19 | 数据智能分析方法、装置、计算机设备及存储介质 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/168,925 Continuation US20210158973A1 (en) | 2019-08-19 | 2021-02-05 | Intelligent data analysis method and device, computer device, and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020215671A1 true WO2020215671A1 (zh) | 2020-10-29 |
Family
ID=69075500
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/116942 Ceased WO2020215671A1 (zh) | 2019-08-19 | 2019-11-11 | 数据智能分析方法、装置、计算机设备及存储介质 |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US20210158973A1 (zh) |
| JP (1) | JP7165809B2 (zh) |
| CN (1) | CN110675959B (zh) |
| SG (1) | SG11202008324YA (zh) |
| WO (1) | WO2020215671A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112395401A (zh) * | 2020-11-17 | 2021-02-23 | 中国平安人寿保险股份有限公司 | 自适应负样本对采样方法、装置、电子设备及存储介质 |
| CN113159181A (zh) * | 2021-04-23 | 2021-07-23 | 湖南大学 | 基于改进的深度森林的工业控制网络异常检测方法和系统 |
| CN113268921A (zh) * | 2021-05-13 | 2021-08-17 | 西安交通大学 | 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质 |
| CN116266211A (zh) * | 2021-12-17 | 2023-06-20 | 中国移动通信有限公司研究院 | 数据关联性分析方法、装置、电子设备及可读存储介质 |
Families Citing this family (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111738286A (zh) * | 2020-03-17 | 2020-10-02 | 北京京东乾石科技有限公司 | 故障判定及其模型训练方法、装置、设备及存储介质 |
| CN111581877B (zh) * | 2020-03-25 | 2025-01-21 | 中国平安人寿保险股份有限公司 | 样本模型训练方法、样本生成方法、装置、设备及介质 |
| CN111986763B (zh) * | 2020-09-03 | 2024-05-14 | 深圳平安智慧医健科技有限公司 | 疾病数据分析方法、装置、电子设备及存储介质 |
| CN112134862B (zh) * | 2020-09-11 | 2023-09-08 | 国网电力科学研究院有限公司 | 基于机器学习的粗细粒度混合网络异常检测方法及装置 |
| CN112434208B (zh) * | 2020-12-03 | 2024-05-07 | 百果园技术(新加坡)有限公司 | 一种孤立森林的训练及其网络爬虫的识别方法与相关装置 |
| CN112579587B (zh) * | 2020-12-29 | 2024-07-02 | 纽扣互联(北京)科技有限公司 | 数据清洗方法及装置、设备和存储介质 |
| CN112862179A (zh) * | 2021-02-03 | 2021-05-28 | 国网山西省电力公司吕梁供电公司 | 一种用能行为的预测方法、装置及计算机设备 |
| CN113672366B (zh) * | 2021-08-05 | 2025-04-04 | Oppo广东移动通信有限公司 | 应用程序管理方法、装置、电子设备及存储介质 |
| CN114358422B (zh) * | 2022-01-04 | 2024-12-27 | 中国工商银行股份有限公司 | 研发进度的异常预测方法及装置、存储介质和电子设备 |
| CN114547970B (zh) * | 2022-01-25 | 2024-02-20 | 中国长江三峡集团有限公司 | 一种水电厂顶盖排水系统异常智能诊断方法 |
| CN114581252B (zh) * | 2022-03-03 | 2024-04-05 | 平安科技(深圳)有限公司 | 目标案件的预测方法和装置、电子设备、存储介质 |
| CN115146696A (zh) * | 2022-04-18 | 2022-10-04 | 华北电力大学 | 一种基于多任务学习的综合能源系统运行状态监测方法 |
| CN115129769B (zh) * | 2022-06-01 | 2024-08-20 | 苏州大学 | 一种居民出行调查扩样方法、装置及存储介质 |
| CN115312193B (zh) * | 2022-08-05 | 2025-05-06 | 天津医科大学总医院 | 医学潜在相关指标风险监控系统、方法、终端及存储介质 |
| CN115547508B (zh) * | 2022-11-29 | 2023-03-21 | 联仁健康医疗大数据科技股份有限公司 | 数据校正方法、装置、电子设备及存储介质 |
| CN117251474A (zh) * | 2022-12-19 | 2023-12-19 | 中国南方电网有限责任公司 | 电力系统碳流计算的数据库构建方法、装置、设备和介质 |
| KR102653187B1 (ko) * | 2023-02-23 | 2024-04-01 | 주식회사 쇼퍼하우스 | 웹크롤링 기반 학습용 데이터 전처리 전자 장치 및 그 방법 |
| CN117009813B (zh) * | 2023-07-05 | 2025-10-31 | 上海幻电信息科技有限公司 | 样本拼接训练方法及装置 |
| CN117786560B (zh) * | 2024-02-28 | 2024-05-07 | 通用电梯股份有限公司 | 一种基于多粒度级联森林的电梯故障分类方法及电子设备 |
| CN119272178B (zh) * | 2024-12-10 | 2025-04-29 | 上海融和元储能源有限公司 | 数据处理方法以及装置 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2017186048A1 (zh) * | 2016-04-27 | 2017-11-02 | 第四范式(北京)技术有限公司 | 展示预测模型的方法、装置及调整预测模型的方法、装置 |
| CN108417274A (zh) * | 2018-03-06 | 2018-08-17 | 东南大学 | 流行病预测方法、系统及设备 |
| CN108647249A (zh) * | 2018-04-18 | 2018-10-12 | 平安科技(深圳)有限公司 | 舆情数据预测方法、装置、终端及存储介质 |
| CN109241987A (zh) * | 2018-06-29 | 2019-01-18 | 南京邮电大学 | 基于加权的深度森林的机器学习方法 |
| CN109656918A (zh) * | 2019-01-04 | 2019-04-19 | 平安科技(深圳)有限公司 | 流行病发病指数的预测方法、装置、设备及可读存储介质 |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7547283B2 (en) * | 2000-11-28 | 2009-06-16 | Physiosonics, Inc. | Methods for determining intracranial pressure non-invasively |
| WO2006087854A1 (ja) * | 2004-11-25 | 2006-08-24 | Sharp Kabushiki Kaisha | 情報分類装置、情報分類方法、情報分類プログラム、情報分類システム |
| US9746985B1 (en) * | 2008-02-25 | 2017-08-29 | Georgetown University | System and method for detecting, collecting, analyzing, and communicating event-related information |
| US8255346B2 (en) * | 2009-11-11 | 2012-08-28 | International Business Machines Corporation | Methods and systems for variable group selection and temporal causal modeling |
| ES2388413B1 (es) * | 2010-07-01 | 2013-08-22 | Telefónica, S.A. | Método para la clasificación de videos. |
| CN105608200A (zh) * | 2015-12-28 | 2016-05-25 | 湖南蚁坊软件有限公司 | 一种网络舆论趋势预测分析方法 |
| KR20180052489A (ko) * | 2016-11-10 | 2018-05-18 | 주식회사 레드아이스 | 사용자 경험분석 및 환경요인에 기초한 크로스보더 전자상거래 상품 추천 방법 |
| JP6736530B2 (ja) * | 2017-09-13 | 2020-08-05 | ヤフー株式会社 | 予測装置、予測方法、及び予測プログラム |
| CN107918772B (zh) * | 2017-12-10 | 2021-04-30 | 北京工业大学 | 基于压缩感知理论和gcForest的目标跟踪方法 |
| CN108389631A (zh) * | 2018-02-07 | 2018-08-10 | 平安科技(深圳)有限公司 | 水痘发病预警方法、服务器及计算机可读存储介质 |
| CN108288502A (zh) * | 2018-04-11 | 2018-07-17 | 平安科技(深圳)有限公司 | 疾病预测方法及装置、计算机装置及可读存储介质 |
| CN108648829A (zh) * | 2018-04-11 | 2018-10-12 | 平安科技(深圳)有限公司 | 疾病预测方法及装置、计算机装置及可读存储介质 |
| CN108921702A (zh) * | 2018-06-04 | 2018-11-30 | 北京至信普林科技有限公司 | 基于大数据的园区招商方法及装置 |
-
2019
- 2019-08-19 CN CN201910763137.5A patent/CN110675959B/zh active Active
- 2019-11-11 JP JP2021506707A patent/JP7165809B2/ja active Active
- 2019-11-11 SG SG11202008324YA patent/SG11202008324YA/en unknown
- 2019-11-11 WO PCT/CN2019/116942 patent/WO2020215671A1/zh not_active Ceased
-
2021
- 2021-02-05 US US17/168,925 patent/US20210158973A1/en not_active Abandoned
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2017186048A1 (zh) * | 2016-04-27 | 2017-11-02 | 第四范式(北京)技术有限公司 | 展示预测模型的方法、装置及调整预测模型的方法、装置 |
| CN108417274A (zh) * | 2018-03-06 | 2018-08-17 | 东南大学 | 流行病预测方法、系统及设备 |
| CN108647249A (zh) * | 2018-04-18 | 2018-10-12 | 平安科技(深圳)有限公司 | 舆情数据预测方法、装置、终端及存储介质 |
| CN109241987A (zh) * | 2018-06-29 | 2019-01-18 | 南京邮电大学 | 基于加权的深度森林的机器学习方法 |
| CN109656918A (zh) * | 2019-01-04 | 2019-04-19 | 平安科技(深圳)有限公司 | 流行病发病指数的预测方法、装置、设备及可读存储介质 |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112395401A (zh) * | 2020-11-17 | 2021-02-23 | 中国平安人寿保险股份有限公司 | 自适应负样本对采样方法、装置、电子设备及存储介质 |
| CN112395401B (zh) * | 2020-11-17 | 2024-06-04 | 中国平安人寿保险股份有限公司 | 自适应负样本对采样方法、装置、电子设备及存储介质 |
| CN113159181A (zh) * | 2021-04-23 | 2021-07-23 | 湖南大学 | 基于改进的深度森林的工业控制网络异常检测方法和系统 |
| CN113159181B (zh) * | 2021-04-23 | 2022-06-10 | 湖南大学 | 基于改进的深度森林的工业控制系统异常检测方法和系统 |
| CN113268921A (zh) * | 2021-05-13 | 2021-08-17 | 西安交通大学 | 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质 |
| CN113268921B (zh) * | 2021-05-13 | 2022-12-09 | 西安交通大学 | 凝汽器清洁系数预估方法、系统、电子设备及可读存储介质 |
| CN116266211A (zh) * | 2021-12-17 | 2023-06-20 | 中国移动通信有限公司研究院 | 数据关联性分析方法、装置、电子设备及可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2021532501A (ja) | 2021-11-25 |
| CN110675959A (zh) | 2020-01-10 |
| SG11202008324YA (en) | 2020-11-27 |
| JP7165809B2 (ja) | 2022-11-04 |
| US20210158973A1 (en) | 2021-05-27 |
| CN110675959B (zh) | 2023-07-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020215671A1 (zh) | 数据智能分析方法、装置、计算机设备及存储介质 | |
| CN112365171B (zh) | 基于知识图谱的风险预测方法、装置、设备及存储介质 | |
| US9442929B2 (en) | Determining documents that match a query | |
| TWI740891B (zh) | 利用訓練資料訓練模型的方法和訓練系統 | |
| US20200175314A1 (en) | Predictive data analytics with automatic feature extraction | |
| WO2022217713A1 (zh) | 症候群监测预警方法、装置、计算机设备及存储介质 | |
| WO2021012790A1 (zh) | 页面数据生成方法、装置、计算机设备及存储介质 | |
| US11921681B2 (en) | Machine learning techniques for predictive structural analysis | |
| WO2021164171A1 (zh) | 知识库中数据处理方法、装置、计算机设备和存储介质 | |
| US20230394352A1 (en) | Efficient multilabel classification by chaining ordered classifiers and optimizing on uncorrelated labels | |
| WO2021114613A1 (zh) | 基于人工智能的故障节点识别方法、装置、设备和介质 | |
| CN114547257A (zh) | 类案匹配方法、装置、计算机设备及存储介质 | |
| WO2021218037A1 (zh) | 目标检测方法、装置、计算机设备和存储介质 | |
| EP4318265A1 (en) | Insight mining using machine learning | |
| CN118734247B (zh) | 基于moe架构的智能城市中枢数据融合计算模型训练方法、预警方法及设备 | |
| CN112699668A (zh) | 一种化学信息抽取模型的训练方法、抽取方法、装置、设备及存储介质 | |
| CN115827797A (zh) | 一种基于大数据的环境数据分析整合方法及系统 | |
| WO2019080419A1 (zh) | 标准知识库的构建方法、电子装置及存储介质 | |
| CN116434973A (zh) | 基于人工智能的传染病预警方法、装置、设备及介质 | |
| US12536430B2 (en) | Machine learning techniques for efficient data pattern recognition across structured data objects | |
| WO2023040145A1 (zh) | 基于人工智能的文本分类方法、装置、电子设备及介质 | |
| US11907191B2 (en) | Content based log retrieval by using embedding feature extraction | |
| WO2024199121A1 (zh) | 网络模型的结构化稀疏、图像分类方法及相关装置 | |
| HK40017556B (zh) | 数据智能分析方法、装置、计算机设备及存储介质 | |
| HK40017556A (zh) | 数据智能分析方法、装置、计算机设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19926035 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021506707 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19926035 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19926035 Country of ref document: EP Kind code of ref document: A1 |

