WO2019218482A1 - 基于大数据的人群筛选方法、装置、终端设备及可读存储介质 - Google Patents
基于大数据的人群筛选方法、装置、终端设备及可读存储介质 Download PDFInfo
- Publication number
- WO2019218482A1 WO2019218482A1 PCT/CN2018/097561 CN2018097561W WO2019218482A1 WO 2019218482 A1 WO2019218482 A1 WO 2019218482A1 CN 2018097561 W CN2018097561 W CN 2018097561W WO 2019218482 A1 WO2019218482 A1 WO 2019218482A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- feature value
- sample
- feature
- output
- screening
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/211—Selection of the most significant subset of features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
Definitions
- the present application belongs to the field of data processing technologies, and in particular, to a big data-based crowd screening method, apparatus, terminal device, and computer readable storage medium.
- the related features can be found according to statistics, so that the population selected according to the relevant features is less accurate for the target population. For example, when screening people with COPD, mainly based on the statistical results of the patient population, if there is a large number of people with COPD near a certain area, or if they are located in a certain age group. If there are more people with chronic obstructive pulmonary disease, the region or the age interval is selected as the relevant feature to select the target population. However, in fact, there are many related features of chronic obstructive pulmonary disease, so the actual target population should not be limited to the region or the age group. Screening. In summary, the existing population screening methods can be based on fewer relevant features and low screening accuracy.
- the embodiments of the present application provide a method, a device, a terminal device, and a computer readable storage medium based on big data, so as to solve the problem that the related features can be based on screening in the prior art, and the accuracy of screening is improved. Low problem.
- a first aspect of the embodiments of the present application provides a population data screening method based on big data, including:
- sample information including a sample environmental feature and a sample feature value, wherein the sample feature value is used to describe an individual state of the corresponding sample individual in the sample population;
- a second aspect of the embodiments of the present application provides a big data-based crowd screening apparatus, which may include means for implementing the steps of the above-described big data-based crowd screening method.
- a third aspect of the embodiments of the present application provides a terminal device, including a memory and a processor, where the computer stores computer readable instructions executable on the processor, the processor executing the computer
- the steps of the above-described big data-based crowd screening method are implemented when the instruction is read.
- a fourth aspect of embodiments of the present application provides a computer readable storage medium storing computer readable instructions that, when executed by a processor, implement the above-described big data-based crowd The steps of the screening method.
- the embodiment of the present application obtains a plurality of sample information of a sample population, where each sample information includes a sample environment feature and a sample feature value, and the sample feature value is used to describe an individual state of the sample individual corresponding to the sample information to which the sample feature value belongs.
- the plurality of sample information is matched with the preset pre-processing model, and the fitted pre-processing model is output as a screening model, and then the screening model is used to screen the population to be screened, and the population to be screened is obtained.
- the plurality of individual features are input to the screening model, and the output feature value set corresponding to the plurality of individual features outputted by the screening model is obtained, the feature value set includes a plurality of output feature values, and finally the output features are output.
- the output feature value that meets the preset condition in the value set is added to the target feature value set, and the target population corresponding to the target feature value set is determined from the to-be-selected population.
- the embodiment of the present application trains the screening model by using multiple sample information. Therefore, the screening of the population to be screened according to a plurality of characteristics can be improved. The accuracy of population screening.
- FIG. 1 is a flowchart showing an implementation of a method for screening a population based on big data in the first embodiment of the present application
- FIG. 2 is a flowchart of implementing a population data screening method based on big data in Embodiment 2 of the present application;
- Embodiment 3 is a flowchart of implementing a population data screening method based on big data in Embodiment 3 of the present application;
- FIG. 4 is a flowchart showing an implementation of a method for screening a population based on big data in Embodiment 4 of the present application;
- FIG. 5 is a flowchart of implementing a population data screening method based on big data in Embodiment 5 of the present application.
- FIG. 6 is a structural block diagram of a big data-based crowd screening device in Embodiment 6 of the present application.
- FIG. 7 is a schematic diagram of a terminal device in Embodiment 7 of the present application.
- FIG. 1 is a flowchart of implementing a big data-based crowd screening method according to an embodiment of the present application. As shown in Figure 1, the population screening method includes the following steps:
- S101 Acquire a plurality of sample information of a sample population, where the sample information includes a sample environment feature and a sample feature value, where the sample feature value is used to describe an individual state of the sample object in the sample population corresponding to the sample individual.
- the plurality of sample information of the plurality of sample individuals in the sample population is first acquired, and one sample information corresponds to one sample individual.
- the sample population is determined according to the sample conditions.
- the sample conditions include the sample source and sample selection factors.
- the sample source determines the source of the sample population.
- the sample source is the individual file of a certain city
- the sample selection factor determines the constituent sample selected from the source.
- the sample individuals in the population, such as the sample selection factor are men over the age of 50, so the sample condition consisting of the sample source and the sample selection factor is to screen out males over the age of 50 in the individual files of a certain city. Multiple sample individuals constitute a sample population.
- sample conditions are not limited to the above examples, and may be determined according to actual application scenarios.
- the number of sample information obtained from the sample population may be set, so that the number of sample information is at a preset order of magnitude, and the order of magnitude may be based on accuracy. Ask for free settings.
- the order of magnitude can be one thousand.
- the sample information in the sample population includes a sample environment feature and a sample feature value, wherein the sample environment feature is used to indicate that the sample information corresponds to an environmental state of the sample individual, and may include one or more sub-features, such as the sample environment feature may include the sample individual Age, family drinking water type, marital status, occupation, blood pressure status, respiratory rate, diabetes status, drinking status, smoking status, etc.
- the sub-characteristics included in the sample environmental characteristics can also be determined according to the actual application scenario.
- the sample feature value in the sample information is used to indicate an individual state of the sample individual corresponding to the sample information, wherein the individual state is associated with the purpose of screening the target population in the embodiment of the present application, for example, if an anxiety disorder needs to be screened out
- the range of sample eigenvalues may be limited to two integers of 0 and 1, and the sample eigenvalue of 1 indicates that the corresponding sample individual has been subjected to anxiety. Influence, the sample eigenvalue of 0 indicates that the corresponding sample individual is not affected by anxiety.
- a plurality of blood information samples are obtained from a health file or a specific database.
- health records or feature databases such as hospital databases store different kinds of sample individuals, the stored information is of a high order of magnitude, and the sample environment features contain more complete features, the sample feature values are predetermined, and the accuracy is high. Therefore, the sample population can be determined directly from the health file or a specific database, and then multiple sample information can be obtained based on the sample population.
- S102 Fit the plurality of sample information with a preset pre-processing model, and use the fitted pre-processing model as a screening model.
- sample environment feature under the sample information may contain multiple sub-features, and the individual state is related to multiple sub-features of the corresponding sample individual, that is, multiple sub-features may affect the generation of the sample feature values. For example, if an adult male represented by a sample individual is married, has three children, has a high frequency of occupational changes, and is in a state of drinking and smoking, the probability that the sample individual has an anxiety disorder, that is, the sample characteristic value The probability of being 1 is greater.
- sample environment characteristics of the sample individual include marital status, number of children, occupation, drinking status, and smoking status associated with sample characteristic values of the sample individual.
- the change rule of the individual state is affected by the plurality of sub-features, so in the embodiment of the present application, the sample information and the sample feature value have been determined to be a plurality of sample information and a preset pre-processing model.
- the fitting is performed, the pre-processing model is continuously trained during the fitting process, and finally the fitted pre-processing model is output as the screening model.
- S103 Acquire a plurality of individual features of the to-be-screened population, and input the plurality of individual features into the screening model to obtain an output feature value set corresponding to the plurality of individual features, where the output feature value set includes multiple Output feature values.
- the population to be screened is analyzed.
- multiple individual characteristics of the population to be screened are obtained, and the individual characteristics correspond to the individuals to be screened in the population to be screened.
- the individual features are in the same format as the sample environment features of the sample individuals, ie, the types of sub-features are the same.
- the individual features are processed, and only the child features of the same type as the sample environment features are retained.
- individual characteristics include name, age, gender, drinking status, and smoking status
- sample environmental characteristics include age, drinking status, and smoking status.
- Individual characteristics are processed prior to inputting individual characteristics into the screening model, and only individuals are obtained.
- the sub-characteristics of the characteristics of age, drinking status and smoking status reduce the complexity of subsequent calculations.
- the value of a sub-feature in the processed individual feature is null, the average value or the preset value of the sub-feature in the sample environment feature of the plurality of sample individuals is taken as the value of the sub-feature in the individual feature to prevent Subsequent calculation errors have improved the stability of the calculation.
- the plurality of individual features are input to the screening model, and each individual feature can obtain a corresponding output feature value after being calculated by the screening model, and the output feature value is in the same format as the sample feature value of the sample information, but It is not limited to the range of values of sample feature values. Acquiring a plurality of output feature value output results corresponding to the plurality of individual features to generate an output feature value set.
- S104 Add an output feature value that meets a preset condition in the output feature value set to the target feature value set, and determine a target population corresponding to the target feature value set from the to-be-selected population.
- an output feature value that satisfies a preset condition is obtained from the output feature value set, and the output feature value is added to the target feature value set.
- the preset condition may be configured to add an output feature value that is at a preset ratio in the output feature value set to the target feature value set, such as sorting the plurality of output feature values in the output feature value set. And after the sorting is completed, add the output feature value of the top 10% to the target feature value set.
- the preset condition may also be set according to the feature threshold, and the method for determining the feature threshold is specifically described later. Since the output feature value corresponds to the individual to be screened in the population to be selected, the target population in the population to be screened can be determined according to the target feature value set, and the screening of the population to be screened is completed.
- FIG. 1 shows that, in the embodiment of the present application, a plurality of sample information corresponding to a plurality of sample individuals in a sample population is first acquired, wherein the sample information includes a sample environment feature and a sample feature value, and the sample feature value is used for Describe the individual state of the sample individual, fit the multiple sample information with the preset pre-processing model, and output the fitted pre-processing model as a screening model to screen the population to be screened and obtain multiple populations to be screened.
- a plurality of individual features of the individual to be screened inputting a plurality of individual features into the screening model, obtaining an output feature value set formed by combining the plurality of output feature values, and finally adding the output feature values satisfying the preset condition in the output feature value set
- a plurality of individuals to be selected corresponding to the target feature value set are determined from the to-be-selected population, and the target population is automatically filtered, and the accuracy of screening the target population is improved.
- FIG. 2 is a flowchart of implementing a big data-based crowd screening method according to Embodiment 2 of the present application.
- this embodiment refines S102 to obtain S201-S202, which are as follows:
- S201 input the plurality of sample information to the preprocessing model to train the preprocessing model, wherein a sample environment feature of the sample information is used as an input parameter of the preprocessing model, and the sample is The sample feature value of the information is used as a reference parameter of the pre-processing model.
- the sample environment feature in the sample information is used as an input parameter of the preprocessing model, and the sample feature value in the sample information is used as a reference parameter of the preprocessing model, specifically, based on
- the sample information constructs a sample information set (Characters environ1 , Value environ1 ), (Characters environ2 , Value environ2 )... (Characters environn , Value environn ), wherein Characters environi represents the sample environment characteristic of the i-th sample information, It may include multiple sub-features, Value environi represents the sample feature value of the i-th sample information, and n represents the total number of pieces of sample information obtained from the sample population.
- the sample information set is input to the pre-processing model to train the pre-processing model.
- the calculation formula of the input parameter for the pre-processing model is:
- the f() in the formula indicates a function that exists in the function space.
- the function space refers to a set of functions of a given kind from one set to another, that is, the f() function is initially in an unknown state, and K means pre- There are K above f() functions in the processing model, and all the results calculated by the f() function need to be accumulated to obtain the final predicted value.
- the f() function is learned by the sequential learning method, so that the finally obtained K f() functions are maximally matched.
- Data from multiple sample information For example, on the basis that the input parameter is Characters environi , the prediction of the predicted value of the t-round is performed, and when the prediction of the prediction of the t-th round is performed, the prediction result of the prediction value of the t- 1th round is retained, that is, training according to the order Preprocess the model to make predictive values
- the difference between the value and the value of the environi is gradually reduced, as follows:
- Value environi concentrated sample information is input parameter Characters environi corresponding reference parameters, that is, the sample value of the sample feature information, in the present embodiment, the application as a parameter of the optimization function.
- the ⁇ (f t ) in the optimization function formula is a regular term, and D is a constant term.
- the regular term controls the training degree of the optimization function to prevent over-fitting of the sample information set and the pre-processing model;
- the constant term is a constant, and the constant is set.
- the term is to limit the range of values of the optimization function.
- the process of optimizing the optimization function is a process of determining the appropriate f() function to minimize the value of the above error function.
- the expanded optimization function is:
- the output value obtained by the simplification function depends on the values of g i and h i , so that the appropriate f() function can be quickly determined, which improves the simplicity of training, in the embodiment of the present application.
- the pre-processing model is trained by the above-described order learning method and the optimization function (the simplification function is also possible).
- the trained pre-processing model is output as a screening model (mainly the trained f() function).
- the screening of the population to be screened is required, the individual characteristics of the individual to be screened, Characters environx, are input into the screening model, and the optimized calculation formula in the screening model can be obtained. Calculated to obtain predicted values
- the sample environment feature in the sample information is used as an input parameter of the preprocessing model, and the sample feature value of the sample information is used as a reference parameter of the preprocessing model, thereby
- the sample information is input to the pre-processing model to train the pre-processing model, and finally the trained pre-processing model is output as the screening model, which improves the fitting degree of the screening model and the plurality of sample information, and improves the screening model.
- the accuracy of crowd screening is used as an input parameter of the preprocessing model, and the sample feature value of the sample information is used as a reference parameter of the preprocessing model, thereby
- the sample information is input to the pre-processing model to train the pre-processing model, and finally the trained pre-processing model is output as the screening model, which improves the fitting degree of the screening model and the plurality of sample information, and improves the screening model.
- the accuracy of crowd screening is used as an input parameter of the preprocessing model, and the sample feature value of the sample information is used as a reference
- FIG. 3 is a flowchart of implementing a big data-based crowd screening method according to Embodiment 3 of the present application.
- the embodiment obtains S301 to S302 after S104 is refined, and the details are as follows:
- S301 Acquire a feature threshold, where the feature threshold is used to determine whether the output feature value meets the preset condition.
- the screening model after the screening model is determined, it is required to acquire a plurality of individual features of the population to be screened, and automatically input a plurality of individual features into the screening model to obtain an output feature value set including a plurality of output feature values.
- the feature threshold is first acquired, and the output feature value that satisfies the preset condition is an output feature value that is greater than or equal to the feature threshold.
- FIG. 4 is a flowchart of implementing a big data-based crowd screening method according to Embodiment 4 of the present application.
- the S301 is refined to obtain S401 to S402, which are as follows:
- S401 Perform big data analysis on the plurality of sample information, and determine that the sample feature value is a quantity ratio of the sample information of the first feature value occupied by the plurality of sample information.
- the value range of the sample feature values may be different.
- the value of the sample feature value is the first feature value or the second feature value, and the first feature value is greater than the first feature value. The case of the two eigenvalues will be described.
- the sample feature value takes the value of the first feature value to indicate that the corresponding sample individual is affected by the anxiety disorder
- the sample feature value takes the second feature value to indicate that the corresponding sample individual is not affected by the anxiety disorder
- the screening purpose is To select the target population affected by the anxiety disorder in the population to be screened, firstly, the big data analysis is performed on the plurality of sample information, and the first sample quantity whose sample feature value is the first eigenvalue is extracted, thereby obtaining the same The proportion of this quantity in all sample information.
- S402 Calculate the feature threshold according to the quantity ratio, the first feature value, and the second feature value.
- the value range of the sample feature value is limited to the first feature value and the second feature value, calculating a difference between the first feature value and the second feature value, and subtracting the difference and the quantity from the first feature value
- the foregoing calculation method is only applicable to the case where the value of the sample feature value is the first feature value or the second feature value, and the first feature value is greater than the second feature value, and the other possible sample feature values may be in the above calculation method.
- the extension of the present application is not described in detail in the embodiment of the present application.
- FIG. 5 is a flowchart of implementing a big data-based crowd screening method according to Embodiment 5 of the present application.
- S501 to S503 are obtained, which are as follows:
- S501 Input a sample environment feature of the plurality of sample information to the screening model, and obtain a plurality of result feature values corresponding to sample environment features of the plurality of sample information output by the screening model.
- a plurality of sample information is input to the screening model.
- the sample environment feature of the plurality of sample information is input as an input parameter to the screening model, but the sample feature value of the plurality of sample information is not used as the reference parameter of the screening model, but the plurality of sample information is directly used by the calculation formula of the screening model.
- the sample environment features are calculated to obtain a plurality of output parameters, that is, a plurality of result feature values corresponding to the sample environment features.
- the sample feature information calculated by the screening model differs from the original sample feature value of the sample information.
- S502 Sort the plurality of result feature values to generate a sequence of result feature values.
- multiple result feature values are sorted according to the numerical value to generate a numerical sequence, and the first column of the numerical sequence is the result feature value with the largest numerical value. It is worth mentioning that if the first result feature value of the plurality of result feature values is the same as the second result feature value, the feature value is generated according to the input order of the sample information corresponding to the first result feature value and the second result feature value. sequence.
- a sequence writing mechanism is set, and when a certain feature eigenvalue needs to be written into the eigenvalue sequence, it is determined whether there is an existing eigenvalue in the eigenvalue sequence that is the same as the eigenvalue of the result, if the result feature does not exist If the value of the existing result is the same, the minimum value of the feature value sequence is found to be greater than the value of the existing feature value, and the value of the existing feature value is written after the existing feature value.
- the result feature value is written at a position after the existing result feature value, if there are multiple identical feature values If there is a result feature value, the existing result feature value at the end of the plurality of existing result feature values is searched, and the result feature value is written at the position after the existing result feature value.
- S503 Acquire a preset screening ratio, and find the result feature value corresponding to the screening ratio in the sequence of the result feature values, and output the value as the feature threshold.
- the preset screening ratio is obtained, and the screening position corresponding to the screening ratio is searched in the eigenvalue sequence, and the result eigenvalue at the screening position is output as the feature threshold. For example, if 300 result feature values have been written in the sequence of feature values, and the filter ratio is 10%, then the filter position is the 30th bit, and the value of the feature value of the 30th bit in the sequence of feature values is extracted as a feature. Threshold.
- the screening ratio may be determined according to a big data analysis method, for example, a plurality of national sample information may be counted nationwide, and each national sample information includes a sample environmental feature and a sample feature value, and first, the sample feature value is selected and filtered.
- the number of samples corresponding to the target is used to determine the proportion of the number of screening samples in all national sample information.
- the ratio is used as the screening ratio.
- the general applicability of the generated feature thresholds is higher through the big data analysis method.
- the screening ratio can also be artificially preset.
- step S302 an output feature value of the output feature value set that is greater than or equal to the feature threshold is extracted, and the extracted output feature value is added to the target feature value set.
- the output feature value greater than or equal to the feature threshold is extracted from the output feature value set, and the extracted output feature value corresponds to the selected target individual, so the extracted output feature value is added to the target feature value. Aggregating, so as to subsequently determine a target population corresponding to the target feature value from the population to be screened.
- the feature threshold is obtained by using different methods, and the output feature value greater than or equal to the feature threshold is extracted from the output feature value set, and the extracted output is extracted.
- the feature value is added to the target feature value set, and the simplicity and efficiency of the target feature value set generation are improved by setting the feature threshold.
- FIG. 6 is a structural block diagram of a big data-based crowd screening device provided by an embodiment of the present application.
- the crowd Screening devices include:
- a first acquiring unit 61 configured to acquire a plurality of sample information of a sample population, where the sample information includes a sample environment feature and a sample feature value, where the sample feature value is used to describe a sample corresponding to the sample information in the sample population Individual state of the individual;
- a fitting unit 62 configured to fit the plurality of sample information with a preset pre-processing model, and use the fitted pre-processing model as a screening model;
- a second acquiring unit 63 configured to acquire a plurality of individual features of the to-be-screened population, and input the plurality of individual features into the screening model to obtain an output feature value set corresponding to the plurality of individual features,
- the output feature value set includes a plurality of output feature values
- the target determining unit 64 is configured to add an output feature value that meets a preset condition in the output feature value set to the target feature value set, and determine a target crowd corresponding to the target feature value set from the to-be-selected population .
- the fitting unit 62 includes:
- An input unit configured to input the plurality of sample information to the pre-processing model to train the pre-processing model, wherein a sample environment feature of the sample information is used as an input parameter of the pre-processing model, a sample feature value of the sample information as a reference parameter of the preprocessing model;
- an output unit configured to output the trained pre-processing model as the screening model.
- the target determining unit 64 includes:
- a threshold acquiring unit configured to acquire a feature threshold, where the feature threshold is used to determine whether the output feature value meets the preset condition
- an extracting unit configured to extract an output feature value that is greater than or equal to the feature threshold in the output feature value set, and add the extracted output feature value to the target feature value set.
- the sample feature value is a first feature value or a second feature value
- the threshold acquiring unit includes:
- An analyzing unit configured to perform big data analysis on the plurality of sample information, and determine a ratio of a quantity in which the sample feature value is a value of the sample information of the first feature value occupied by the plurality of sample information;
- a calculating unit configured to calculate the feature threshold according to the quantity ratio and the first feature value and the second feature value.
- the threshold obtaining unit includes:
- a feature input unit configured to input a sample environment feature of the plurality of sample information to the screening model, and acquire a plurality of result feature values corresponding to sample environment features of the plurality of sample information output by the screening model ;
- a sorting unit configured to sort the plurality of result feature values to generate a sequence of the result feature values
- the threshold output unit is configured to obtain a preset screening ratio, and find the result feature value corresponding to the screening ratio in the sequence of the result feature values, and output the value as the feature threshold.
- FIG. 7 is a schematic diagram of a terminal device according to an embodiment of the present application.
- the terminal device 7 of this embodiment includes a processor 70 and a memory 71 in which computer readable instructions 72 executable on the processor 70 are stored, such as based on big data. Crowd screening program.
- the processor 70 executes the computer readable instructions 72 to implement the steps in the various embodiments of the various big data based crowd screening methods described above, such as steps S101 through S104 shown in FIG.
- the processor 70 when executing the computer readable instructions 72, implements the functions of the various units in the population screening apparatus embodiment described above, such as the functions of the units 61-64 shown in FIG.
- the computer readable instructions 72 may be partitioned into one or more modules/units that are stored in the memory 71 and executed by the processor 70, To complete this application.
- the one or more modules/units may be a series of computer readable instruction segments capable of performing a particular function for describing the execution of the computer readable instructions 72 in the terminal device 7.
- the computer readable instructions 72 can be segmented into a first acquisition unit, a fitting unit, a second acquisition unit, and a target determination unit, each unit having a specific function as described above.
- the terminal device may include, but is not limited to, a processor 70 and a memory 71. It will be understood by those skilled in the art that FIG. 7 is only an example of the terminal device 7, and does not constitute a limitation of the terminal device 7, and may include more or less components than those illustrated, or combine some components or different components.
- the terminal device may further include an input/output device, a network access device, a bus, and the like.
- the processor 70 may be a central processing unit (CPU), or may be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), Field-Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
- the general purpose processor may be a microprocessor or the processor or any conventional processor or the like.
- the memory 71 may be an internal storage unit of the terminal device 7, such as a hard disk or a memory of the terminal device 7.
- the memory 71 may also be an external storage device of the terminal device 7, for example, a plug-in hard disk provided on the terminal device 7, a smart memory card (SMC), and a secure digital (SD). Card, flash card, etc. Further, the memory 71 may also include both an internal storage unit of the terminal device 7 and an external storage device.
- the memory 71 is configured to store the computer readable instructions and other programs and data required by the terminal device.
- the memory 71 can also be used to temporarily store data that has been output or is about to be output.
- each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
- the integrated unit if implemented in the form of a software functional unit and sold or used as a standalone product, may be stored in a computer readable storage medium.
- a computer readable storage medium A number of instructions are included to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present application.
- the foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and the like, which can store program codes. .
Landscapes
- Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Biology (AREA)
- Evolutionary Computation (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Image Analysis (AREA)
Abstract
基于大数据的人群筛选方法、终端设备及计算机可读存储介质,包括:获取样本人群的多个样本信息,样本信息包括样本环境特征和样本特征值;将多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型(S102);获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值(S103);将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群(S104)。以上方法实现了根据多个特征对待筛选人群进行筛选,提升了人群筛选的准确率。
Description
本申请要求于2018年05月14日提交中国专利局、申请号为201810455659.4、发明名称为“基于大数据的人群筛选方法及终端设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请属于数据处理技术领域,尤其涉及基于大数据的人群筛选方法、装置、终端设备及计算机可读存储介质。
在现实生活中,往往存在从基数庞大的人群中筛选出部分目标人群的需求,筛选一般以某个特征为依据,比如说以某个县所有人为基础,筛选出年龄超过五十岁的目标人群。而统计学是关于认识客观现象总体数量特征和数量关系的科学,在需要筛选出因某个状态改变而产生特殊性的目标人群时,需要运用统计学,查找出与该状态改变关联的相关特征,从而根据相关特征筛选出目标人群。
但是,在现有的人群筛选方法中,根据统计学能够查找出的相关特征较少,从而根据相关特征筛选出的人群为目标人群的准确度较低。比如在对患有慢阻肺的人群进行筛选时,主要根据已患病人群的统计结果,如果统计出某一个区域附近患有慢阻肺的人数较多,或位于某一个年龄区间的患有慢阻肺的人数较多,则将该区域或该年龄区间作为相关特征筛选出目标人群,但实际上慢阻肺的相关特征较多,故实际的目标人群应当不限于该区域或该年龄区间进行筛选。综上,现有的人群筛选方法可依据的相关特征少,筛选准确度低。
有鉴于此,本申请实施例提供了基于大数据的人群筛选方法、装置、终端设备及计算机可读存储介质,以解决现有技术中进行人群筛选时可依据的相关特征少,筛选的准确度低的问题。
本申请实施例的第一方面提供了一种基于大数据的人群筛选方法,包括:
获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;
将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;
获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;
将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
本申请实施例的第二方面提供了一种基于大数据的人群筛选装置,可以包括用于实现上述基于大数据的人群筛选方法的步骤的单元。
本申请实施例的第三方面提供了一种终端设备,包括存储器以及处理器,所述存储器中存储有可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现上述基于大数据的人群筛选方法的步骤。
本申请实施例的第四方面提供了一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,所述计算机可读指令被处理器执行时实现上述基于大数据的人群筛选方法的步骤。
本申请实施例通过获取样本人群的多个样本信息,其中,每个样本信息包括样本环境特征和样本特征值,样本特征值用于描述样本特征值所属的样本信息对应的样本个体的个体状态,获取完毕后,将多个样本信息与预设的预处理模型进行拟合,并将拟合完成的预处理模型作为筛选模型进行输出,然后通过筛选模型进行对待筛选人群的筛选,获取待筛选人群的多个个体特征,将多个个体特征输入至筛选模型,得到经筛选模型计算后输出的与多个个体特征对应的输出特征值集合,特征值集合包括多个输出特征值,最后将输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,从待筛选人群中确定出与目标特征值集合对应的目标人群,本申请实施例通过基于多个样本信息训练出筛选模型,从而可根据多个特征对待筛选人群进行筛选,提升了人群筛选的准确率。
图1是本申请实施例一中基于大数据的人群筛选方法的实现流程图;
图2是本申请实施例二中基于大数据的人群筛选方法的实现流程图;
图3是本申请实施例三中基于大数据的人群筛选方法的实现流程图;
图4是本申请实施例四中基于大数据的人群筛选方法的实现流程图;
图5是本申请实施例五中基于大数据的人群筛选方法的实现流程图;
图6是本申请实施例六中基于大数据的人群筛选装置的结构框图;
图7是本申请实施例七中终端设备的示意图。
为了对本申请的技术特征、目的和效果有更加清楚的理解,现对照附图详细说明本申请的具体实施方式。
请参阅图1,图1是本申请实施例提供的一种基于大数据的人群筛选方法的实现流程图。如图1所示,该人群筛选方法包括以下步骤:
S101:获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态。
在本申请实施例中,为了筛选出需求的目标人群,首先获取样本人群中多个样本个体的多个样本信息,其中一个样本信息对应一个样本个体。样本人群根据样本条件确定,样本条件包括样本来源和样本选择因素等,样本来源决定样本人群选取的来源地,比如样本来源为某个市的个体档案,样本选择因素决定从来源地选择的构成样本人群的样本个体,比如样本选择因素为年龄在五十岁以上的男性,故样本来源和样本选择因素组成的样本条件即是在某个市的个体档案中筛选出年龄在五十岁以上的男性的多个样本个体组成样本人群。当然,样本条件并不限于上述例子,即可根据实际应用场景进行确定。为了提升对目标人群筛选的准确度,故在样本人群对应的样本条件确定后,可设置从样本人群中获取的样本信息的数量,使样本信息的数量处于预设的数量级,数量级可根据准确度要求自由设置。比如数量级可以为一千个。样本人群中的样本信息包括样本环境特征和样本特征值,其中,样本环境特征用于指示该样本信息对应样本个体的环境状态,可能包含一个或多个子特征,比如样本环境特征可以包括样本个体的年龄、家庭饮水类型、婚姻状况、职业、血压状况、呼吸频率、糖尿病状况、饮酒状况和吸烟状况等,同样地,样本环境特征包含的子特征也可以根据实际应用场景进行确定。样本信息中的样本特征值用于指示与样本信息对应的样本个体的个体状态,其中,个体状态与本申请实施例中筛选目标人群的目的相关联,举例来说,若需要筛选出受焦虑症影响的目标人群,则个体状态与个体是否患有焦虑症相关,则可将样本特征值的取值范围限定为0和1两个整数,样本特征值为1指示对应的样本个体已受焦虑症影响,样本特征值为0指示对应的样本个体未受焦虑症影响。
优选地,从健康档案或特定数据库获取多个血液信息样本。通常来说,健康档案或特点数据库如医院数据库中存放有不同种类的样本个体,存储信息的数量级高,并且样本环境特征包含的特征较为齐全,样本特征值已预先确定,并且准确度较高,故可直接从健康档案或特定数据库中确定样本人群,再基于样本人群进行多个样本信息的获取。
S102:将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模 型作为筛选模型。
由于样本信息下的样本环境特征可能包含多个子特征,而个体状态与对应的样本个体的多个子特征有关,即多个子特征会对样本特征值的生成产生影响。举例来说,若某样本个体代表的一个成年男性处于已婚状态,育有三个子女,职业变更频率高,并且处于饮酒和吸烟的状态,则该样本个体患焦虑症的概率,即样本特征值为1的概率较大。在上述例子中,样本个体的样本环境特征包括的婚姻状态、子女个数、职业、饮酒状况和吸烟状况与该样本个体的样本特征值相关联。但是,在普遍情况下,无法获知个体状态受多个子特征影响的变化规律,故在本申请实施例中,将样本环境特征和样本特征值已确定的多个样本信息与预设的预处理模型进行拟合,在拟合过程中不断训练预处理模型,最后将拟合完成的预处理模型输出为筛选模型。
S103:获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值。
由多个样本信息生成筛选模型后,对待筛选人群进行分析,首先,获取待筛选人群的多个个体特征,个体特征与待筛选人群中的待筛选个体对应。一般地,个体特征与样本个体的样本环境特征的格式相同,即包含子特征的种类相同。可选地,在获取到个体特征后,对个体特征进行处理,仅保留个体特征中与样本环境特征相同种类的子特征。比如个体特征包括姓名、年龄、性别、饮酒状况和吸烟状况,而样本环境特征包括年龄、饮酒状况和吸烟状况,则在将个体特征输入至筛选模型之前,先对个体特征进行处理,只获取个体特征中种类为年龄、饮酒状况和吸烟状况的子特征,降低了后续计算的复杂度。另外,若处理后的个体特征中某个子特征的值为空,则取多个样本个体的样本环境特征中该子特征的平均值或预设的值作为个体特征中该子特征的值,防止后续计算出错,提升了计算的稳定性。
确定多个个体特征后,将多个个体特征输入至筛选模型,每个个体特征在经筛选模型计算后可得到对应的输出特征值,输出特征值与样本信息的样本特征值的格式相同,但并不限定于样本特征值的取值范围。获取多个个体特征对应的多个输出特征值输出结果,生成输出特征值集合。
S104:将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
输出特征值集合确定后,从输出特征值集合中获取满足预设条件的输出特征值,并将该输出特征值添加至目标特征值集合。在本申请实施例中,预设条件可设定将输出特征值集合中处于预设比例的输出特征值添加至目标特征值集合,比如将输出特征值集合中的多个输出特征值进行大小排序,并在排序完成后,将比例处于前10%的输出特征值添加进目标特征值 集合。此外,预设条件还可根据特征阈值进行设定,特征阈值的确定方法在后文进行具体阐述。由于输出特征值对应待筛选人群中的待筛选个体,故根据目标特征值集合可确定出待筛选人群中的目标人群,完成对待筛选人群的筛选。
通过图1所示实施例可知,在本申请实施例中,首先获取样本人群中多个样本个体对应的多个样本信息,其中,样本信息包括样本环境特征和样本特征值,样本特征值用于描述样本个体的个体状态,将多个样本信息与预设的预处理模型进行拟合,并将拟合完成的预处理模型输出为筛选模型,进行对待筛选人群的筛选,获取待筛选人群多个待筛选个体的多个个体特征,将多个个体特征输入至筛选模型,得到由多个输出特征值组合成的输出特征值集合,最后将输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,从待筛选人群中确定与目标特征值集合对应的多个待筛选个体,作为目标人群,实现了自动筛选,并且提升了对目标人群筛选的准确度。
请参阅图2,图2是本申请实施例二提供的一种基于大数据的人群筛选方法的实现流程图。相对于图1对应的实施例,本实施例对S102进行细化后得到S201~S202,详述如下:
S201:将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数。
在将多个样本信息输入至预处理模型时,将样本信息中的样本环境特征作为预处理模型的输入参数,将样本信息中的样本特征值作为预处理模型的参照参数,具体地,基于多个样本信息构建样本信息集,为(Characters
environ1,Value
environ1),(Characters
environ2,Value
environ2)……(Characters
environn,Value
environn),其中,Characters
environi代表第i个样本信息的样本环境特征,在其下可能包括多个子特征,Value
environi代表第i个样本信息的样本特征值,n代表从样本人群获取的多个样本信息的总个数。将样本信息集输入至预处理模型,以训练预处理模型,本申请实施例中,预处理模型对输入参数的计算公式为:
在上述公式中,
代表对输入参数为Characters
environi的预测值,即是将Characters
environi作为输入参数输入至预处理模型后,预处理模型计算后的输出结果。公式中的f()指示一个存在于函数空间的函数,函数空间指的是从一个集合到另一个集合的给定种类的函数的集合,即f()函数最初处于未知状态,K则表示预处理模型中存在K个上述 的f()函数,需要将所有的f()函数计算出的结果累加后,才能得到最终的预测值。
在实际对预处理模型的训练过程,依赖于上述公式,在本申请实施例中,采用次序学习的方法对f()函数进行学习,以使最终得到的K个f()函数最大限度地符合多个样本信息中的数据。举例来说,在输入参数为Characters
environi的基础上,进行t轮的预测值预测,并在进行第t轮的预测值预测时,保留第t-1轮的预测值预测结果,即依照次序训练预处理模型,使得预测值
与真实值(Value
environi)之间的差距逐渐减小,具体见下:
上述公式中的
是在给出输入参数为Characters
environi的基础上,进行第t轮的预测后的预测值。为了确定在次序学习过程中所需求的f()函数,使其尽量贴近于样本信息集,故构建优化函数,具体公式见下:
在上述公式中,Value
environi是样本信息集中与输入参数Characters
environi对应的参照参数,即是样本信息中的样本特征值,在本申请实施例中作为优化函数的参数。优化函数公式中的Ω(f
t)为正则项,D为常数项,其中,正则项控制优化函数的训练程度,防止样本信息集和预处理模型过拟合;常数项为一个常量,设置常数项是为了限制优化函数的数值范围。值得一提的是,
为误差函数,对优化函数进行优化的过程,即是确定合适的f()函数使得上述误差函数的值尽量减小的过程。
展开后的优化函数为:
由于常数项实质并不影响优化函数的优化过程,故提取出展开后的优化函数中的常数项,可生成展开后的优化函数在第t轮的化简函数,公式如下:
在最终的化简函数中,化简函数得到的输出值依赖于g
i和h
i的值,故能够很快确定合适的f()函数,提升了训练的简便性,在本申请实施例中,通过上述的次序学习的方法以及优化函数(化简函数亦可)以训练预处理模型。
S202:将训练后的所述预处理模型输出为所述筛选模型。
当样本信息集中所有的输入参数和参照参数全部输入预处理模型,并且预处理模型训练完成后,将训练完成的预处理模型作为筛选模型(主要是训练完成的f()函数)进行输出。当需要进行待筛选人群的筛选时,将待筛选个体的个体特征Characters
environx输入筛选模型,即可通过筛选模型中优化后的计算公式
经过计算得到预测值
通过图2所示实施例可知,在本申请实施例中,将样本信息中的样本环境特征作为预处理模型的输入参数,将样本信息的样本特征值作为预处理模型的参照参数,从而将多个样本信息输入至预处理模型,以训练预处理模型,最后将训练完成的预处理模型作为筛选模型进 行输出,提升了筛选模型与多个样本信息的贴合度,并且提升了通过筛选模型进行人群筛选的准确性。
请参阅图3,图3是本申请实施例三提供的一种基于大数据的人群筛选方法的实现流程图。相对于图1对应的实施例,本实施例对S104细化后得到S301~S302,详述如下:
S301:获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件。
在本申请实施例中,筛选模型确定后,需要获取待筛选人群的多个个体特征,并自动将多个个体特征输入至筛选模型,得到包含多个输出特征值的输出特征值集合。为了获得输出特征值集合中满足预设条件的输出特征值,首先获取特征阈值,满足预设条件的输出特征值即为大于或等于特征阈值的输出特征值。
请参阅图4,图4是本申请实施例四提供的一种基于大数据的人群筛选方法的实现流程图。相对于图3对应的实施例,本实施例在样本特征值为第一特征值或第二特征值的基础上,对S301细化后得到S401~S402,详述如下:
S401:对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例。
对于不同的应用场景,样本特征值的取值范围可能会出现不同,在本申请实施例中,以样本特征值的取值为第一特征值或第二特征值,并且第一特征值大于第二特征值的情况进行说明。举例来说,样本特征值取值为第一特征值表示对应的样本个体受到焦虑症影响,样本特征值取值为第二特征值表示对应的样本个体未受到焦虑症影响,而筛选目的是从待筛选人群中筛选出受到焦虑症影响的目标人群,则首先对多个样本信息进行大数据分析,提取出样本特征值取值为第一特征值的第一样本数量,从而得到第一样本数量在所有样本信息中占有的数量比例。
S402:根据所述数量比例、所述第一特征值和所述第二特征值计算出所述特征阈值。
由于样本特征值的取值范围限定于第一特征值和第二特征值,故计算第一特征值和第二特征值之间的差值,并将第一特征值减去该差值与数量比例的乘积,得到特征阈值。举例来说,若第一特征值为2,第一数量比例为30%,第二特征值为1,第一数量比例为70%,则差值为2-1=1,特征阈值为2-1×30%=1.7。当然,上述计算方法仅适用于样本特征值的取值为第一特征值或第二特征值,且第一特征值大于第二特征值的情况,对于其他可能的样本特征值可在上述计算方法的基础上进行延伸,本申请实施例不再进行赘述。
请参阅图5,图5是本申请实施例五提供的一种基于大数据的人群筛选方法的实现流程图。相对于图3对应的实施例,本实施例对S301细化后得到S501~S503,详述如下:
S501:将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型 输出的与所述多个样本信息的样本环境特征对应的多个结果特征值。
在本申请实施例中,当筛选模型生成后,将多个样本信息输入至筛选模型。具体将多个样本信息的样本环境特征作为输入参数输入至筛选模型,但是不将多个样本信息的样本特征值作为筛选模型的参照参数,而是通过筛选模型的计算公式直接对多个样本信息的样本环境特征进行计算得到多个输出参数,即多个与样本环境特征对应的结果特征值。一般来说,样本信息经筛选模型计算后的结果特征值与样本信息原有的样本特征值存在差异。
S502:对所述多个结果特征值进行排序,生成结果特征值序列。
获取到多个结果特征值后,对多个结果特征值按照数值大小进行排序,生成数值序列,数值序列的列首即为数值最大的结果特征值。值得一提的是,若多个结果特征值中的第一结果特征值与第二结果特征值相同,则按照第一结果特征值和第二结果特征值对应的样本信息的输入次序生成特征值序列。具体地,设置次序写入机制,当某一个结果特征值需要写入特征值序列时,判断特征值序列中是否存在与该结果特征值相同的已有结果特征值,若不存在与该结果特征值相同的已有结果特征值,则按照该结果特征值的数值大小查找出特征值序列中最小的大于该结果特征值的已有结果特征值,并在该已有结果特征值后的位置写入该结果特征值;若存在与该结果特征值相同的已有结果特征值,则在该已有结果特征值后的位置写入该结果特征值,若存在多个与该结果特征值相同的已有结果特征值,则查找多个已有结果特征值末尾的已有结果特征值,并在该已有结果特征值后的位置写入该结果特征值。
S503:获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
特征值序列生成后,获取预设的筛选比例,并在特征值序列中查找与该筛选比例对应的筛选位置,将处于该筛选位置的结果特征值输出为特征阈值。举例来说,特征值序列中已写入300个结果特征值,而筛选比例为10%,则筛选位置为第30位,则提取特征值序列中位于第30位的结果特征值的数值作为特征阈值。可选地,筛选比例可根据大数据分析方法进行确定,例如可统计全国范围的多个全国样本信息,每个全国样本信息包括样本环境特征和样本特征值,首先样本特征值取值为与筛选目的对应的数值的筛选样本数量,从而确定筛选样本数量在所有全国样本信息中占有的比例,将该比例作为筛选比例,通过大数据分析方法可使得生成的特征阈值的普遍适用性较高。当然,筛选比例也可人为预先设定。
参阅图3,在步骤S302中:提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
特征阈值确定后,从输出特征值集合中提取出大于或等于特征阈值的输出特征值,则提取出的输出特征值与筛选的目标个体对应,故将提取出的输出特征值添加至目标特征值集合, 以便后续从待筛选人群中确定与目标特征值对应的目标人群。
通过图3所示实施例可知,在本申请实施例中,通过采用不同的方法获取特征阈值,并从输出特征值集合中提取出大于或等于特征阈值的输出特征值,并将提取出的输出特征值添加至目标特征值集合,通过设定特征阈值,提升了目标特征值集合生成的简便性和效率。
对应于上文实施例所述的一种基于大数据的人群筛选方法,图6示出了本申请实施例提供的一种基于大数据的人群筛选装置的一个结构框图,参照图6,该人群筛选装置包括:
第一获取单元61,用于获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;
拟合单元62,用于将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;
第二获取单元63,用于获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;
目标确定单元64,用于将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
可选地,所述拟合单元62,包括:
输入单元,用于将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数;
输出单元,用于将训练后的所述预处理模型输出为所述筛选模型。
可选地,所述目标确定单元64,包括:
阈值获取单元,用于获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件;
提取单元,用于提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
可选地,样本特征值为第一特征值或第二特征值,所述阈值获取单元,包括:
分析单元,用于对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例;
计算单元,用于根据所述数量比例和所述第一特征值和所述第二特征值计算出所述特征阈值。
可选地,所述阈值获取单元,包括:
特征输入单元,用于将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型输出的与所述多个样本信息的样本环境特征对应的多个结果特征值;
排序单元,用于对所述多个结果特征值进行排序,生成结果特征值序列;
阈值输出单元,用于获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
图7是本申请实施例提供的终端设备的示意图。如图7所示,该实施例的终端设备7包括:处理器70以及存储器71,所述存储器71中存储有可在所述处理器70上运行的计算机可读指令72,例如基于大数据的人群筛选程序。所述处理器70执行所述计算机可读指令72时实现上述各个基于大数据的人群筛选方法实施例中的步骤,例如图1所示的步骤S101至S104。或者,所述处理器70执行所述计算机可读指令72时实现上述人群筛选装置实施例中各单元的功能,例如图6所示单元61至64的功能。
示例性的,所述计算机可读指令72可以被分割成一个或多个模块/单元,所述一个或者多个模块/单元被存储在所述存储器71中,并由所述处理器70执行,以完成本申请。所述一个或多个模块/单元可以是能够完成特定功能的一系列计算机可读指令段,该指令段用于描述所述计算机可读指令72在所述终端设备7中的执行过程。例如,所述计算机可读指令72可以被分割成第一获取单元、拟合单元、第二获取单元及目标确定单元,各单元具体功能如上所述。
所述终端设备可包括,但不仅限于,处理器70、存储器71。本领域技术人员可以理解,图7仅仅是终端设备7的示例,并不构成对终端设备7的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件,例如所述终端设备还可以包括输入输出设备、网络接入设备、总线等。
所称处理器70可以是中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
所述存储器71可以是所述终端设备7的内部存储单元,例如终端设备7的硬盘或内存。所述存储器71也可以是所述终端设备7的外部存储设备,例如所述终端设备7上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,所述存储器71还可以既包括所述终端设备7的内部存 储单元也包括外部存储设备。所述存储器71用于存储所述计算机可读指令以及所述终端设备所需的其他程序和数据。所述存储器71还可以用于暂时地存储已经输出或者将要输出的数据。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,以上实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围。
Claims (20)
- 一种基于大数据的人群筛选方法,其特征在于,包括:获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
- 如权利要求1所述的人群筛选方法,其特征在于,所述将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型,包括:将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数;将训练后的所述预处理模型输出为所述筛选模型。
- 如权利要求1所述的人群筛选方法,其特征在于,所述将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,包括:获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件;提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
- 如权利要求3所述的人群筛选方法,其特征在于,所述样本特征值为第一特征值或第二特征值,所述获取特征阈值,包括:对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例;根据所述数量比例、所述第一特征值和所述第二特征值计算出所述特征阈值。
- 如权利要求3所述的人群筛选方法,其特征在于,所述获取特征阈值,包括:将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型输出的与所述多个样本信息的样本环境特征对应的多个结果特征值;对所述多个结果特征值进行排序,生成结果特征值序列;获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
- 一种基于大数据的人群筛选装置,其特征在于,包括:第一获取单元,用于获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;拟合单元,用于将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;第二获取单元,用于获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;目标确定单元,用于将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
- 如权利要求6所述的人群筛选装置,其特征在于,所述拟合单元,包括:输入单元,用于将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数;输出单元,用于将训练后的所述预处理模型输出为所述筛选模型。
- 如权利要求6所述的人群筛选装置,其特征在于,所述目标确定单元,包括:阈值获取单元,用于获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件;提取单元,用于提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
- 如权利要求8所述的人群筛选装置,其特征在于,所述症状特征值为第一特征值或第二特征值,所述阈值获取单元,包括:分析单元,用于对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例;计算单元,用于根据所述数量比例和所述第一特征值和所述第二特征值计算出所述特征阈值。
- 如权利要求8所述的人群筛选装置,其特征在于,所述阈值获取单元,包括:特征输入单元,用于将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型输出的与所述多个样本信息的样本环境特征对应的多个结果特征值;排序单元,用于对所述多个结果特征值进行排序,生成结果特征值序列;阈值输出单元,用于获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
- 一种终端设备,其特征在于,包括存储器以及处理器,所述存储器中存储有可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现如下步骤:获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
- 根据权利要求11所述的终端设备,其特征在于,所述将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型,包括:将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数;将训练后的所述预处理模型输出为所述筛选模型。
- 根据权利要求11所述的终端设备,其特征在于,所述将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,包括:获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件;提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
- 根据权利要求13所述的终端设备,其特征在于,所述样本特征值为第一特征值或第二特征值,所述获取特征阈值,包括:对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例;根据所述数量比例、所述第一特征值和所述第二特征值计算出所述特征阈值。
- 根据权利要求13所述的终端设备,其特征在于,所述获取特征阈值,包括:将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型输出的与所述多个样本信息的样本环境特征对应的多个结果特征值;对所述多个结果特征值进行排序,生成结果特征值序列;获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,其特征在于,所述计算机可读指令被至少一个处理器执行时实现如下步骤:获取样本人群的多个样本信息,所述样本信息包括样本环境特征和样本特征值,所述样本特征值用于描述所述样本信息在所述样本人群中对应样本个体的个体状态;将所述多个样本信息与预设的预处理模型进行拟合,并将拟合后的所述预处理模型作为筛选模型;获取待筛选人群的多个个体特征,并将所述多个个体特征输入至所述筛选模型,得到与所述多个个体特征对应的输出特征值集合,所述输出特征值集合包括多个输出特征值;将所述输出特征值集合中满足预设条件的输出特征值添加至目标特征值集合,并从所述待筛选人群中确定与所述目标特征值集合对应的目标人群。
- 根据权利要求16所述的计算机可读存储介质,其特征在于,所述计算机可读指令被至少一个处理器执行时实现如下步骤:将所述多个样本信息输入至所述预处理模型,以训练所述预处理模型,其中,将所述样本信息的样本环境特征作为所述预处理模型的输入参数,将所述样本信息的样本特征值作为所述预处理模型的参照参数;将训练后的所述预处理模型输出为所述筛选模型。
- 根据权利要求16所述的计算机可读存储介质,其特征在于,所述计算机可读指令被至少一个处理器执行时实现如下步骤:获取特征阈值,所述特征阈值用于判断所述输出特征值是否满足所述预设条件;提取出所述输出特征值集合中大于或等于所述特征阈值的输出特征值,并将提取出的输出特征值添加至所述目标特征值集合。
- 根据权利要求18所述的计算机可读存储介质,其特征在于,所述样本特征值为第一特征值或第二特征值,所述计算机可读指令被至少一个处理器执行时实现如下步骤:对所述多个样本信息进行大数据分析,确定样本特征值取值为所述第一特征值的所述样本信息在所述多个样本信息中占有的数量比例;根据所述数量比例、所述第一特征值和所述第二特征值计算出所述特征阈值。
- 根据权利要求18所述的计算机可读存储介质,其特征在于,所述计算机可读指令被至少一个处理器执行时实现如下步骤:将所述多个样本信息的样本环境特征输入至所述筛选模型,并获取所述筛选模型输出的与所述多个样本信息的样本环境特征对应的多个结果特征值;对所述多个结果特征值进行排序,生成结果特征值序列;获取预设的筛选比例,并在所述结果特征值序列中查找出与所述筛选比例对应的所述结果特征值,输出为所述特征阈值。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810455659.4 | 2018-05-14 | ||
| CN201810455659.4A CN108629381A (zh) | 2018-05-14 | 2018-05-14 | 基于大数据的人群筛选方法及终端设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019218482A1 true WO2019218482A1 (zh) | 2019-11-21 |
Family
ID=63693185
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/097561 Ceased WO2019218482A1 (zh) | 2018-05-14 | 2018-07-27 | 基于大数据的人群筛选方法、装置、终端设备及可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN108629381A (zh) |
| WO (1) | WO2019218482A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113610168A (zh) * | 2021-08-11 | 2021-11-05 | 平安科技(深圳)有限公司 | 数据处理方法、装置、设备及介质 |
| CN114334696A (zh) * | 2021-12-30 | 2022-04-12 | 中国电信股份有限公司 | 质量检测方法及装置、电子设备和计算机可读存储介质 |
| CN115146138A (zh) * | 2022-08-04 | 2022-10-04 | 重庆农村商业银行股份有限公司 | 一种用户筛选方法、装置、设备及可读存储介质 |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109726242A (zh) * | 2018-12-29 | 2019-05-07 | 陕西西部资信股份有限公司 | 数据处理方法及系统 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106214120A (zh) * | 2016-08-19 | 2016-12-14 | 靳晓亮 | 一种青光眼的早期筛查方法 |
| CN106706627A (zh) * | 2017-03-06 | 2017-05-24 | 温鹏 | 原血红素和β‑葡萄糖醛酸苷酶联合在鼻咽部上皮细胞异质性增生检测中的应用及试剂盒 |
| CN107895596A (zh) * | 2016-12-19 | 2018-04-10 | 平安科技(深圳)有限公司 | 风险预测方法及系统 |
-
2018
- 2018-05-14 CN CN201810455659.4A patent/CN108629381A/zh not_active Withdrawn
- 2018-07-27 WO PCT/CN2018/097561 patent/WO2019218482A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106214120A (zh) * | 2016-08-19 | 2016-12-14 | 靳晓亮 | 一种青光眼的早期筛查方法 |
| CN107895596A (zh) * | 2016-12-19 | 2018-04-10 | 平安科技(深圳)有限公司 | 风险预测方法及系统 |
| CN106706627A (zh) * | 2017-03-06 | 2017-05-24 | 温鹏 | 原血红素和β‑葡萄糖醛酸苷酶联合在鼻咽部上皮细胞异质性增生检测中的应用及试剂盒 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113610168A (zh) * | 2021-08-11 | 2021-11-05 | 平安科技(深圳)有限公司 | 数据处理方法、装置、设备及介质 |
| CN113610168B (zh) * | 2021-08-11 | 2024-05-14 | 平安科技(深圳)有限公司 | 数据处理方法、装置、设备及介质 |
| CN114334696A (zh) * | 2021-12-30 | 2022-04-12 | 中国电信股份有限公司 | 质量检测方法及装置、电子设备和计算机可读存储介质 |
| CN114334696B (zh) * | 2021-12-30 | 2024-03-05 | 中国电信股份有限公司 | 质量检测方法及装置、电子设备和计算机可读存储介质 |
| CN115146138A (zh) * | 2022-08-04 | 2022-10-04 | 重庆农村商业银行股份有限公司 | 一种用户筛选方法、装置、设备及可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN108629381A (zh) | 2018-10-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20240012846A1 (en) | Systems and methods for parsing log files using classification and a plurality of neural networks | |
| US20180253657A1 (en) | Real-time credit risk management system | |
| WO2019214248A1 (zh) | 一种风险评估方法、装置、终端设备及存储介质 | |
| CN111914090A (zh) | 一种企业行业分类识别及其特征污染物识别的方法及装置 | |
| WO2018157805A1 (zh) | 一种自动问答处理方法及自动问答系统 | |
| WO2017101506A1 (zh) | 信息处理方法及装置 | |
| CN114329034A (zh) | 基于细粒度语义特征差异的图像文本匹配判别方法及系统 | |
| CN108710907B (zh) | 手写体数据分类方法、模型训练方法、装置、设备及介质 | |
| CN108960269B (zh) | 数据集的特征获取方法、装置及计算设备 | |
| CN112785095A (zh) | 贷款预测方法、装置、电子设备和计算机可读存储介质 | |
| KR102280490B1 (ko) | 상담 의도 분류용 인공지능 모델을 위한 훈련 데이터를 자동으로 생성하는 훈련 데이터 구축 방법 | |
| WO2019218482A1 (zh) | 基于大数据的人群筛选方法、装置、终端设备及可读存储介质 | |
| US11657222B1 (en) | Confidence calibration using pseudo-accuracy | |
| CN112749737A (zh) | 图像分类方法及装置、电子设备、存储介质 | |
| CN112100374B (zh) | 文本聚类方法、装置、电子设备及存储介质 | |
| CN112528022A (zh) | 主题类别对应的特征词提取和文本主题类别识别方法 | |
| CN110634060A (zh) | 一种用户信用风险的评估方法、系统、装置及存储介质 | |
| CN110008972A (zh) | 用于数据增强的方法和装置 | |
| CN109657710B (zh) | 数据筛选方法、装置、服务器及存储介质 | |
| CN121051306B (zh) | 基于对比学习的跨域推荐方法、系统、设备及存储介质 | |
| CN114328916B (zh) | 事件抽取、及其模型的训练方法,及其装置、设备和介质 | |
| CN113806452B (zh) | 信息处理方法、装置、电子设备及存储介质 | |
| US20230214451A1 (en) | System and method for finding data enrichments for datasets | |
| CN114416977A (zh) | 文本难度分级评估方法及装置、设备和存储介质 | |
| WO2022237065A1 (zh) | 分类模型的训练方法、视频分类方法及相关设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18919012 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS (EPO FORM 1205A DATED 26.02.2021) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18919012 Country of ref document: EP Kind code of ref document: A1 |




