WO2017143919A1 - 一种建立数据识别模型的方法及装置 - Google Patents
一种建立数据识别模型的方法及装置 Download PDFInfo
- Publication number
- WO2017143919A1 WO2017143919A1 PCT/CN2017/073444 CN2017073444W WO2017143919A1 WO 2017143919 A1 WO2017143919 A1 WO 2017143919A1 CN 2017073444 W CN2017073444 W CN 2017073444W WO 2017143919 A1 WO2017143919 A1 WO 2017143919A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- training
- model
- samples
- sample set
- establishing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
- G06F18/2155—Generating training patterns; Bootstrap methods, e.g. bagging or boosting characterised by the incorporation of unlabelled data, e.g. multiple instance learning [MIL], semi-supervised techniques using expectation-maximisation [EM] or naïve labelling
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
- G06F18/2148—Generating training patterns; Bootstrap methods, e.g. bagging or boosting characterised by the process organisation or structure, e.g. boosting cascade
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/217—Validation; Performance evaluation; Active pattern learning techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
- G06F18/2415—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on parametric or probabilistic models, e.g. based on likelihood ratio or false acceptance rate versus a false rejection rate
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0499—Feedforward networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/018—Certifying business or products
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/02—Marketing; Price estimation or determination; Fundraising
- G06Q30/0201—Market modelling; Market analysis; Collecting market data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/06—Buying, selling or leasing transactions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/06—Buying, selling or leasing transactions
- G06Q30/0601—Electronic shopping [e-shopping]
- G06Q30/0609—Qualifying participants for shopping transactions
Definitions
- the invention belongs to the technical field of data processing, and in particular relates to a method and a device for establishing a data recognition model.
- the credit of the merchant is an important indicator for the consumer to decide whether to consume.
- the online e-commerce platform is also ranked according to the credit level of the merchant.
- the credit of the merchant is gradually accumulated according to the number and rating of the transaction, and the newly opened store has no credit, and the ranking will be lower. Consumers are more willing to choose higher-credit merchants or higher-selling commodities for their own rights and interests.
- the ranking of the business ranking is directly related to whether the consumer can search for the merchant. If the search is not available, the consumer cannot enter the merchant's store for consumption.
- E-commerce platforms such as Xiaoweijinhuahua and credit business use the training recognition model to identify whether the transaction is a false transaction.
- the TOP catch rate is used to measure the accuracy of the identification of false transactions.
- the so-called catch rate is also called the recall rate, which refers to the ratio of the identified false transactions to the total number of false transactions.
- the TOP catch rate is an index used to evaluate the model obtained by training.
- the transaction record is sorted according to the false transaction probability obtained by the model, and then the sorted transaction records are grouped to calculate the catch rate of each group.
- the TOP catch rate remains stable and can reach the set standard, then the model is judged to be reliable and can be used for subsequent identification.
- the E-commerce platform such as Xiaoweijinfu trains the recognition model
- it is generally first processed by the feature engineering after the training sample is trained by the logistic regression algorithm, and then the test sample is used to calculate the catch rate. Rate to determine whether the recognition model obtained by the training is reliable.
- the recognition model obtained by training now uses the logistic regression model.
- the training samples to be sampled proportionally, no positive samples are distinguished, resulting in noise entering the logistic regression algorithm, which cannot effectively improve the TOP catch rate and ensure stability.
- the linear model has been unable to learn more dimensional information, the model is single, and the effect is limited.
- a method for establishing a data recognition model for establishing a data recognition model based on training samples including positive and negative samples, the method for establishing a data recognition model includes:
- the training model is used for logistic regression training to obtain the first model
- the first model obtained by the training is used to identify the positive sample, and the second training sample set is selected from the positive samples having the recognition result after the first model is identified;
- the first training sample set obtained after sampling and the second training sample set are subjected to deep neural network DNN training to obtain a final data recognition model.
- the method for establishing a data identification model includes: before performing proportional sampling or performing logistic regression training, further comprising:
- Feature engineering preprocessing is performed on the training samples.
- the method for establishing a data identification model includes: before using the training sample for logistic regression training, the method further includes:
- Feature screening is performed on the training samples, and the feature filtering removes the feature that the information value is less than the set threshold by calculating the information value of the feature.
- the method before selecting the second training sample set from the positive samples having the recognition result after the first model is identified, the method further includes:
- the first training sample set is used for DNN training to obtain a second model.
- selecting the second training sample set from the positive samples having the recognition result after the first model is identified includes:
- the first model obtained by the training is evaluated to obtain a ROC curve corresponding to the first model
- the second model obtained by the training is evaluated to obtain a corresponding ROC curve of the second model
- the method of the present invention preferably selects the second training sample set to select a sample more in line with the training requirements and improve the stability of the final data recognition model.
- the present invention also provides an apparatus for establishing a data identification model for establishing a data identification model based on training samples including positive and negative samples, the apparatus comprising:
- a first training module configured to perform a logistic regression training using a training sample to obtain a first model
- a sampling module configured to sample the training samples proportionally, to obtain a first training sample set
- a selection module configured to identify the positive sample by using the first model obtained by the training, and select a second training sample set from the positive samples having the recognition result after the first model is identified;
- the final model training module is configured to perform deep neural network DNN training by using the first training sample set obtained after sampling and the second training sample set to obtain a final data recognition model.
- the device further includes:
- the pre-processing module is used to perform feature engineering pre-processing on the training samples before performing proportional sampling or performing logistic regression training.
- the device further includes:
- the feature screening module is configured to perform feature screening on the training sample before performing the logistic regression training using the training sample, and the feature filtering removes the feature that the information value is less than the set threshold by calculating the information value of the feature.
- the device of the present invention further comprises:
- the second training module is configured to perform DNN training by using the first training sample set to obtain a second model.
- the selection module selects the second training sample set from the positive samples having the recognition result after the first model is identified, the following operations are performed:
- the first model obtained by the training is evaluated to obtain a ROC curve corresponding to the first model
- the second model obtained by the training is evaluated to obtain a corresponding ROC curve of the second model
- the invention provides a method and a device for establishing a data recognition model, which performs feature engineering preprocessing and feature screening on all training samples, and uses the first model identification result obtained by logistic regression training and DNN using the first training sample set.
- the second training sample set is selected from all the positive samples with the recognition result, and the final data recognition model is obtained by combining the deep neural network training, thereby improving the stability of the model.
- FIG. 1 is a flow chart of a method for establishing a data identification model according to the present invention
- FIG. 3 is a schematic structural diagram of an apparatus for establishing a data identification model according to the present invention.
- the method for establishing a data identification model in this embodiment includes:
- Step S1 Perform feature engineering preprocessing on the training samples.
- the sample is first subjected to feature engineering preprocessing, that is, data replacement and cleaning are performed on the features of the sample, and the meaningless features are eliminated. For example, data replacement is performed on missing features in the sample.
- Step S2 Perform feature screening on the pre-processed training samples, perform logistic regression training using the trained training samples, and use the first model obtained by the training to identify the positive samples.
- the positive samples and the negative samples are included in all the training samples. This example is described by taking a false transaction as an example. A positive sample indicates a sample of a false transaction, and a negative sample indicates a sample that is not a false transaction.
- model recognition because some features have little to do with the final recognition result, if these features are used as variables, the model recognition results will be worse, or in general, the number of features should be much smaller than the number of samples. Therefore, it is necessary to use feature filtering. Screen out features that are not important or even negative. There are many methods for feature filtering, such as nearest neighbor algorithm, partial least squares, and so on.
- This embodiment preferably filters the features of the sample by employing an information value IV (information value). By calculating the information value corresponding to each feature of the sample, the sample feature whose feature value is smaller than the set threshold is removed, and the influence on the sample distribution is reduced.
- the information value corresponding to the sample feature is calculated according to the characteristics of all the training samples, and the characteristics of a training sample include ⁇ feature 1, feature 2, ..., feature m ⁇ , for which one feature i, i belongs to (1 to m), where m is the number of features. All training samples correspond to the value of feature i ⁇ i1, i2, ..., in ⁇ , where n is the total number of training samples.
- the value of feature i for example, the value of feature i is divided into a group, so that fenturei is divided into K groups, and the information value IV of the feature feature i is calculated according to the following formula:
- Disgood ki is the number of negative samples in the sample group
- Disbad ki is the number of positive samples in the sample group.
- the present embodiment is not limited to the number of negative samples, which is the number of positive samples, i.e., may also be used Disgood ki represents the number of positive samples
- Disbad ki represents the number of negative samples. Therefore, the feature can be filtered according to the information value corresponding to the feature, and the feature whose corresponding information value is less than the set threshold is discarded, and the feature that affects the result is retained for subsequent training, thereby improving the reliability of the training model.
- all the training samples after feature selection are used for logistic regression training to obtain the first model, which is the recognition model used in the prior art scheme.
- the present invention is further trained on this basis to obtain a more reliable model.
- all the training samples after feature selection are used for logistic regression training to obtain the stability of the first model. Some of the samples can be selected for subsequent training, so that the model obtained by the subsequent training has better stability.
- the stability of the measurement model is generally based on the TOP catch rate indicator, and the TOP catch rate can be calculated based on the false transaction probability obtained by the model identification sample.
- all the positive samples are identified by using the first model obtained by training, and the probability that each training sample corresponds to a false transaction is obtained, and all positive samples and their identified probabilities are training set B, that is, The first model is identified with a positive sample of the recognition result.
- a part of the training samples are selected from the training set B as the subsequent training according to the recognition result.
- step S3 the pre-processed training samples are sampled proportionally, and the first training sample set obtained after sampling is used for DNN training to obtain a second model.
- the accurate identification samples may be directly selected from the training set B as the second training sample set used for the subsequent training.
- the pre-processed all training samples are preferably sampled to obtain the training set A (the first training sample set), for example, the ratio of the positive and negative samples is 1:10.
- the training set A the first training sample set
- the ratio of the positive and negative samples is 1:10.
- first select all positive samples then select enough negative samples from the negative samples to maintain a 1:10 ratio.
- the first training sample set obtained after sampling is used for DNN training, and a second model can be obtained.
- DNN Deep Neural Networks
- DNN Deep Neural Networks
- the recognition result of the second model is not stable enough.
- the training of the second training sample set in the subsequent steps can obtain a stable final data recognition model.
- feature engineering preprocessing is performed on all training samples, and feature screening is used to filter out features that are not important or even negative, so that the model obtained by training is more reliable.
- the training model may be pre-processed and feature-filtered when the first model is trained and the second model is trained, or the feature screening may be performed only when the first model is trained. Feature screening is not performed in the second model. It is easy to understand that even if feature engineering preprocessing and feature screening are not performed, the recognition effect of the trained model can be improved, so that the recognition effect of the trained model is better than the prior art, and will not be described here.
- Step S4 Select a second training sample set from the positive samples having the recognition result after the first model is identified, according to the result of performing DNN training using the first training sample set and the result of identifying the positive sample by using the first model.
- the ROC curve is a graphical method for displaying the true rate and false positive rate of the model. It is commonly used to evaluate the effect of the model. Each point on the ROC curve has three values, which are the True Positive Rate (TPR). False Positive Rate (FPR) and threshold probability.
- TPR True Positive Rate
- FPR False Positive Rate
- the True Positive Rate (TPR) is the ratio of the positive and positive samples predicted by the model to positive; the false positive rate (FPR) is the negative and negative samples predicted positive by the model.
- the ratio of the actual number; the threshold probability is a decision threshold for determining that the prediction result is positive, and is determined to be positive if the result of the sample prediction is greater than the threshold probability, otherwise it is determined to be negative.
- the better the prediction effect of the model the closer the TPR is to 1, and the closer the FPR is to zero.
- a part of the training samples are selected from the training set B for subsequent training, and the specific methods selected include:
- the second model obtained by the training is evaluated to obtain a corresponding ROC curve of the second model
- the first model obtained by the training is evaluated to obtain a ROC curve corresponding to the first model
- the number of samples in the selected second training sample set is smaller than the number of positive samples in the first training sample set, and the maximum number of positive samples in the first training sample set is not exceeded, so as to ensure the proportion of positive and negative samples, Preventing too many positive samples leads to poor overall model performance.
- Selecting the second training sample set may also select a certain number of sample second training sample sets from the training set B according to the probability from the largest to the smallest according to the probability obtained by the model evaluation. Or, according to experience, a threshold is set, and a sample whose probability is greater than the threshold is selected from the training set B as a second training sample set.
- the invention is preferably based on the ROC curve The choice of intersection points will ensure better results in subsequent training.
- Step S5 performing DNN training using the first training sample set and the second training sample set to obtain a final data recognition model.
- the first training sample set and the second training sample set are used for DNN training to obtain a final data recognition model.
- the DNN deep learning training model is not described here.
- the ROC curve shown in Fig. 2 shows that the final data recognition model obtained by the training in this embodiment is far better than the first model effect obtained directly by logistic regression training.
- the upper curve in FIG. 2 is the ROC curve corresponding to the final data recognition model trained in the embodiment, and the lower curve is the ROC curve corresponding to the first model obtained directly through the logistic regression training.
- the embodiment further provides an apparatus for establishing a data identification model, which is used for establishing a data identification model according to a training sample including positive and negative samples, and the apparatus includes:
- a first training module configured to perform a logistic regression training using a training sample to obtain a first model
- a sampling module configured to sample the training samples proportionally, to obtain a first training sample set
- a selection module configured to identify the positive sample by using the first model obtained by the training, and select a second training sample set from the positive samples having the recognition result after the first model is identified;
- the final model training module is configured to perform deep neural network DNN training by using the first training sample set obtained after sampling and the second training sample set to obtain a final data recognition model.
- the device further includes:
- the pre-processing module is used to perform feature engineering pre-processing on the training samples before performing proportional sampling or performing logistic regression training.
- the device also includes:
- the feature screening module is configured to perform feature screening on the training sample before performing the logistic regression training using the training sample, and the feature filtering removes the feature that the information value is less than the set threshold by calculating the information value of the feature.
- the device further comprises:
- the second training module is configured to perform DNN training by using the first training sample set to obtain a second model.
- the second training data set is selected by using a preferred method.
- the selection module selects the second training sample set from the positive samples having the recognition result after the first model is identified, the following operations are performed:
- the first model obtained by the training is evaluated to obtain a ROC curve corresponding to the first model
- the second model obtained by the training is evaluated to obtain a corresponding ROC curve of the second model
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Business, Economics & Management (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- Accounting & Taxation (AREA)
- Finance (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Evolutionary Computation (AREA)
- Development Economics (AREA)
- Strategic Management (AREA)
- Life Sciences & Earth Sciences (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Business, Economics & Management (AREA)
- Marketing (AREA)
- Economics (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Entrepreneurship & Innovation (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Health & Medical Sciences (AREA)
- Game Theory and Decision Science (AREA)
- Medical Informatics (AREA)
- Probability & Statistics with Applications (AREA)
- Image Analysis (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Complex Calculations (AREA)
Abstract
一种建立数据识别模型的方法及装置,用于根据包括正、负样本的训练样本建立数据识别模型,该方法包括:对训练样本进行特征工程预处理(S1);对预处理后的训练样本进行特征筛选,采用特征筛选后的训练样本进行逻辑回归训练,采用训练得到的第一模型对正样本进行识别(S2);对预处理后的训练样本按比例采样,采用采样后得到的第一训练样本集进行DNN训练,得到第二模型(S3);根据采用第一训练样本集进行DNN训练的结果与采用第一模型对正样本进行识别的结果,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集(S4);采用第一训练样本集和第二训练样本集进行DNN训练得到最终的数据识别模型(S5)。装置包括第一训练模块、采样模块、选择模块和最终模型训练模块。提高了数据识别模型的稳定性。
Description
本申请要求2016年02月26日递交的申请号为201610110817.3、发明名称为“一种建立数据识别模型的方法及装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本发明属于数据处理技术领域,尤其涉及一种建立数据识别模型的方法及装置。
商家的信用是消费者决定是否消费的重要指标,目前网上电商平台也是按照商家的信用高低进行排名。商家的信用根据交易的数量和评分逐步累积,刚开的店铺没有信用,排名就会靠后。消费者出于对自身权益的考虑,更愿意选择信用较高的商家或者销量较高的商品。而商家排名的先后直接关系到消费者是否能够搜索到商家,搜索不到的情况下,消费者就无法进入商家的店铺进行消费。
因此网上商家都有提升信用的需求,催生了一些专为商家提升信用的网站和个人,通过刷单等虚假交易行为来提升商家的信用。虚假交易行为不利于市场的健康发展,不利于保护消费者的权益,属于电商平台需要严厉打击的行为。
电商平台例如小微金服花呗和信贷业务,在使用时都要利用训练得到的识别模型来识别交易是否是虚假交易。通常在业务上通过TOP抓坏率来衡量对虚假交易的识别是否准确,所谓抓坏率也称为召回率,是指识别出的虚假交易占虚假交易总数的比率。TOP抓坏率是用于对训练得到的模型进行评估的指标,按模型识别得到的虚假交易概率对交易记录进行排序,然后对排序后的交易记录进行分组,计算各组的抓坏率,如果TOP抓坏率保持稳定且能达到设定的标准,则判断模型可靠,可用于后续的识别。
然而目前小微金服等电商平台在训练识别模型时,一般是先对训练样本通过特征工程处理后,经过逻辑回归算法训练得到识别模型,然后采用测试样本来计算抓坏率,根据抓坏率来判断训练得到的识别模型是否可靠。
但是现在训练得到的识别模型是使用逻辑回归模型,对于训练样本按比例采样,没有对正样本进行区分,导致噪音进入逻辑回归算法,无法有效提高TOP抓坏率和保证稳定性。并且随着虚假交易维度越来越多,线性模型已经无法学到更多维度的信息,模型单一,效果受限。
发明内容
本发明的目的是提供一种建立数据识别模型的方法及装置,以解决现有技术逻辑回归模型训练时噪音的影响,以及模型单一、效果不理想等问题。结合机器学习和深度学习进行训练,在判断虚假交易时,有效提高TOP抓坏率,取得很好的效果。
为了实现上述目的,本发明技术方案如下:
一种建立数据识别模型的方法,用于根据包括正、负样本的训练样本建立数据识别模型,所述建立数据识别模型的方法包括:
采用训练样本进行逻辑回归训练,得到第一模型;
对训练样本按比例采样,获得第一训练样本集;
采用训练得到的第一模型对正样本进行识别,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集;
采用采样后得到的第一训练样本集与所述第二训练样本集进行深度神经网络DNN训练,得到最终的数据识别模型。
进一步地,所述建立数据识别模型的方法,在进行按比例采样或进行逻辑回归训练前,还包括:
对训练样本进行特征工程预处理。
进一步地,所述建立数据识别模型的方法,在采用训练样本进行逻辑回归训练之前,还包括:
对训练样本进行特征筛选,所述特征筛选通过计算特征的信息值,去除信息值小于设定阈值的特征。
优选地,所述从第一模型识别后具有识别结果的正样本中选择出第二训练样本集之前,还包括:
采用第一训练样本集进行DNN训练,得到第二模型。
进一步地,所述从第一模型识别后具有识别结果的正样本中选择出第二训练样本集,包括:
对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;
对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;
根据第一模型与第二模型ROC曲线的交点对应的阈值概率,从第一模型识别后具有识别结果的正样本中选择出概率小于所述阈值概率的样本作为第二训练样本集。
本发明优选地选择第二训练样本集的方法能够选择出更加符合训练要求的样本,提高最终数据识别模型的稳定性。
本发明还提出了一种建立数据识别模型的装置,用于根据包括正、负样本的训练样本建立数据识别模型,所述装置包括:
第一训练模块,用于采用训练样本进行逻辑回归训练,得到第一模型;
采样模块,用于对训练样本按比例采样,获得第一训练样本集;
选择模块,用于采用训练得到的第一模型对正样本进行识别,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集;
最终模型训练模块,用于采用采样后得到的第一训练样本集与所述第二训练样本集进行深度神经网络DNN训练,得到最终的数据识别模型。
进一步地,所述装置还包括:
预处理模块,用于在进行按比例采样或进行逻辑回归训练前,对训练样本进行特征工程预处理。
进一步地,所述装置还包括:
特征筛选模块,用于在采用训练样本进行逻辑回归训练之前,对训练样本进行特征筛选,所述特征筛选通过计算特征的信息值,去除信息值小于设定阈值的特征。
优选地,本发明所述装置还包括:
第二训练模块,用于采用第一训练样本集进行DNN训练,得到第二模型。
进一步地,所述选择模块从第一模型识别后具有识别结果的正样本中选择出第二训练样本集时,执行如下操作:
对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;
对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;
根据第一模型与第二模型ROC曲线的交点对应的阈值概率,从第一模型识别后具有识别结果的正样本中选择出概率小于所述阈值概率的样本作为第二训练样本集。
本发明提出的一种建立数据识别模型的方法及装置,通过对全部训练样本进行特征工程预处理以及特征筛选,并根据逻辑回归训练得到的第一模型识别结果和采用第一训练样本集进行DNN训练的结果,从具有识别结果的所有正样本中选择出第二训练样本集,来结合深度神经网络训练得到最终的数据识别模型,提高了模型的稳定性。
图1为本发明建立数据识别模型的方法流程图;
图2为本发明数据识别模型评估效果对照图;
图3为本发明建立数据识别模型的装置结构示意图。
下面结合附图和实施例对本发明技术方案做进一步详细说明,以下实施例不构成对本发明的限定。
如图1所示,本实施例一种建立数据识别模型的方法,包括:
步骤S1、对训练样本进行特征工程预处理。
对于获取的全部训练样本,由于样本中的特征有些值缺失,或者偏差超出正常的范围,会影响到后续的训练,通常需要对样本进行特征工程处理。本实施例首先对样本进行特征工程预处理,即对样本的特征进行数据替换和清洗,剔除无意义特征。例如对样本中缺失的特征进行数据替换等。
步骤S2、对预处理后的训练样本进行特征筛选,采用特征筛选后的训练样本进行逻辑回归训练,采用训练得到的第一模型对正样本进行识别。
全部训练样本中包括正样本和负样本,本实施例以虚假交易为例来进行说明,正样本表示是虚假交易的样本,负样本表示不是虚假交易的样本。
在模型识别中,因为有些特征与最终识别结果关系不大,若把这些特征作为变量会使得模型识别结果变差,或一般情况下应使特征数大大小于样本数,所以有必要采用特征筛选来筛选掉不重要甚至有负作用的特征。进行特征筛选的方法很多,例如有最近邻算法、偏最小二乘法等。本实施例优选地通过采用信息值IV(information value)来对样本的特征进行筛选。通过计算样本每个特征对应的信息值,将特征对应的信息值小于设定阈值的样本特征去除,减少其对样本分布的影响。
本实施例计算样本特征对应的信息值是根据所有训练样本的特征来计算,假设一条训练样本的特征包括{feature 1、feature 2、…、feature m},对于其中的一个特征feature i,i属于(1~m),m为特征数量。所有训练样本对应该feature i的值为{i1,i2,…,in},n为训练样本总数。
则可以根据feature i的值进行分组,例如将feature i的值为a的划分为一组,这样将fenturei分为K组,根据如下公式计算特征feature i的信息值IV:
其中,Disgoodki为样本组中负样本数量,Disbadki为样本组中正样本数量。本实施例不限定哪个为负样本数量,哪个为正样本数量,即也可以用Disgoodki表示正样本数量,Disbadki表示负样本数量。从而可以根据特征对应的信息值来筛选特征,将对应信息值小于设定阈值的特征舍弃,保留对结果有影响的特征用来进行后续的训练,提高训练模型的可靠性。
在进行特征筛选后,采用特征筛选后的全部训练样本进行逻辑回归训练得到第一模型,该模型即为现有技术方案中采用的识别模型。本发明在此基础上进一步训练以得到更加可靠的模型。一般来说采用特征筛选后的全部训练样本进行逻辑回归训练得到第一模型稳定性比较好,可以选择其中的一些样本来进行后续的训练,以使得后续训练得到的模型具有较好的稳定性。衡量模型稳定性一般采用TOP抓坏率指标,TOP抓坏率可以根据模型识别样本得到的虚假交易概率来进行计算。
为此,本实施例采用训练得到的第一模型对所有正样本进行识别,得到每个训练样本对应的为虚假交易的概率,记所有正样本及其识别得到的概率为训练集合B,即通过第一模型识别后具有识别结果的正样本。在后续步骤中根据识别结果从训练集合B中选择一部分训练样本作为后续的训练用。
步骤S3、对预处理后的训练样本按比例采样,采用采样后得到的第一训练样本集进行DNN训练,得到第二模型。
为了从训练集合B中选择一部分训练样本作为后续的训练用,可以直接从训练集合B中选择识别准确的样本作为后续训练采用的第二训练样本集。
本实施例优选地对预处理后的全部训练样本按比例采样得到训练集合A(第一训练样本集),例如正负样本的比例为1:10。在操作中,先选择出所有的正样本,然后从负样本中选择足够多的负样本,保持1:10的比例。然后采用采样后得到的第一训练样本集进行DNN训练,可以得到一个第二模型。深度神经网络DNN(Deep Neural Networks)是近年来机器学习领域中的研究热点,DNN训练广泛应用在语音识别及其他数据分类上,关于DNN训练的内容这里不再赘述。
在后续步骤中根据第二模型的训练结果与第一模型的训练结果从训练集合B中选择
第二训练样本集。
根据实验得到的经验,第二模型的识别结果稳定性不够。而结合第二训练样本集在后续步骤中进行训练能够得到稳定性好的最终数据识别模型。
需要说明的是,本实施例对全部训练样本进行特征工程预处理,以及采用特征筛选来筛选掉不重要甚至有负作用的特征,都是为了训练得到的模型更加可靠。在具体的实施例中,可以在训练得到第一模型和训练得到第二模型时都需要对训练样本进行预处理和特征筛选,也可以仅在训练得到第一模型时进行特征筛选,而在训练第二模型时不进行特征筛选。容易理解的是,即使不进行特征工程预处理及特征筛选,也能提高训练得到的模型的识别效果,使得训练得到的模型的识别效果好于现有技术,这里不再赘述。
步骤S4、根据采用第一训练样本集进行DNN训练的结果与采用第一模型对正样本进行识别的结果,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集。
ROC曲线是显示模型真正率和假正率的一种图形化方法,常用来评估模型的效果,ROC曲线上每个点对应有三个值,分别为纵坐标真正率(True Positive Rate,TPR)、横坐标假正率(False Positive Rate,FPR)和阈值概率。真正率(True Positive Rate,TPR)是指被模型预测为正的正样本与正样本实际数量的比率;假正率(False Positive Rate,FPR)是指被模型预测为正的负样本与负样本实际数量的比率;阈值概率是用来判定预测结果为正的判定阈值,如果样本预测的结果大于该阈值概率则判定为正,否则判定为负。模型的预测效果越好,其TPR越接近于1,FPR越接近于0。
本实施例从训练集合B中选择一部分训练样本作为后续的训练用,选择的具体方法包括:
对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;
对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;
根据第一模型与第二模型ROC曲线的交点对应的阈值概率,选择训练集合B中概率小于该阈值概率的样本,作为第二训练样本集。
需要说明的是,选择的第二训练样本集中的样本数量小于第一训练样本集中的正样本数量,最多不超过第一训练样本集中的正样本数量,这样是为了保证正负样本的比例,以防止正样本过多导致模型整体效果变差。
选择第二训练样本集还可以根据模型评估得到的概率,从训练集合B中按照概率从大到小顺序选择一定数量的样本第二训练样本集。或者根据经验设定一个阈值,从训练集合B中选择概率大于该阈值的样本作为第二训练样本集。本发明优选地根据ROC曲线
的交点进行选择,能够保证在后续的训练中得到更好的结果。
步骤S5、采用第一训练样本集和第二训练样本集进行DNN训练得到最终的数据识别模型。
最后采用第一训练样本集和第二训练样本集进行DNN训练得到最终的数据识别模型,关于DNN深度学习训练模型,这里不再赘述。如图2所示的ROC曲线表明,本实施例训练得到的最终的数据识别模型效果远远好于直接通过逻辑回归训练得到的第一模型效果。图2中上面的曲线为本实施例训练得到的最终的数据识别模型对应的ROC曲线,下面的曲线为直接通过逻辑回归训练得到的第一模型对应的ROC曲线。
通过对最终数据识别模型TOP抓坏率的计算,可以发现本实施例提出的建立数据识别模型的方法大大提高了模型的稳定性。
如图3所示,本实施例还提出了一种建立数据识别模型的装置,用于根据包括正、负样本的训练样本建立数据识别模型,该装置包括:
第一训练模块,用于采用训练样本进行逻辑回归训练,得到第一模型;
采样模块,用于对训练样本按比例采样,获得第一训练样本集;
选择模块,用于采用训练得到的第一模型对正样本进行识别,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集;
最终模型训练模块,用于采用采样后得到的第一训练样本集与所述第二训练样本集进行深度神经网络DNN训练,得到最终的数据识别模型。
与上述方法对应地,容易理解的是,本装置还包括:
预处理模块,用于在进行按比例采样或进行逻辑回归训练前,对训练样本进行特征工程预处理。
以及,本装置还包括:
特征筛选模块,用于在采用训练样本进行逻辑回归训练之前,对训练样本进行特征筛选,所述特征筛选通过计算特征的信息值,去除信息值小于设定阈值的特征。
优选地,本装置还包括:
第二训练模块,用于采用第一训练样本集进行DNN训练,得到第二模型。
则本实施例采用优选的方法来选择第二训练数据集,选择模块从第一模型识别后具有识别结果的正样本中选择出第二训练样本集时,执行如下操作:
对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;
对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;
根据第一模型与第二模型ROC曲线的交点对应的阈值概率,从第一模型识别后具有识别结果的正样本中选择出概率小于所述阈值概率的样本作为第二训练样本集。
以上实施例仅用以说明本发明的技术方案而非对其进行限制,在不背离本发明精神及其实质的情况下,熟悉本领域的技术人员当可根据本发明作出各种相应的改变和变形,但这些相应的改变和变形都应属于本发明所附的权利要求的保护范围。
Claims (10)
- 一种建立数据识别模型的方法,用于根据包括正、负样本的训练样本建立数据识别模型,其特征在于,所述建立数据识别模型的方法包括:采用训练样本进行逻辑回归训练,得到第一模型;对训练样本按比例采样,获得第一训练样本集;采用训练得到的第一模型对正样本进行识别,从第一模型识别后具有识别结果的正样本中选择出第二训练样本集;采用采样后得到的第一训练样本集与所述第二训练样本集进行深度神经网络DNN训练,得到最终的数据识别模型。
- 根据权利要求1所述的建立数据识别模型的方法,其特征在于,所述建立数据识别模型的方法,在进行按比例采样或进行逻辑回归训练前,还包括:对训练样本进行特征工程预处理。
- 根据权利要求2所述的建立数据识别模型的方法,其特征在于,所述建立数据识别模型的方法,在采用训练样本进行逻辑回归训练之前,还包括:对训练样本进行特征筛选,所述特征筛选通过计算特征的信息值,去除信息值小于设定阈值的特征。
- 根据权利要求1所述的建立数据识别模型的方法,其特征在于,所述从第一模型识别后具有识别结果的正样本中选择出第二训练样本集之前,还包括:采用第一训练样本集进行DNN训练,得到第二模型。
- 根据权利要求4所述的建立数据识别模型的方法,其特征在于,所述从第一模型识别后具有识别结果的正样本中选择出第二训练样本集,包括:对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;根据第一模型与第二模型ROC曲线的交点对应的阈值概率,从第一模型识别后具有识别结果的正样本中选择出概率小于所述阈值概率的样本作为第二训练样本集。
- 一种建立数据识别模型的装置,用于根据包括正、负样本的训练样本建立数据识别模型,其特征在于,所述装置包括:第一训练模块,用于采用训练样本进行逻辑回归训练,得到第一模型;采样模块,用于对训练样本按比例采样,获得第一训练样本集;选择模块,用于采用训练得到的第一模型对正样本进行识别,从第一模型识别后具 有识别结果的正样本中选择出第二训练样本集;最终模型训练模块,用于采用采样后得到的第一训练样本集与所述第二训练样本集进行深度神经网络DNN训练,得到最终的数据识别模型。
- 根据权利要求6所述的建立数据识别模型的装置,其特征在于,所述装置还包括:预处理模块,用于在进行按比例采样或进行逻辑回归训练前,对训练样本进行特征工程预处理。
- 根据权利要求7所述的建立数据识别模型的装置,其特征在于,所述装置还包括:特征筛选模块,用于在采用训练样本进行逻辑回归训练之前,对训练样本进行特征筛选,所述特征筛选通过计算特征的信息值,去除信息值小于设定阈值的特征。
- 根据权利要求6所述的建立数据识别模型的装置,其特征在于,所述装置还包括:第二训练模块,用于采用第一训练样本集进行DNN训练,得到第二模型。
- 根据权利要求9所述的建立数据识别模型的装置,其特征在于,所述选择模块从第一模型识别后具有识别结果的正样本中选择出第二训练样本集时,执行如下操作:对训练得到的第一模型进行评估,得到第一模型对应的ROC曲线;对训练得到的第二模型进行评估,得到第二模型对应的ROC曲线;根据第一模型与第二模型ROC曲线的交点对应的阈值概率,从第一模型识别后具有识别结果的正样本中选择出概率小于所述阈值概率的样本作为第二训练样本集。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/112,637 US11551036B2 (en) | 2016-02-26 | 2018-08-24 | Methods and apparatuses for building data identification models |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610110817.3A CN107133628A (zh) | 2016-02-26 | 2016-02-26 | 一种建立数据识别模型的方法及装置 |
| CN201610110817.3 | 2016-02-26 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US16/112,637 Continuation US11551036B2 (en) | 2016-02-26 | 2018-08-24 | Methods and apparatuses for building data identification models |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017143919A1 true WO2017143919A1 (zh) | 2017-08-31 |
Family
ID=59684712
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/073444 Ceased WO2017143919A1 (zh) | 2016-02-26 | 2017-02-14 | 一种建立数据识别模型的方法及装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11551036B2 (zh) |
| CN (1) | CN107133628A (zh) |
| TW (1) | TWI739798B (zh) |
| WO (1) | WO2017143919A1 (zh) |
Cited By (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109559214A (zh) * | 2017-09-27 | 2019-04-02 | 阿里巴巴集团控股有限公司 | 虚拟资源分配、模型建立、数据预测方法及装置 |
| CN109636242A (zh) * | 2019-01-03 | 2019-04-16 | 深圳壹账通智能科技有限公司 | 企业评分方法、装置、介质及电子设备 |
| CN109685527A (zh) * | 2018-12-14 | 2019-04-26 | 拉扎斯网络科技(上海)有限公司 | 检测商户虚假交易的方法、装置、系统及计算机存储介质 |
| CN110263824A (zh) * | 2019-05-29 | 2019-09-20 | 阿里巴巴集团控股有限公司 | 模型的训练方法、装置、计算设备及计算机可读存储介质 |
| CN110348523A (zh) * | 2019-07-15 | 2019-10-18 | 北京信息科技大学 | 一种基于Stacking的恶意网页集成识别方法及系统 |
| CN110472137A (zh) * | 2019-07-05 | 2019-11-19 | 中国平安人寿保险股份有限公司 | 识别模型的负样本构建方法、装置和系统 |
| CN111667028A (zh) * | 2020-07-09 | 2020-09-15 | 腾讯科技(深圳)有限公司 | 一种可靠负样本确定方法和相关装置 |
| CN111931848A (zh) * | 2020-08-10 | 2020-11-13 | 中国平安人寿保险股份有限公司 | 数据的特征提取方法、装置、计算机设备及存储介质 |
| CN112350956A (zh) * | 2020-10-23 | 2021-02-09 | 新华三大数据技术有限公司 | 一种网络流量识别方法、装置、设备及机器可读存储介质 |
| CN112561082A (zh) * | 2020-12-22 | 2021-03-26 | 北京百度网讯科技有限公司 | 生成模型的方法、装置、设备以及存储介质 |
| CN114067415A (zh) * | 2021-11-26 | 2022-02-18 | 北京百度网讯科技有限公司 | 回归模型的训练方法、对象评估方法、装置、设备和介质 |
| US11551036B2 (en) | 2016-02-26 | 2023-01-10 | Alibaba Group Holding Limited | Methods and apparatuses for building data identification models |
Families Citing this family (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107391760B (zh) * | 2017-08-25 | 2018-05-25 | 平安科技(深圳)有限公司 | 用户兴趣识别方法、装置及计算机可读存储介质 |
| CN107798390B (zh) * | 2017-11-22 | 2023-03-21 | 创新先进技术有限公司 | 一种机器学习模型的训练方法、装置以及电子设备 |
| US11539716B2 (en) * | 2018-07-31 | 2022-12-27 | DataVisor, Inc. | Online user behavior analysis service backed by deep learning models trained on shared digital information |
| CN109241770B (zh) * | 2018-08-10 | 2021-11-09 | 深圳前海微众银行股份有限公司 | 基于同态加密的信息值计算方法、设备及可读存储介质 |
| CN109325357B (zh) * | 2018-08-10 | 2021-12-14 | 深圳前海微众银行股份有限公司 | 基于rsa的信息值计算方法、设备及可读存储介质 |
| CN109242165A (zh) * | 2018-08-24 | 2019-01-18 | 蜜小蜂智慧(北京)科技有限公司 | 一种模型训练及基于模型训练的预测方法及装置 |
| CN110009509B (zh) * | 2019-01-02 | 2021-02-19 | 创新先进技术有限公司 | 评估车损识别模型的方法及装置 |
| CN109919931B (zh) * | 2019-03-08 | 2020-12-25 | 数坤(北京)网络科技有限公司 | 冠脉狭窄度评价模型训练方法及评价系统 |
| CN110163652B (zh) * | 2019-04-12 | 2021-07-13 | 上海上湖信息技术有限公司 | 获客转化率预估方法及装置、计算机可读存储介质 |
| CN110363534B (zh) * | 2019-06-28 | 2023-11-17 | 创新先进技术有限公司 | 用于识别异常交易的方法及装置 |
| CN111160485B (zh) * | 2019-12-31 | 2022-11-29 | 中国民用航空总局第二研究所 | 基于回归训练的异常行为检测方法、装置及电子设备 |
| CN111340102B (zh) * | 2020-02-24 | 2022-03-01 | 支付宝(杭州)信息技术有限公司 | 评估模型解释工具的方法和装置 |
| CN113762579B (zh) * | 2021-01-07 | 2025-11-18 | 北京沃东天骏信息技术有限公司 | 一种模型训练方法、装置、计算机存储介质及设备 |
| US12577871B2 (en) | 2022-08-03 | 2026-03-17 | Schlumberger Technology Corporation | Linear cut generation method for sensor inversion constraint imposition |
| WO2024030525A1 (en) | 2022-08-03 | 2024-02-08 | Schlumberger Technology Corporation | Automated record quality determination and processing for pollutant emission quantification |
| CN115544341A (zh) * | 2022-09-14 | 2022-12-30 | 中国银联股份有限公司 | 信息处理方法、装置、设备及存储介质 |
| EP4619907A4 (en) * | 2022-12-15 | 2026-03-25 | Services Petroliers Schlumberger | METHANE EMISSIONS MONITORING BASED ON MACHINE LEARNING |
| CN115905548B (zh) * | 2023-03-03 | 2024-05-10 | 美云智数科技有限公司 | 水军识别方法、装置、电子设备及存储介质 |
| AU2024284054A1 (en) | 2023-06-09 | 2026-01-08 | Schlumberger Technology B.V. | Emission detecting camera placement planning using 3d models |
| US12254622B2 (en) | 2023-06-16 | 2025-03-18 | Schlumberger Technology Corporation | Computing emission rate from gas density images |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101799875A (zh) * | 2010-02-10 | 2010-08-11 | 华中科技大学 | 一种目标检测方法 |
| CN103902968A (zh) * | 2014-02-26 | 2014-07-02 | 中国人民解放军国防科学技术大学 | 一种基于AdaBoost分类器的行人检测模型训练方法 |
| US20150095017A1 (en) * | 2013-09-27 | 2015-04-02 | Google Inc. | System and method for learning word embeddings using neural language models |
| CN104966097A (zh) * | 2015-06-12 | 2015-10-07 | 成都数联铭品科技有限公司 | 一种基于深度学习的复杂文字识别方法 |
Family Cites Families (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7096207B2 (en) * | 2002-03-22 | 2006-08-22 | Donglok Kim | Accelerated learning in machine vision using artificially implanted defects |
| US9443141B2 (en) * | 2008-06-02 | 2016-09-13 | New York University | Method, system, and computer-accessible medium for classification of at least one ICTAL state |
| CN102147851B (zh) * | 2010-02-08 | 2014-06-04 | 株式会社理光 | 多角度特定物体判断设备及多角度特定物体判断方法 |
| US20150112765A1 (en) * | 2013-10-22 | 2015-04-23 | Linkedln Corporation | Systems and methods for determining recruiting intent |
| US9978362B2 (en) * | 2014-09-02 | 2018-05-22 | Microsoft Technology Licensing, Llc | Facet recommendations from sentiment-bearing content |
| CN104636732B (zh) * | 2015-02-12 | 2017-11-07 | 合肥工业大学 | 一种基于序列深信度网络的行人识别方法 |
| CN104702492B (zh) * | 2015-03-19 | 2019-10-18 | 百度在线网络技术(北京)有限公司 | 垃圾消息模型训练方法、垃圾消息识别方法及其装置 |
| WO2017004448A1 (en) * | 2015-07-02 | 2017-01-05 | Indevr, Inc. | Methods of processing and classifying microarray data for the detection and characterization of pathogens |
| CN105184226A (zh) * | 2015-08-11 | 2015-12-23 | 北京新晨阳光科技有限公司 | 数字识别方法和装置及神经网络训练方法和装置 |
| CN107133628A (zh) | 2016-02-26 | 2017-09-05 | 阿里巴巴集团控股有限公司 | 一种建立数据识别模型的方法及装置 |
| US20170249594A1 (en) * | 2016-02-26 | 2017-08-31 | Linkedln Corporation | Job search engine for recent college graduates |
-
2016
- 2016-02-26 CN CN201610110817.3A patent/CN107133628A/zh active Pending
-
2017
- 2017-02-08 TW TW106104133A patent/TWI739798B/zh active
- 2017-02-14 WO PCT/CN2017/073444 patent/WO2017143919A1/zh not_active Ceased
-
2018
- 2018-08-24 US US16/112,637 patent/US11551036B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101799875A (zh) * | 2010-02-10 | 2010-08-11 | 华中科技大学 | 一种目标检测方法 |
| US20150095017A1 (en) * | 2013-09-27 | 2015-04-02 | Google Inc. | System and method for learning word embeddings using neural language models |
| CN103902968A (zh) * | 2014-02-26 | 2014-07-02 | 中国人民解放军国防科学技术大学 | 一种基于AdaBoost分类器的行人检测模型训练方法 |
| CN104966097A (zh) * | 2015-06-12 | 2015-10-07 | 成都数联铭品科技有限公司 | 一种基于深度学习的复杂文字识别方法 |
Cited By (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11551036B2 (en) | 2016-02-26 | 2023-01-10 | Alibaba Group Holding Limited | Methods and apparatuses for building data identification models |
| US10891161B2 (en) | 2017-09-27 | 2021-01-12 | Advanced New Technologies Co., Ltd. | Method and device for virtual resource allocation, modeling, and data prediction |
| CN109559214A (zh) * | 2017-09-27 | 2019-04-02 | 阿里巴巴集团控股有限公司 | 虚拟资源分配、模型建立、数据预测方法及装置 |
| EP3617983A4 (en) * | 2017-09-27 | 2020-05-06 | Alibaba Group Holding Limited | METHOD AND DEVICE FOR ALLOCATING VIRTUAL RESOURCES, MODELING AND DATA PREDICTION |
| US10691494B2 (en) | 2017-09-27 | 2020-06-23 | Alibaba Group Holding Limited | Method and device for virtual resource allocation, modeling, and data prediction |
| CN109685527B (zh) * | 2018-12-14 | 2024-03-29 | 拉扎斯网络科技(上海)有限公司 | 检测商户虚假交易的方法、装置、系统及计算机存储介质 |
| CN109685527A (zh) * | 2018-12-14 | 2019-04-26 | 拉扎斯网络科技(上海)有限公司 | 检测商户虚假交易的方法、装置、系统及计算机存储介质 |
| CN109636242A (zh) * | 2019-01-03 | 2019-04-16 | 深圳壹账通智能科技有限公司 | 企业评分方法、装置、介质及电子设备 |
| CN110263824A (zh) * | 2019-05-29 | 2019-09-20 | 阿里巴巴集团控股有限公司 | 模型的训练方法、装置、计算设备及计算机可读存储介质 |
| CN110263824B (zh) * | 2019-05-29 | 2023-09-05 | 创新先进技术有限公司 | 模型的训练方法、装置、计算设备及计算机可读存储介质 |
| CN110472137A (zh) * | 2019-07-05 | 2019-11-19 | 中国平安人寿保险股份有限公司 | 识别模型的负样本构建方法、装置和系统 |
| CN110472137B (zh) * | 2019-07-05 | 2023-07-25 | 中国平安人寿保险股份有限公司 | 识别模型的负样本构建方法、装置和系统 |
| CN110348523A (zh) * | 2019-07-15 | 2019-10-18 | 北京信息科技大学 | 一种基于Stacking的恶意网页集成识别方法及系统 |
| CN111667028B (zh) * | 2020-07-09 | 2024-03-12 | 腾讯科技(深圳)有限公司 | 一种可靠负样本确定方法和相关装置 |
| CN111667028A (zh) * | 2020-07-09 | 2020-09-15 | 腾讯科技(深圳)有限公司 | 一种可靠负样本确定方法和相关装置 |
| CN111931848A (zh) * | 2020-08-10 | 2020-11-13 | 中国平安人寿保险股份有限公司 | 数据的特征提取方法、装置、计算机设备及存储介质 |
| CN112350956A (zh) * | 2020-10-23 | 2021-02-09 | 新华三大数据技术有限公司 | 一种网络流量识别方法、装置、设备及机器可读存储介质 |
| CN112350956B (zh) * | 2020-10-23 | 2022-07-01 | 新华三大数据技术有限公司 | 一种网络流量识别方法、装置、设备及机器可读存储介质 |
| CN112561082A (zh) * | 2020-12-22 | 2021-03-26 | 北京百度网讯科技有限公司 | 生成模型的方法、装置、设备以及存储介质 |
| CN114067415A (zh) * | 2021-11-26 | 2022-02-18 | 北京百度网讯科技有限公司 | 回归模型的训练方法、对象评估方法、装置、设备和介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| TWI739798B (zh) | 2021-09-21 |
| US11551036B2 (en) | 2023-01-10 |
| US20180365522A1 (en) | 2018-12-20 |
| CN107133628A (zh) | 2017-09-05 |
| TW201732662A (zh) | 2017-09-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017143919A1 (zh) | 一种建立数据识别模型的方法及装置 | |
| CN111626821B (zh) | 基于集成特征选择实现客户分类的产品推荐方法及系统 | |
| WO2017133492A1 (zh) | 一种风险评估方法和系统 | |
| WO2017143921A1 (zh) | 一种多重抽样模型训练方法及装置 | |
| CN111583012B (zh) | 融合文本信息的信用债发债主体违约风险评估方法 | |
| KR102362872B1 (ko) | 인공지능 학습을 위한 클린 라벨 데이터 정제 방법 | |
| CN108319672B (zh) | 基于云计算的移动终端不良信息过滤方法及系统 | |
| CN106485528A (zh) | 检测数据的方法和装置 | |
| CN111506798A (zh) | 用户筛选方法、装置、设备及存储介质 | |
| CN109902731B (zh) | 一种基于支持向量机的性能故障的检测方法及装置 | |
| CN109859199B (zh) | 一种sd-oct图像的淡水无核珍珠质量检测的方法 | |
| CN110930038A (zh) | 一种贷款需求识别方法、装置、终端及存储介质 | |
| CN119722294B (zh) | 信贷风险检测方法及装置、电子设备、程序产品 | |
| CN104850868A (zh) | 一种基于k-means和神经网络聚类的客户细分方法 | |
| CN119541701A (zh) | 一种基于小样本学习的食品安全风险评估的方法 | |
| CN114596152A (zh) | 基于无监督模型预测发债主体违约的方法、设备及存储介质 | |
| CN107016416B (zh) | 基于邻域粗糙集和pca融合的数据分类预测方法 | |
| CN111090833A (zh) | 一种数据处理方法、系统及相关设备 | |
| CN115330401A (zh) | 违规商户识别模型构建方法及装置、违规商户识别方法 | |
| CN120526222A (zh) | 一种流水图像检验方法、装置、设备及介质 | |
| CN113673595A (zh) | 一种数据处理方法、装置及设备 | |
| CN119338512A (zh) | 一种针对两融交易型客户的流失测算方法及装置 | |
| CN115510976A (zh) | 基于模糊样本分析的深度学习模型去偏方法 | |
| CN114398942A (zh) | 一种基于集成的个人所得税异常检测方法及装置 | |
| CN113590925A (zh) | 一种用户确定方法、装置、设备及计算机存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17755748 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17755748 Country of ref document: EP Kind code of ref document: A1 |