WO2022095352A1 - 基于智能决策的异常用户识别方法、装置及计算机设备 - Google Patents
基于智能决策的异常用户识别方法、装置及计算机设备 Download PDFInfo
- Publication number
- WO2022095352A1 WO2022095352A1 PCT/CN2021/090422 CN2021090422W WO2022095352A1 WO 2022095352 A1 WO2022095352 A1 WO 2022095352A1 CN 2021090422 W CN2021090422 W CN 2021090422W WO 2022095352 A1 WO2022095352 A1 WO 2022095352A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- samples
- unlabeled
- sample
- user identification
- user
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/50—Monitoring users, programs or devices to maintain the integrity of platforms, e.g. of processors, firmware or operating systems
- G06F21/55—Detecting local intrusion or implementing counter-measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/243—Classification techniques relating to the number of classes
- G06F18/24323—Tree-organised classifiers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q20/00—Payment architectures, schemes or protocols
- G06Q20/38—Payment protocols; Details thereof
- G06Q20/40—Authorisation, e.g. identification of payer or payee, verification of customer or shop credentials; Review and approval of payers, e.g. check credit lines or negative lists
- G06Q20/401—Transaction verification
- G06Q20/4016—Transaction verification involving fraud or risk level assessment in transaction processing
Definitions
- the present application relates to the technical field of artificial intelligence, and in particular, to a method, device, computer equipment and storage medium for identifying abnormal users based on intelligent decision-making.
- the rule model is organized into empirical rules based on the abnormal users that have been found, and is based on human subjective judgment, with poor coverage and low recognition accuracy.
- Blacklist recognition is to obtain blacklist data from the outside, and track and monitor abnormal users in the blacklist. Blacklist recognition cannot deal with new abnormal users that appear at any time, and the accuracy is still low.
- the purpose of the embodiments of the present application is to propose a method, device, computer equipment and storage medium for identifying abnormal users based on intelligent decision, so as to solve the problem of low accuracy of identifying abnormal users.
- the embodiment of the present application provides a method for identifying abnormal users based on intelligent decision-making, which adopts the following technical solutions:
- the original data set includes blacklist data, authentic user data and original user data
- the embodiment of the present application also provides an abnormal user identification device based on intelligent decision-making, which adopts the following technical solutions:
- a data set acquisition module configured to acquire an original data set, wherein the original data set includes blacklist data, authentic user data and original user data;
- a data reorganization module used for data reorganization of the original data set to obtain labeled samples and unlabeled samples
- a first training module configured to input the labeled samples into a first user identification model, so as to perform first training on the first user identification model through the labeled samples to obtain a second user identification model;
- a data enhancement module configured to perform data enhancement on the unlabeled samples to obtain an enhanced unlabeled sample set corresponding to the unlabeled samples
- the second training module is used to perform second training on the second user identification model through the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples to obtain an abnormal user identification model;
- the sample input module is used for inputting the user sample to be identified into the abnormal user identification model to obtain the user identification result.
- an embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the following steps when executing the computer-readable instructions:
- the original data set includes blacklist data, authentic user data and original user data
- the embodiments of the present application further provide a computer-readable storage medium, where the computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the following steps are implemented:
- the original data set includes blacklist data, authentic user data and original user data
- the embodiments of the present application mainly have the following beneficial effects: after obtaining the original data set, data recombination is performed through data comparison to obtain labeled samples and unlabeled samples; Carry out the first training to obtain a second user identification model with a certain abnormal user identification ability; perform data enhancement on unlabeled samples to obtain an enhanced unlabeled sample set, which is changed from the original prediction of one unlabeled sample to multiple similar unlabeled samples.
- the labeled samples are predicted to improve the generalization ability of the second user identification model; the second user identification model is comprehensively trained through the labeled samples and the enhanced unlabeled sample set, and the model further extracts information from the unlabeled samples for learning, and finally
- the abnormal user identification model is obtained, and the abnormal user identification model can accurately output the user identification result according to the user sample to be identified, thereby improving the accuracy of the abnormal user identification.
- FIG. 1 is an exemplary system architecture diagram to which the present application can be applied;
- Fig. 2 is a flowchart of an embodiment of a method for identifying abnormal users based on intelligent decision-making according to the present application
- Fig. 3 is a flow chart of a specific implementation manner of step S202 in Fig. 2;
- Fig. 4 is a flow chart of a specific implementation manner of step S2023 in Fig. 3;
- Fig. 5 is a flow chart of a specific implementation manner of step S205 in Fig. 2;
- FIG. 6 is a schematic structural diagram of an embodiment of an abnormal user identification device based on intelligent decision-making according to the present application
- FIG. 7 is a schematic structural diagram of an embodiment of a computer device according to the present application.
- the system architecture 100 may include terminal devices 101 , 102 , and 103 , a network 104 and a server 105 .
- the network 104 is a medium used to provide a communication link between the terminal devices 101 , 102 , 103 and the server 105 .
- the network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, among others.
- the user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages and the like.
- Various communication client applications may be installed on the terminal devices 101 , 102 and 103 , such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, and the like.
- the terminal devices 101, 102, and 103 can be various electronic devices that have a display screen and support web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, dynamic Picture Experts Compression Standard Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) Players, Laptops and Desktops, etc.
- MP3 players Moving Picture Experts Group Audio Layer III, dynamic Picture Experts Compression Standard Audio Layer 3
- MP4 Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4
- the server 105 may be a server that provides various services, such as a background server that provides support for the pages displayed on the terminal devices 101 , 102 , and 103 .
- the intelligent decision-based abnormal user identification method provided by the embodiments of the present application is generally executed by the server, and accordingly, the intelligent decision-based abnormal user identification device is generally set in the server.
- terminal devices, networks and servers in FIG. 1 are merely illustrative. There can be any number of terminal devices, networks and servers according to implementation needs.
- FIG. 2 a flow chart of an embodiment of the method for identifying abnormal users based on intelligent decision-making according to the present application is shown.
- the described method for identifying abnormal users based on intelligent decision-making includes the following steps:
- Step S201 obtaining an original data set, wherein the original data set includes blacklist data, authentic user data and original user data.
- Abnormal user identification in this application involves intelligent decision-making in artificial intelligence.
- the electronic device for example, the server shown in FIG. 1
- the abnormal user identification method based on intelligent decision runs can communicate with the terminal through wired connection or wireless connection.
- the above wireless connection methods may include but are not limited to 3G/4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or developed in the future .
- the blacklist data can be the user data corresponding to the identified abnormal users;
- the authentic user data can be the user data that has passed the security authentication and is determined to be a non-abnormal user;
- the original user data can be the platform’s operation and production activities. The full amount of user data recorded.
- the server reads the original data set from the database, and the original data set includes blacklist data, authentic user data and original user data.
- the blacklist data may be acquired from the outside in advance and provided by a third-party data party.
- the platform will carry out strict identity authentication on some users during operation and production activities, and the user data corresponding to the users who have completed the identity authentication is the authentic user data.
- the blacklist data records the wool party determined by a third party, including virtual mobile phone numbers that cannot pass verification methods such as man-machine verification.
- Verified user data may be user data determined by the platform to be non-abnormal users through verification methods such as face recognition and binding bank cards.
- the above-mentioned original data set can also be stored in a node of a blockchain.
- the blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm.
- Blockchain essentially a decentralized database, is a series of data blocks associated with cryptographic methods. Each data block contains a batch of network transaction information to verify its Validity of information (anti-counterfeiting) and generation of the next block.
- the blockchain can include the underlying platform of the blockchain, the platform product service layer, and the application service layer.
- Step S202 performing data reorganization on the original data set to obtain labeled samples and unlabeled samples.
- compare the blacklist data with the user IDs of the original user data such as user names or mobile phone numbers
- compare the user IDs of the authentic user data and the original user data compare the user IDs of the authentic user data and the original user data , and the user corresponding to the duplicate user ID is added to the white sample.
- Black samples and white samples constitute labeled samples, and the data that has not been matched repeatedly in the original user data is regarded as unlabeled samples.
- the labeled samples and unlabeled samples also include the user data of the users in the samples.
- the label of the sample can identify whether the user is an abnormal user. For example, there are user A and user B in the labeled sample, the label of user A is 1, indicating that user A is an abnormal user; the label of user B is 0, indicating that user B is a non-abnormal user; user C has no label and cannot know user C. Whether it is an abnormal user or a non-exceptional user.
- Step S203 input the labeled samples into the first user identification model, so as to perform first training on the first user identification model through the labeled samples to obtain a second user identification model.
- the first user identification model may be a user identification model that has not completed the first training.
- the labeled samples are input into the first user identification model, the user data in the labeled samples is used as the model input, the sample labels are used as the expected output of the model, and the first user identification model is trained according to the model input and expected output (ie first training) to obtain a second user identification model.
- Step S204 performing data enhancement on the unlabeled samples to obtain an enhanced unlabeled sample set corresponding to the unlabeled samples.
- unlabeled samples are also added to model training.
- Unlabeled samples have no labels, which may bring large errors in training.
- data enhancement is performed on unlabeled samples, that is, similar data of unlabeled samples is generated, and the data scale of unlabeled samples is expanded. Get an enhanced unlabeled sample set.
- linear interpolation is used to obtain enhanced unlabeled samples:
- (a new ,b new ,...m new ) are enhanced unlabeled samples generated by interpolation, (a i ,b i ,...m i ) unlabeled samples, (a j ,b j ,... .,m j ,) is another randomly selected unlabeled sample, and the value of ⁇ ranges from 0 to 1.
- Step S205 performing second training on the second user identification model by using the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples to obtain an abnormal user identification model.
- Both the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples are input into the second user identification model.
- Each enhanced unlabeled sample in the enhanced unlabeled sample set has a user prediction result, and the user prediction result with the highest occurrence probability is used as the user prediction result of the unlabeled sample.
- the user prediction result of the unlabeled sample in the previous round of training is in In this round of training, it is used as a pseudo-label.
- the cross-entropy loss is calculated according to the user prediction results and labels of labeled samples, the user prediction results of unlabeled samples and pseudo-labels, and the model parameters are adjusted to reduce the cross-entropy loss until the model converges, and an abnormal user identification model is obtained.
- Step S206 input the user sample to be identified into the abnormal user identification model to obtain a user identification result.
- the server receives the user samples to be identified, inputs the user samples to be identified into the abnormal user identification model, and obtains a user identification result, which shows whether the user is an abnormal user.
- data recombination is performed through data comparison to obtain labeled samples and unlabeled samples; the labeled samples are input into the first user identification model for the first training to obtain a certain abnormal user identification ability.
- the second user identification model is based on the second user identification model; the unlabeled sample is enhanced by data enhancement to obtain an enhanced unlabeled sample set, and the original prediction of one unlabeled sample is changed to prediction of multiple similar unlabeled samples in order to improve the second user identification model.
- the second user identification model is comprehensively trained through the labeled samples and the enhanced unlabeled sample set, and the model further extracts information from the unlabeled samples for learning, and finally obtains the abnormal user identification model.
- the abnormal user identification model can be based on The user sample to be identified accurately outputs the user identification result, which improves the accuracy of identifying abnormal users.
- step S202 may include:
- Step S2021 compare the blacklist data and the authentic user data with the original user data, respectively, to determine a list of labeled users and an initial unlabeled sample.
- the user IDs are compared to determine the users whose blacklist data and verification user data are duplicated with the original user data, and obtain a list of tagged users; Unlabeled samples.
- Step S2022 filling the labeled user list with data according to the original data set to obtain an initial labeled sample.
- the tagged user list includes black users and white users.
- the black users are obtained by comparing the blacklist data with the original user data
- the white users are obtained by comparing the authentic user data with the original user data.
- the server reads the characteristics of each dimension in the blacklist data and the original user data of the black user, and adds the characteristics of each dimension to the list of labeled users; reads the real user data of the white user and the original user data in each dimension of the data. feature, add the features of each dimension to the list of labeled users, and get the initial labeled samples. Missing features can be filled with features; features with conflicting data are subject to blacklist data or authentic user data.
- Step S2023 Perform feature screening on the initial labeled samples and the initial unlabeled samples to obtain labeled samples and unlabeled samples.
- the feature dimensions of the initial labeled samples and the initial unlabeled samples are many, and the features of the same dimension can be screened from the initial labeled samples and the initial unlabeled samples to obtain the labeled samples and the unlabeled samples.
- the filtered features may include the number of occurrences of the terminal identifier of the user terminal in the write-off record within a preset time, and the activity of the network address of the user terminal within the preset time. Number of times, write-off time, service type, settlement price, etc.
- step S2023 may include:
- Step S20231 Input the initial labeled samples into the first user identification model, so as to perform third training on the first user identification model through the initial labeled samples to obtain a third user identification model.
- the initial labeled samples and the initial unlabeled samples contain full-dimensional features
- the initial labeled samples are input into the first user identification model
- the first user identification model is trained from the full features to obtain a third user identification model.
- Step S20232 Input the initial unlabeled sample into the third user identification model to obtain a pseudo-label of the initial unlabeled sample.
- the initial unlabeled samples are input into the third user identification model for identification processing, and the pseudo-labels of the initial unlabeled samples are obtained.
- Feature screening in this application requires labels, so pseudo-labels need to be added to the initial unlabeled samples first.
- Step S20233 Perform feature screening on the initial labeled samples and the initial unlabeled samples with pseudo-labels through random forest to obtain labeled samples and unlabeled samples, and determine the screened features as target features.
- the feature contribution degree of each feature is calculated by random forest.
- the feature contribution degree measures the importance of the feature.
- a preset number of features are selected, and the initial labeled samples and the initial unlabeled samples are not deleted.
- the data to the feature is deleted, and the labeled samples and unlabeled samples are obtained.
- pseudo-labels are first added to the initial unlabeled samples, so as to screen important features and obtain labeled samples and unlabeled samples, which ensures the smooth implementation of model training.
- step S20233 may include: taking the initial labeled samples and the initial unlabeled samples with pseudo-labels as the samples to be screened, and performing random sampling with replacement for several times to obtain several feature screening training sets; screening training sets based on several features , generate several decision trees to obtain a random forest; calculate the first out-of-bag data error of each decision tree in the random forest according to the out-of-bag data, wherein the out-of-bag data comes from the feature screening training set corresponding to each decision tree; randomly change the out-of-bag data feature in the data, and calculate the second out-of-bag data error of each decision tree; calculate the feature contribution degree of each feature according to the calculated second out-of-bag data error and the first out-of-bag data error; according to the calculated feature contribution degree Perform feature screening on initial labeled samples and initial unlabeled samples with pseudo-labels to obtain labeled samples and unlabeled samples, and determine the filtered features as target features.
- both the initial labeled samples and the initial unlabeled samples with pseudo-labels will be randomly sampled several times as labeled samples to be screened, and the characteristics of the samples can be randomly sampled after each sampling.
- the random sampling with replacement of the samples to be screened may be booststrapping sampling.
- Booststrapping sampling refers to sampling the original sample multiple times with replacement, and each sampling obtains a new sample, and after repeating the operation for many times Obtain multiple new samples, which can represent the sample distribution of the original samples.
- the training set is screened for each feature, and a decision tree is generated respectively, and the generated K decision trees constitute a random forest.
- a full split is performed based on information gain/information gain ratio/Gini index.
- a random forest is established, the feature contribution degree of each feature is calculated, and important features are screened out according to the feature contribution degree, and labeled samples and unlabeled samples are obtained, so that the model can perform targeted training on important features, improving the performance of the model. training efficiency.
- the above step S204 may include: for each unlabeled sample, determining the adjacent sample set of the unlabeled sample according to the Euclidean distance between the unlabeled samples, wherein the adjacent sample set includes a preset number of adjacent samples; for each For the adjacent samples, the extended sample points are selected on the feature space connection line between the adjacent samples and the unlabeled samples; according to the selected extended sample points and the unlabeled samples, an enhanced unlabeled sample set corresponding to the unlabeled samples is constructed.
- the unlabeled samples can be regarded as points in the feature space, and the dimension of the feature space is the same as that of the unlabeled samples. For each unlabeled sample, determine the Euclidean distance between the unlabeled sample and other unlabeled samples, sort the Euclidean distance from small to large, select a preset number of unlabeled samples, and obtain the adjacent sample set. Unlabeled samples can be regarded as adjacent samples of the original unlabeled samples.
- the feature dimension of the unlabeled sample is m
- (a new ,b new ,...,m new ) is the coordinate of the expanded sample point in the feature space
- (a,b,...,m) is the unlabeled sample
- the coordinates in the feature space, a n , b n , ..., m n represent the coordinates of the adjacent samples in each dimension in the feature space
- rand(0-1) is the adjustment factor, which adjusts the expansion of the sample points to the unlabeled samples.
- the expanded sample corresponding to the unlabeled sample is obtained according to the coordinates of the expanded sample point in the feature space.
- the unlabeled sample and the corresponding expanded sample can be used as the enhanced unlabeled sample, and the combination is Enhanced unlabeled sample set.
- the adjacent samples of the unlabeled samples are determined according to the Euclidean distance in the feature space, and the expanded sample points are generated according to the unlabeled samples and the adjacent samples. enhanced.
- step S205 may include:
- Step S2051 input the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples into the second user identification model, to obtain the user prediction results of the labeled samples, and the user prediction results of each enhanced unlabeled sample in the enhanced unlabeled sample set .
- the server inputs the labeled samples and the enhanced unlabeled sample set into the second user identification model, and obtains the user prediction result of the labeled samples; there are multiple enhanced unlabeled samples in the enhanced unlabeled sample set, and each enhanced unlabeled sample is There are corresponding user prediction results.
- Step S2052 Determine the user prediction result of the unlabeled sample according to the user prediction result of each enhanced unlabeled sample.
- the user prediction results of the user prediction results of the enhanced unlabeled samples are classified, and a category of user prediction results with the highest frequency is used as the user prediction results of the unlabeled samples corresponding to the enhanced unlabeled sample set.
- Step S2053 the user prediction result of the unlabeled sample in the second training in the previous round is used as the pseudo-label of the unlabeled sample in the current second training to calculate the regularized cross-entropy loss of the labeled sample and the unlabeled sample.
- the regularized cross-entropy loss is the loss function of the second user identification model.
- the second training consists of multiple rounds of training, and each round of training outputs user prediction results of unlabeled samples.
- the user prediction result of the unlabeled sample in the second training of the previous round is used as the pseudo-label of the unlabeled sample.
- the time-varying parameters are as follows:
- T 1 and T 2 represent the training rounds of the second training
- ⁇ f is the maximum value of the time-varying parameter. It can be seen from the time-varying parameters that the information extracted by the second user identification model from the unlabeled samples is gradually enhanced, which corresponds to the gradual improvement of the identification accuracy of the second user identification model with the deepening of training, which ensures the final abnormal user identification. accuracy of the model.
- Step S2054 adjust the parameters of the second user identification model according to the regularized cross-entropy loss until the model converges, and obtain an abnormal user identification model.
- the server adjusts the model parameters with the goal of minimizing the regularized cross-entropy loss until the second user identification model converges, and an abnormal user identification model is obtained.
- the user identification model of the present application is built based on the LGBM algorithm.
- LGBM LightBGM
- LGBM is an optimization framework that implements the GBDT algorithm. Its main idea is to use weak classifiers (decision trees) to iteratively train to obtain the optimal model. LGBM traverses each feature through multiple rounds of iteration, and then traverses all possible segmentation points for each feature to find the optimal segmentation point j of the optimal feature m, and each iteration generates a weak classifier based on a decision tree. , each classifier is trained on the residuals of the previous round of classifiers. Weak classifiers need to satisfy low variance and high bias. The process of LGBM algorithm training is to continuously improve the accuracy of the final classifier by reducing the bias.
- the enhanced unlabeled sample set is input into the second user identification model to obtain the user identification result of the unlabeled sample
- the regularized cross-entropy loss is calculated in combination with the user identification result of the labeled sample
- the model parameters are adjusted according to the loss, so that the first
- the second user identification model is further trained according to unlabeled samples, which ensures the accuracy of the obtained abnormal user identification model.
- step S206 may include: acquiring a user sample to be identified; performing feature screening on the user sample to be identified according to a preset target feature; inputting the user sample to be identified after feature screening into an abnormal user identification model to obtain a user identification result.
- the sample to be identified can be input by the user at the terminal.
- the target feature is determined according to the feature contribution degree, and the feature screening is performed on the user samples to be identified according to the target feature, and the features other than the target feature are removed.
- the user samples to be identified after feature screening are input into the abnormal user identification model to obtain the user identification result.
- the samples are firstly screened according to the preset target features to obtain samples whose feature dimensions conform to the model, which ensures the accuracy of the user identification results.
- the intelligent decision-based abnormal user identification method in this application involves machine learning and predictive analysis in the field of artificial intelligence, and may also involve fraud detection in financial technology.
- the computer-readable instructions can be stored in a computer-readable storage medium.
- the computer-readable instructions when executed, may include the processes of the above-mentioned method embodiments.
- the aforementioned storage medium may be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (Read-Only Memory, ROM), or a random access memory (Random Access Memory, RAM) or the like.
- the present application provides an embodiment of an abnormal user identification device 300 based on intelligent decision-making, which is similar to the method embodiment shown in FIG. 2 .
- the apparatus can be specifically applied to various electronic devices.
- the intelligent decision-based abnormal user identification device 300 in this embodiment includes: a data set acquisition module 301 , a data reorganization module 302 , a first training module 303 , a data enhancement module 304 , and a second training module 305 and sample input module 306, where:
- the data set obtaining module 301 is configured to obtain an original data set, wherein the original data set includes blacklist data, authentic user data and original user data.
- the data reorganization module 302 is used to reorganize the original data set to obtain labeled samples and unlabeled samples.
- the first training module 303 is configured to input the labeled samples into the first user identification model, so as to perform the first training on the first user identification model through the labeled samples to obtain the second user identification model.
- the data enhancement module 304 is configured to perform data enhancement on the unlabeled samples to obtain an enhanced unlabeled sample set corresponding to the unlabeled samples.
- the second training module 305 is configured to perform second training on the second user identification model by using the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples to obtain an abnormal user identification model.
- the sample input module 306 is configured to input the user sample to be identified into the abnormal user identification model to obtain the user identification result.
- the labeled samples are input into the first user identification model for first training, and the identification of users with certain abnormality is obtained.
- the second user identification model with the ability; data enhancement of unlabeled samples to enhance the unlabeled sample set, from the original prediction of one unlabeled sample to the prediction of multiple similar unlabeled samples, in order to improve the second user identification
- the generalization ability of the model; the second user identification model is comprehensively trained through the labeled samples and the enhanced unlabeled sample set, and the model further extracts information from the unlabeled samples for learning, and finally obtains the abnormal user identification model.
- the abnormal user identification model can The user identification result is accurately output according to the user sample to be identified, and the accuracy of abnormal user identification is improved.
- the data reorganization module 302 includes: a data comparison submodule, a data filling submodule, and a feature screening submodule, wherein:
- the data comparison sub-module is used to compare the blacklist data and the authentic user data with the original user data respectively to determine the list of labeled users and the initial unlabeled samples.
- the data filling sub-module is used to fill the labeled user list with data according to the original data set to obtain the initial labeled samples.
- the feature screening submodule is used to perform feature screening on the initial labeled samples and the initial unlabeled samples to obtain labeled samples and unlabeled samples.
- the feature screening sub-module includes: a training unit, an input unit, and a screening unit, wherein:
- the training unit is used for inputting the initial labeled samples into the first user identification model, so as to perform third training on the first user identification model through the initial labeled samples to obtain a third user identification model.
- the input unit is used to input the initial unlabeled sample into the third user identification model to obtain the pseudo-label of the initial unlabeled sample.
- the screening unit is used to perform feature screening on the initial labeled samples and the initial unlabeled samples with pseudo-labels through random forest to obtain labeled samples and unlabeled samples, and determine the filtered features as target features.
- pseudo-labels are first added to the initial unlabeled samples, so as to screen important features and obtain labeled samples and unlabeled samples, which ensures the smooth implementation of model training.
- the screening unit includes: a sampling subunit, a generation subunit, a first calculation subunit, a second calculation subunit, a contribution calculation subunit, and a feature screening subunit, wherein:
- the sampling sub-unit is used to take the initial labeled samples and the initial unlabeled samples with pseudo-labels as samples to be screened for several times with replacement and random sampling to obtain several feature screening training sets.
- Generating subunits is used to filter the training set based on several features, and generate several decision trees to obtain random forests.
- the first calculation subunit is used to calculate the first out-of-bag data error of each decision tree in the random forest according to the out-of-bag data, wherein the out-of-bag data comes from the feature screening training set corresponding to each decision tree.
- the second calculation subunit is used to randomly change the features in the out-of-bag data and calculate the second out-of-bag data error of each decision tree.
- the contribution calculation subunit is used to calculate the feature contribution degree of each feature according to the second out-of-bag data error and the first out-of-bag data error obtained by calculation.
- the feature screening subunit is used to perform feature screening on the initial labeled samples and the initial unlabeled samples with pseudo-labels according to the calculated feature contribution, to obtain labeled samples and unlabeled samples, and determine the filtered features as target characteristics.
- a random forest is established, the feature contribution degree of each feature is calculated, and important features are screened out according to the feature contribution degree, and labeled samples and unlabeled samples are obtained, so that the model can perform targeted training on important features, improving the performance of the model. training efficiency.
- the data enhancement module 303 includes: a sample determination submodule, a sample point selection submodule, and a sample set construction submodule, wherein:
- the sample determination sub-module is configured to, for each unlabeled sample, determine the adjacent sample set of the unlabeled sample according to the Euclidean distance between the unlabeled samples, wherein the adjacent sample set includes a preset number of adjacent samples.
- the sample point selection sub-module is used for each adjacent sample to select the extended sample point on the feature space connecting line between the adjacent sample and the unlabeled sample.
- the sample set construction sub-module is used to construct an enhanced unlabeled sample set corresponding to the unlabeled samples according to the selected expanded sample points and unlabeled samples.
- the adjacent samples of the unlabeled samples are determined according to the Euclidean distance in the feature space, and the expanded sample points are generated according to the unlabeled samples and the adjacent samples. enhanced.
- the second training module 304 includes: a sample input sub-module, a result determination sub-module, a loss calculation sub-module, and a parameter adjustment sub-module, wherein:
- the sample input sub-module is used to input the labeled samples and the enhanced unlabeled sample set corresponding to the unlabeled samples into the second user identification model to obtain the user prediction results of the labeled samples, and the enhanced unlabeled samples in the enhanced unlabeled sample set. user prediction results.
- the result determination sub-module is used for determining the user prediction result of the unlabeled sample according to the user prediction result of each enhanced unlabeled sample.
- the loss calculation submodule is used to use the user prediction results of the unlabeled samples in the second training round as the pseudo-label of the unlabeled samples in the current second training to calculate the regularized cross-entropy loss of the labeled samples and the unlabeled samples. .
- the parameter adjustment sub-module is used to adjust the parameters of the second user identification model according to the regularized cross-entropy loss, until the model converges, and an abnormal user identification model is obtained.
- the enhanced unlabeled sample set is input into the second user identification model to obtain the user identification result of the unlabeled sample
- the regularized cross-entropy loss is calculated in combination with the user identification result of the labeled sample
- the model parameters are adjusted according to the loss, so that the first
- the second user identification model is further trained according to unlabeled samples, which ensures the accuracy of the obtained abnormal user identification model.
- the sample input module 306 includes: a sample acquisition sub-module, a screening sub-module and an identification input sub-module, wherein:
- the sample acquisition sub-module is used to acquire a sample of the user to be identified.
- the screening sub-module is used for feature screening of user samples to be identified according to preset target features.
- the identification input sub-module is used to input the user samples to be identified after feature screening into the abnormal user identification model to obtain the user identification result.
- the samples are firstly screened according to the preset target features to obtain samples whose feature dimensions conform to the model, which ensures the accuracy of the user identification results.
- FIG. 7 is a block diagram of the basic structure of a computer device according to this embodiment.
- the computer device 4 includes a memory 41, a processor 42, and a network interface 43 that communicate with each other through a system bus. It should be pointed out that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all of the shown components, and that more or less components may be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculation and/or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, special-purpose Integrated circuit (Application Specific Integrated Circuit, ASIC), programmable gate array (Field-Programmable Gate Array, FPGA), digital processor (Digital Signal Processor, DSP), embedded equipment, etc.
- ASIC Application Specific Integrated Circuit
- FPGA Field-Programmable Gate Array
- DSP Digital Signal Processor
- the computer equipment may be a desktop computer, a notebook computer, a palmtop computer, a cloud server and other computing equipment.
- the computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touch pad or a voice control device.
- the memory 41 includes at least one type of computer-readable storage medium.
- the computer-readable storage medium may be non-volatile or volatile.
- the computer-readable storage medium includes flash memory, hard disk, and multimedia card. , card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), programmable Program read only memory (PROM), magnetic memory, magnetic disk, optical disk, etc.
- the memory 41 may be an internal storage unit of the computer device 4 , such as a hard disk or a memory of the computer device 4 .
- the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, flash memory card (Flash Card), etc.
- the memory 41 may also include both the internal storage unit of the computer device 4 and its external storage device.
- the memory 41 is generally used to store the operating system and various application software installed on the computer device 4 , such as computer-readable instructions for an abnormal user identification method based on intelligent decision-making.
- the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
- the processor 42 may be a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. This processor 42 is typically used to control the overall operation of the computer device 4 . In this embodiment, the processor 42 is configured to execute computer-readable instructions stored in the memory 41 or process data, for example, computer-readable instructions for executing the intelligent decision-based abnormal user identification method.
- CPU Central Processing Unit
- controller central processing unit
- microcontroller a microcontroller
- microprocessor microprocessor
- This processor 42 is typically used to control the overall operation of the computer device 4 .
- the processor 42 is configured to execute computer-readable instructions stored in the memory 41 or process data, for example, computer-readable instructions for executing the intelligent decision-based abnormal user identification method.
- the network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
- the computer device provided in this embodiment can execute the steps of the above-mentioned intelligent decision-based abnormal user identification method.
- the steps of the intelligent decision-based abnormal user identification method herein may be the steps in the intelligent decision-based abnormal user identification method of the above-mentioned various embodiments.
- data recombination is performed through data comparison to obtain labeled samples and unlabeled samples; the labeled samples are input into the first user identification model for the first training to obtain a certain abnormal user identification ability.
- the second user identification model is based on the second user identification model; the unlabeled sample is enhanced by data enhancement to obtain an enhanced unlabeled sample set, and the original prediction of one unlabeled sample is changed to prediction of multiple similar unlabeled samples in order to improve the second user identification model.
- the second user identification model is comprehensively trained through the labeled samples and the enhanced unlabeled sample set, and the model further extracts information from the unlabeled samples for learning, and finally obtains the abnormal user identification model.
- the abnormal user identification model can be based on The user sample to be identified accurately outputs the user identification result, which improves the accuracy of identifying abnormal users.
- the present application also provides another embodiment, that is, to provide a computer-readable storage medium, where the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to The at least one processor is caused to perform the steps of the above-mentioned intelligent decision-based abnormal user identification method.
- data recombination is performed through data comparison to obtain labeled samples and unlabeled samples; the labeled samples are input into the first user identification model for the first training to obtain a certain abnormal user identification ability.
- the second user identification model is based on the second user identification model; the unlabeled sample is enhanced by data enhancement to obtain an enhanced unlabeled sample set, and the original prediction of one unlabeled sample is changed to prediction of multiple similar unlabeled samples in order to improve the second user identification model.
- the second user identification model is comprehensively trained through the labeled samples and the enhanced unlabeled sample set, and the model further extracts information from the unlabeled samples for learning, and finally obtains the abnormal user identification model.
- the abnormal user identification model can be based on The user sample to be identified accurately outputs the user identification result, which improves the accuracy of identifying abnormal users.
- the method of the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is better implementation.
- the technical solution of the present application can be embodied in the form of a software product in essence or in a part that contributes to the prior art, and the computer software product is stored in a storage medium (such as ROM/RAM, magnetic disk, CD-ROM), including several instructions to make a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) execute the methods described in the various embodiments of this application.
- a storage medium such as ROM/RAM, magnetic disk, CD-ROM
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- Business, Economics & Management (AREA)
- Computer Security & Cryptography (AREA)
- General Engineering & Computer Science (AREA)
- Bioinformatics & Computational Biology (AREA)
- Accounting & Taxation (AREA)
- Evolutionary Biology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Evolutionary Computation (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computer Hardware Design (AREA)
- Finance (AREA)
- Strategic Management (AREA)
- General Business, Economics & Management (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
一种基于智能决策的异常用户识别方法、装置、计算机设备及存储介质,涉及人工智能领域,该方法包括:获取原始数据集;对原始数据集进行数据重组,得到有标签样本和无标签样本;将有标签样本输入第一用户识别模型,以对第一用户识别模型进行第一训练,得到第二用户识别模型;对无标签样本进行数据增强,得到与无标签样本对应的增强无标签样本集;通过有标签样本以及与无标签样本对应的增强无标签样本集,对第二用户识别模型进行第二训练,得到异常用户识别模型;将待识别用户样本输入异常用户识别模型,得到用户识别结果。此外,该方法还涉及区块链技术,原始数据集可存储于区块链中。该方法提高了异常用户识别的准确性。
Description
本申请要求于2020年11月03日提交中国专利局、申请号为202011211553.3,发明名称为“基于智能决策的异常用户识别方法、装置及计算机设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及人工智能技术领域,尤其涉及一种基于智能决策的异常用户识别方法、装置、计算机设备及存储介质。
随着互联网技术的发展,越来越多的用户通过互联网获取、享受各种信息服务,而提供信息服务的平台会记录得到大量的用户信息。提供信息服务的平台经常会遇到各种异常用户,例如羊毛党,羊毛党会利用虚假信息获取大量利益,给平台带来巨大损失,同时,还可能出现异常用户进行网络欺诈以及网络攻击,因此平台需要能够对这些异常用户进行识别。
然而,发明人意识到,传统的异常用户识别技术,通常是通过规则模型或黑名单进行识别。规则模型是基于已发现的异常用户整理成经验性规则,是以人的主观判断为基准,覆盖性差,识别的准确性较低。黑名单识别是从外部获取黑名单数据,对黑名单中出现的异常用户进行跟踪和监测,黑名单识别无法应对随时出现的新异常用户,准确性依然较低。
发明内容
本申请实施例的目的在于提出一种基于智能决策的异常用户识别方法、装置、计算机设备及存储介质,以解决异常用户识别准确性较低的问题。
为了解决上述技术问题,本申请实施例提供一种基于智能决策的异常用户识别方法,采用了如下所述的技术方案:
获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;
对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;
将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;
对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;
通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;
将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
为了解决上述技术问题,本申请实施例还提供一种基于智能决策的异常用户识别装置,采用了如下所述的技术方案:
数据集获取模块,用于获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;
数据重组模块,用于对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;
第一训练模块,用于将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;
数据增强模块,用于对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;
第二训练模块,用于通过所述有标签样本以及与所述无标签样本对应的增强无标签样 本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;
样本输入模块,用于将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
为了解决上述技术问题,本申请实施例还提供一种计算机设备,包括存储器和处理器,所述存储器中存储有计算机可读指令,所述处理器执行所述计算机可读指令时实现如下步骤:
获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;
对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;
将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;
对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;
通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;
将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
为了解决上述技术问题,本申请实施例还提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,所述计算机可读指令被处理器执行时实现如下步骤:
获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;
对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;
将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;
对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;
通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;
将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
与现有技术相比,本申请实施例主要有以下有益效果:获取原始数据集后,通过数据比对进行数据重组得到有标签样本以及无标签样本;将有标签样本输入第一用户识别模型以进行第一训练,得到具有一定异常用户识别能力的第二用户识别模型;对无标签样本进行数据增强得到增强无标签样本集,由原本对一个无标签样本的预测改为对多个相似的无标签样本进行预测,以便提升第二用户识别模型的泛化能力;通过有标签样本和增强无标签样本集对第二用户识别模型进行综合训练,模型进一步从无标签样本中提取信息进行学习,最终得到异常用户识别模型,异常用户识别模型能够根据待识别用户样本准确输出用户识别结果,提高了异常用户识别的准确性。
为了更清楚地说明本申请中的方案,下面将对本申请实施例描述中所需要使用的附图作一个简单介绍,显而易见地,下面描述中的附图是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是本申请可以应用于其中的示例性系统架构图;
图2是根据本申请的基于智能决策的异常用户识别方法的一个实施例的流程图;
图3是图2中步骤S202的一种具体实施方式的流程图;
图4是图3中步骤S2023的一种具体实施方式的流程图;
图5是图2中步骤S205的一种具体实施方式的流程图;
图6是根据本申请的基于智能决策的异常用户识别装置的一个实施例的结构示意图;
图7是根据本申请的计算机设备的一个实施例的结构示意图。
除非另有定义,本文所使用的所有的技术和科学术语与属于本申请的技术领域的技术人员通常理解的含义相同;本文中在申请的说明书中所使用的术语只是为了描述具体的实施例的目的,不是旨在于限制本申请;本申请的说明书和权利要求书及上述附图说明中的术语“包括”和“具有”以及它们的任何变形,意图在于覆盖不排他的包含。本申请的说明书和权利要求书或上述附图中的术语“第一”、“第二”等是用于区别不同对象,而不是用于描述特定顺序。
在本文中提及“实施例”意味着,结合实施例描述的特定特征、结构或特性可以包含在本申请的至少一个实施例中。在说明书中的各个位置出现该短语并不一定均是指相同的实施例,也不是与其它实施例互斥的独立的或备选的实施例。本领域技术人员显式地和隐式地理解的是,本文所描述的实施例可以与其它实施例相结合。
为了使本技术领域的人员更好地理解本申请方案,下面将结合附图,对本申请实施例中的技术方案进行清楚、完整地描述。
如图1所示,系统架构100可以包括终端设备101、102、103,网络104和服务器105。网络104用以在终端设备101、102、103和服务器105之间提供通信链路的介质。网络104可以包括各种连接类型,例如有线、无线通信链路或者光纤电缆等等。
用户可以使用终端设备101、102、103通过网络104与服务器105交互,以接收或发送消息等。终端设备101、102、103上可以安装有各种通讯客户端应用,例如网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
终端设备101、102、103可以是具有显示屏并且支持网页浏览的各种电子设备,包括但不限于智能手机、平板电脑、电子书阅读器、MP3播放器(Moving Picture Experts Group Audio Layer III,动态影像专家压缩标准音频层面3)、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器、膝上型便携计算机和台式计算机等等。
服务器105可以是提供各种服务的服务器,例如对终端设备101、102、103上显示的页面提供支持的后台服务器。
需要说明的是,本申请实施例所提供的基于智能决策的异常用户识别方法一般由服务器执行,相应地,基于智能决策的异常用户识别装置一般设置于服务器中。
应该理解,图1中的终端设备、网络和服务器的数目仅仅是示意性的。根据实现需要,可以具有任意数目的终端设备、网络和服务器。
继续参考图2,示出了根据本申请的基于智能决策的异常用户识别方法的一个实施例的流程图。所述的基于智能决策的异常用户识别方法,包括以下步骤:
步骤S201,获取原始数据集,其中,原始数据集包括黑名单数据、验真用户数据以及原始用户数据。
本申请中的异常用户识别涉及人工智能中的智能决策。在本实施例中,基于智能决策的异常用户识别方法运行于其上的电子设备(例如图1所示的服务器)可以通过有线连接方式或者无线连接方式与终端进行通信。需要指出的是,上述无线连接方式可以包括但不限于3G/4G连接、WiFi连接、蓝牙连接、WiMAX连接、Zigbee连接、UWB(ultra wideband)连接、以及其他现在已知或将来开发的无线连接方式。
其中,黑名单数据可以是已确定的异常用户所对应的用户数据;验真用户数据可以是已通过安全认证、确定为非异常用户的用户数据;原始用户数据可以是平台在经营、生产活动中记录的全量用户数据。
具体地,服务器从数据库中读取原始数据集,原始数据集中包括黑名单数据、验真用户数据以及原始用户数据。
在一个实施例中,黑名单数据可以预先从外部获取,由第三方数据方提供。平台在经营、生产活动中会对一些用户进行严格的身份认证,完成身份认证的用户所对应的用户数据即为验真用户数据。举例说明,在羊毛党识别的场景中,黑名单数据记录了第三方确定的羊毛党,包括了无法通过人机验证等验真方式的虚拟手机号码。验真用户数据可以是平台通过人脸识别、绑定银行卡等验真方式确定为非异常用户的用户数据。
需要强调的是,为进一步保证上述原始数据集的私密和安全性,上述原始数据集还可以存储于一区块链的节点中。
本申请所指区块链是分布式数据存储、点对点传输、共识机制、加密算法等计算机技术的新型应用模式。区块链(Blockchain),本质上是一个去中心化的数据库,是一串使用密码学方法相关联产生的数据块,每一个数据块中包含了一批次网络交易的信息,用于验证其信息的有效性(防伪)和生成下一个区块。区块链可以包括区块链底层平台、平台产品服务层以及应用服务层等。
步骤S202,对原始数据集进行数据重组,得到有标签样本以及无标签样本。
具体地,比对黑名单数据和原始用户数据的用户标识(例如用户名或者手机号码),将重复的用户标识所对应的用户加入黑样本;比对验真用户数据和原始用户数据的用户标识,将重复的用户标识所对应的用户加入白样本。
黑样本和白样本构成有标签样本,原始用户数据中未完成重复匹配的数据作为无标签样本,有标签样本和无标签样本还包括样本中用户的用户数据。
样本的标签可以标识用户是否为异常用户。例如,有标签样本中有用户A和用户B,用户A标签为1,表示用户A为异常用户;用户B标签为0,表示用户B为非异常用户;用户C没有标签,无法得知用户C是异常用户还是非异常用户。
步骤S203,将有标签样本输入第一用户识别模型,以通过有标签样本对第一用户识别模型进行第一训练,得到第二用户识别模型。
其中,第一用户识别模型可以是尚未完成第一训练的用户识别模型。
具体地,将有标签样本输入第一用户识别模型,有标签样本中的用户数据将作为模型输入,样本标签作为模型的期望输出,根据模型输入和期望输出对第一用户识别模型进行训练(即第一训练),得到第二用户识别模型。
步骤S204,对无标签样本进行数据增强,得到与无标签样本对应的增强无标签样本集。
具体地,无标签样本也要加入模型训练。无标签样本没有标签,在训练中可能带来较大的误差,为提高模型的泛化能力,对无标签样本进行数据增强,即生成无标签样本的相似数据,扩充无标签样本的数据规模,得到增强无标签样本集。
在一个实施例中,基于邻域风险最小化原则,使用线性插值得到增强无标签样本:
(a
new,b
new,...m
new)=λ(a
i,b
i,...m
i)+(1-λ)*(a
j,b
j,...,m
j,) (1)
其中,(a
new,b
new,...m
new)是插值生成的增强无标签样本,(a
i,b
i,...m
i)无标签样本,(a
j,b
j,...,m
j,)是随机选取的另一个无标签样本,λ取值取指范围介于0到1。
步骤S205,通过有标签样本以及与无标签样本对应的增强无标签样本集,对第二用户识别模型进行第二训练,得到异常用户识别模型。
将有标签样本以及与无标签样本对应的增强无标签样本集均输入第二用户识别模型。增强无标签样本集中每个增强无标签样本均有用户预测结果,将出现概率最高的一类用户预测结果作为无标签样本的用户预测结果,无标签样本在上一轮训练中的用户预测结果在本轮训练中作为伪标签。
根据有标签样本的用户预测结果和标签、无标签样本的用户预测结果和伪标签计算交叉熵损失,以减小交叉熵损失为目标调整模型参数直至模型收敛,得到异常用户识别模型。
步骤S206,将待识别用户样本输入异常用户识别模型,得到用户识别结果。
具体地,在模型应用时,服务器接收待识别用户样本,将待识别用户样本输入异常用户识别模型,得到用户识别结果,用户识别结果显示用户是否为异常用户。
本实施例中,获取原始数据集后,通过数据比对进行数据重组得到有标签样本以及无标签样本;将有标签样本输入第一用户识别模型以进行第一训练,得到具有一定异常用户识别能力的第二用户识别模型;对无标签样本进行数据增强得到增强无标签样本集,由原本对一个无标签样本的预测改为对多个相似的无标签样本进行预测,以便提升第二用户识别模型的泛化能力;通过有标签样本和增强无标签样本集对第二用户识别模型进行综合训练,模型进一步从无标签样本中提取信息进行学习,最终得到异常用户识别模型,异常用户识别模型能够根据待识别用户样本准确输出用户识别结果,提高了异常用户识别的准确性。
进一步的,如图3所示,上述步骤S202可以包括:
步骤S2021,将黑名单数据和验真用户数据分别与原始用户数据进行数据比对,以确定有标签用户列表及初始无标签样本。
具体地,比对用户标识,以确定黑名单数据和验证用户数据中与原始用户数据相重复的用户,得到有标签用户列表;原始用户数据中未实现重复匹配的用户所对应的用户数据作为初始无标签样本。
步骤S2022,根据原始数据集对有标签用户列表进行数据填充,得到初始有标签样本。
具体地,有标签用户列表包括黑用户以及白用户,黑用户由黑名单数据与原始用户数据比对得到,白用户由验真用户数据与原始用户数据比对得到。服务器读取黑用户在黑名单数据和原始用户数据中各维度的特征,将各维度的特征添加到有标签用户列表中;读取白用户在验真用户数据以及原始用户数据中每一维度的特征,将各维度的特征添加到有标签用户列表中,得到初始有标签样本。缺失的特征可以进行特征填充;数据冲突的特征以黑名单数据或验真用户数据为准。
步骤S2023,对初始有标签样本和初始无标签样本进行特征筛选,得到有标签样本以及无标签样本。
具体地,初始有标签样本和初始无标签样本特征维度较多,可以从初始有标签样本和初始无标签样本中筛选出相同维度的特征,得到有标签样本以及无标签样本。
例如,在卡券核销相关的羊毛党检测场景中,筛选到的特征可以包括核销记录中用户终端的终端标识在预设时间内出现次数、用户终端的网络地址在预设时间内的活跃次数、核销时间、服务类型、结算价格等。
本实施例中,在数据比对中通过确定重复用户和特征筛选,对原始数据集完成数据重组,得到用于模型训练的有标签样本和无标签样本。
进一步的,如图4所示,上述步骤S2023可以包括:
步骤S20231,将初始有标签样本输入第一用户识别模型,以通过初始有标签样本对第一用户识别模型进行第三训练,得到第三用户识别模型。
具体地,初始有标签样本和初始无标签样本包含全维度的特征,将初始有标签样本输入第一用户识别模型,从全特征对第一用户识别模型进行训练,得到第三用户识别模型。
步骤S20232,将初始无标签样本输入第三用户识别模型,得到初始无标签样本的伪标签。
具体地,将初始无标签样本输入第三用户识别模型进行识别处理,得到初始无标签样本的伪标签。本申请中的特征筛选需要标签,因此需要先给初始无标签样本添加伪标签。
步骤S20233,通过随机森林对初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
具体地,通过随机森林计算各特征的特征贡献度,特征贡献度衡量了特征的重要性,根据特征贡献度选取预设数量的特征,将初始有标签样本和初始无标签样本中未被删选到特征的数据删除,得到有标签样本和无标签样本。
本实施例中,先给初始无标签样本添加伪标签,以便筛选重要特征,得到有标签样本和无标签样本,保证了模型训练的顺利实现。
进一步的,上述步骤S20233可以包括:将初始有标签样本和带有伪标签的初始无标签样本作为待筛选样本进行若干次有放回随机采样,得到若干特征筛选训练集;基于若干特征筛选训练集,生成若干决策树以得到随机森林;根据袋外数据计算随机森林中各决策树的第一袋外数据误差,其中,袋外数据来自各决策树所对应的特征筛选训练集;随机改变袋外数据中的特征,并计算各决策树的第二袋外数据误差;根据计算得到的第二袋外数据误差和第一袋外数据误差计算各特征的特征贡献度;根据计算得到的特征贡献度对初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
具体地,初始有标签样本和带有伪标签的初始无标签样本都将作为有标签的待筛选样本进行若干次有放回随机采样,每次采样之后还可以再对样本的特征进行随机采样,得到若干特征筛选训练集。在一个实施例中,对待筛选样本的有放回随机采样可以是booststrapping采样,booststrapping采样是指对原样本进行多次有放回的抽样,每次抽样均得到一个新样本,重复操作多次后得到多个新样本,多个新样本可以代表原样本的样本分布。
针对每个特征筛选训练集,分别生成决策树,生成的K棵决策树构成随机森林。在生成每棵决策树时,根据信息增益/信息增益比/基尼指数进行完全分裂。
在根据特征筛选训练集建立决策树时,特征筛选训练集中有一部分样本并没有参与决策树的建立,这部分样本即为决策树的袋外数据,袋外数据通常用于评估决策树性能,计算预测错误率,即袋外数据误差。
将袋外数据输入决策树,根据分类结果和样本标签计算袋外数据误差,得到第一袋外数据误差error
1、error
2、...、error
K。随机改变袋外数据中特征的特征值再输入决策树,再计算袋外数据误差,得到第二袋外数据误差error
1'、error'
2、...error'
K;根据第二袋外数据误差和第一袋外数据误差计算各特征的特征贡献度:
根据特征贡献度对特征按降序排序,筛选预设数量的特征(或者剔除相应比例的特征,得到新的待筛选样本,用新的待筛选样本重复上述过程,直至得到最终预设数量的特征),根据筛选到的特征对初始有标签样本和初始无标签样本进行数据重组,留下筛选到的特征所对应的用户数据,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
本实施例中,建立随机森林并计算各特征的特征贡献度,根据特征贡献度进行特征筛选出重要特征,得到有标签样本以及无标签样本,使得模型可以对重要特征进行针对性训练,提高了训练效率。
进一步的,上述步骤S204可以包括:对于每个无标签样本,根据无标签样本间的欧氏距离确定无标签样本的临近样本集,其中,临近样本集包括预设数量的临近样本;对于每个临近样本,在临近样本与无标签样本的特征空间连线上,选取扩充样本点;根据选取的扩充样本点以及无标签样本,构建得到与无标签样本对应的增强无标签样本集。
具体地,无标签样本可视作特征空间中的点,特征空间的维度与无标签样本特征维度相同。对于每个无标签样本,确定无标签样本与其他无标签样本的欧氏距离,将欧氏距离从小到大进行排序,选取预设数量的无标签样本,得到临近样本集,临近样本集中的各无标签样本可视作原无标签样本的临近样本。
无标签样本与临近样本间存在特征空间连线,在特征空间连线上随机选取预设数量的点,得到扩充样本点:
(a
new,b
new,...m
new)=(a,b,...m)+rand(0-1)*((a
n-a,),(b
n-b,)...(m
n-m,)) (3)
其中,无标签样本特征维度为m,(a
new,b
new,...,m
new)是扩充样本点在特征空间中的坐标,(a,b,...,m)是无标签样本在特征空间中的坐标,a
n、b
n、...、m
n表示临近样本在特征空 间中各维度的坐标,rand(0-1)为调节因子,调节扩充样本点到无标签样本的距离
每个临近样本选取完扩充样本点后,根据扩充样本点在特征空间中的坐标得到与无标签样本对应的扩充样本,无标签样本以及与之对应的扩充样本可以作为增强无标签样本,组合为增强无标签样本集。
本实施例中,在特征空间中根据欧氏距离确定无标签样本的临近样本,根据无标签样本和临近样本生成扩充样本点,即可生成与无标签样本相似的多个扩充样本,实现了数据增强。
进一步的,如图5所示,上述步骤S205可以包括:
步骤S2051,将有标签样本以及与无标签样本对应的增强无标签样本集输入第二用户识别模型,得到有标签样本的用户预测结果,以及增强无标签样本集中各增强无标签样本的用户预测结果。
具体地,服务器将有标签样本和增强无标签样本集输入第二用户识别模型,得到有标签样本的用户预测结果;增强无标签样本集中有多个增强无标签样本,每个增强无标签样本均有对应的用户预测结果。
步骤S2052,根据各增强无标签样本的用户预测结果,确定无标签样本的用户预测结果。
具体地,对增强无标签样本的用户预测结果的用户预测结果进行分类,将频数最高的一类用户预测结果,作为与增强无标签样本集所对应的无标签样本的用户预测结果。
步骤S2053,将前轮第二训练中无标签样本的用户预测结果,作为当前第二训练中无标签样本的伪标签,以计算有标签样本和无标签样本的正则化交叉熵损失。
其中,正则化交叉熵损失为第二用户识别模型的损失函数。
具体地,第二训练由多轮训练构成,每轮训练均输出无标签样本的用户预测结果。在进行当前轮次的第二训练时,将前轮第二训练中无标签样本的用户预测结果作为无标签样本的伪标签。联合有标签样本的标签和用户预测结果,以及无标签样本的伪标签和用户预测结果,计算正则化交叉熵损失:
其中,
为有标签样本的交叉熵损失,
为有标签样本的样本标签,f
i
m为有标签样本的用户预测结果,n为有标签样本的样本数量;正则化项
为无标签样本的交叉熵损失;
为无标签样本的伪标签,
为无标签样本的用户预测结果,n'为无标签样本的样本数量,C为样本的类别数量,α(t)为时变参数。
在一个实施例中,时变参数如下:
其中,T
1和T
2表示第二训练的训练轮次,α
f为时变参数的最大值。由时变参数可知第二用户识别模型从无标签样本中提取到的信息逐渐增强,对应于随着训练的深入,第二用户识别模型的识别准确性逐渐提升,保证了最终得到的异常用户识别模型的准确性。
易知在第二训练的第一轮训练中,会输出无标签样本的用户预测结果,但不会计算正则化交叉熵损失。
步骤S2054,根据正则化交叉熵损失对第二用户识别模型进行参数调整,直至模型收敛,得到异常用户识别模型。
服务器以最小化正则化交叉熵损失为目标调整模型参数,直至第二用户识别模型收敛,得到异常用户识别模型。
在一个实施例中,本申请的用户识别模型基于LGBM算法搭建。LGBM(LightBGM)是一个实现GBDT算法的优化框架,其主要思想是利用弱分类器(决策树)迭代训练以得到最优模型。LGBM通过多轮迭代,遍历每个特征,然后对每个特征遍历它所有可能的切分点,找到最优特征m的最优切分点j,每轮迭代产生一个基于决策树的弱分类器,每个分类器在上一轮分类器的残差基础上进行训练。弱分类器要满足低方差和高偏差。LGBM算法训练的过程是通过降低偏差来不断提高最终分类器的精度。
本实施例中,将增强无标签样本集输入第二用户识别模型得到无标签样本的用户识别结果,结合有标签样本的用户识别结果计算正则化交叉熵损失,并根据损失调整模型参数,使第二用户识别模型根据无标签样本进一步训练,保证了得到的异常用户识别模型的准确性。
进一步的,上述步骤S206可以包括:获取待识别用户样本;根据预设的目标特征对待识别用户样本进行特征筛选;将特征筛选后的待识别用户样本输入异常用户识别模型,得到用户识别结果。
具体地,待识别样本可以由用户在终端输入。在特征筛选时根据特征贡献度确定了目标特征,根据目标特征对待识别用户样本进行特征筛选,去除目标特征以外的特征。再将特征筛选后的待识别用户样本输入异常用户识别模型,得到用户识别结果。
本实施例中,获取到待识别用户样本后,先根据预设的目标特征对样本进行特征筛选,得到特征维度符合模型的样本,保证了用户识别结果的准确性。
本申请中基于智能决策的异常用户识别方法涉及人工智能领域中的机器学习和预测分析,还可以涉及金融科技中的欺诈检测。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,该计算机可读指令可存储于一计算机可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,前述的存储介质可为磁碟、光盘、只读存储记忆体(Read-Only Memory,ROM)等非易失性存储介质,或随机存储记忆体(Random Access Memory,RAM)等。
应该理解的是,虽然附图的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,其可以以其他的顺序执行。而且,附图的流程图中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,其执行顺序也不必然是依次进行,而是可以与其他步骤或者其他步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
进一步参考图6,作为对上述图2所示方法的实现,本申请提供了一种基于智能决策的异常用户识别装置300的一个实施例,该装置实施例与图2所示的方法实施例相对应,该装置具体可以应用于各种电子设备中。
如图6所示,本实施例所述的基于智能决策的异常用户识别装置300包括:数据集获取模块301、数据重组模块302、第一训练模块303、数据增强模块304、第二训练模块305以及样本输入模块306,其中:
数据集获取模块301,用于获取原始数据集,其中,原始数据集包括黑名单数据、验真用户数据以及原始用户数据。
数据重组模块302,用于对原始数据集进行数据重组,得到有标签样本以及无标签样本。
第一训练模块303,用于将有标签样本输入第一用户识别模型,以通过有标签样本对第一用户识别模型进行第一训练,得到第二用户识别模型。
数据增强模块304,用于对无标签样本进行数据增强,得到与无标签样本对应的增强无标签样本集。
第二训练模块305,用于通过有标签样本以及与无标签样本对应的增强无标签样本集,对第二用户识别模型进行第二训练,得到异常用户识别模型。
样本输入模块306,用于将待识别用户样本输入异常用户识别模型,得到用户识别结果。
在本实施例中,获取原始数据集后,通过数据比对进行数据重组得到有标签样本以及无标签样本;将有标签样本输入第一用户识别模型以进行第一训练,得到具有一定异常用户识别能力的第二用户识别模型;对无标签样本进行数据增强得到增强无标签样本集,由原本对一个无标签样本的预测改为对多个相似的无标签样本进行预测,以便提升第二用户识别模型的泛化能力;通过有标签样本和增强无标签样本集对第二用户识别模型进行综合训练,模型进一步从无标签样本中提取信息进行学习,最终得到异常用户识别模型,异常用户识别模型能够根据待识别用户样本准确输出用户识别结果,提高了异常用户识别的准确性。
在本实施例的一些可选的实现方式中,数据重组模块302包括:数据比对子模块、数据填充子模块以及特征筛选子模块,其中:
数据比对子模块,用于将黑名单数据和验真用户数据分别与原始用户数据进行数据比对,以确定有标签用户列表及初始无标签样本。
数据填充子模块,用于根据原始数据集对有标签用户列表进行数据填充,得到初始有标签样本。
特征筛选子模块,用于对初始有标签样本和初始无标签样本进行特征筛选,得到有标签样本以及无标签样本。
本实施例中,在数据比对中通过确定重复用户和特征筛选,对原始数据集完成数据重组,得到用于模型训练的有标签样本和无标签样本。
在本实施例的一些可选的实现方式中,特征筛选子模块包括:训练单元、输入单元和筛选单元,其中:
训练单元,用于将初始有标签样本输入第一用户识别模型,以通过初始有标签样本对第一用户识别模型进行第三训练,得到第三用户识别模型。
输入单元,用于将初始无标签样本输入第三用户识别模型,得到初始无标签样本的伪标签。
筛选单元,用于通过随机森林对初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
本实施例中,先给初始无标签样本添加伪标签,以便筛选重要特征,得到有标签样本和无标签样本,保证了模型训练的顺利实现。
在本实施例的一些可选的实现方式中,筛选单元包括:采样子单元、生成子单元、第一计算子单元、第二计算子单元、贡献计算子单元和特征筛选子单元,其中:
采样子单元,用于将初始有标签样本和带有伪标签的初始无标签样本作为待筛选样本进行若干次有放回随机采样,得到若干特征筛选训练集。
生成子单元,用于基于若干特征筛选训练集,生成若干决策树以得到随机森林。
第一计算子单元,用于根据袋外数据计算随机森林中各决策树的第一袋外数据误差,其中,袋外数据来自各决策树所对应的特征筛选训练集。
第二计算子单元,用于随机改变袋外数据中的特征,并计算各决策树的第二袋外数据误差。
贡献计算子单元,用于根据计算得到的第二袋外数据误差和第一袋外数据误差计算各 特征的特征贡献度。
特征筛选子单元,用于根据计算得到的特征贡献度对初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
本实施例中,建立随机森林并计算各特征的特征贡献度,根据特征贡献度进行特征筛选出重要特征,得到有标签样本以及无标签样本,使得模型可以对重要特征进行针对性训练,提高了训练效率。
在本实施例的一些可选的实现方式中,数据增强模块303包括:样本确定子模块、样本点选取子模块以及样本集构建子模块,其中:
样本确定子模块,用于对于每个无标签样本,根据无标签样本间的欧氏距离确定无标签样本的临近样本集,其中,临近样本集包括预设数量的临近样本。
样本点选取子模块,用于对于每个临近样本,在临近样本与无标签样本的特征空间连线上,选取扩充样本点。
样本集构建子模块,用于根据选取的扩充样本点以及无标签样本,构建得到与无标签样本对应的增强无标签样本集。
本实施例中,在特征空间中根据欧氏距离确定无标签样本的临近样本,根据无标签样本和临近样本生成扩充样本点,即可生成与无标签样本相似的多个扩充样本,实现了数据增强。
在本实施例的一些可选的实现方式中,第二训练模块304包括:样本输入子模块、结果确定子模块、损失计算子模块以及参数调整子模块,其中:
样本输入子模块,用于将有标签样本以及与无标签样本对应的增强无标签样本集输入第二用户识别模型,得到有标签样本的用户预测结果,以及增强无标签样本集中各增强无标签样本的用户预测结果。
结果确定子模块,用于根据各增强无标签样本的用户预测结果,确定无标签样本的用户预测结果。
损失计算子模块,用于将前轮第二训练中无标签样本的用户预测结果,作为当前第二训练中无标签样本的伪标签,以计算有标签样本和无标签样本的正则化交叉熵损失。
参数调整子模块,用于根据正则化交叉熵损失对第二用户识别模型进行参数调整,直至模型收敛,得到异常用户识别模型。
本实施例中,将增强无标签样本集输入第二用户识别模型得到无标签样本的用户识别结果,结合有标签样本的用户识别结果计算正则化交叉熵损失,并根据损失调整模型参数,使第二用户识别模型根据无标签样本进一步训练,保证了得到的异常用户识别模型的准确性。
在本实施例的一些可选的实现方式中,样本输入模块306包括:样本获取子模块、筛选子模块以及识别输入子模块,其中:
样本获取子模块,用于获取待识别用户样本。
筛选子模块,用于根据预设的目标特征对待识别用户样本进行特征筛选。
识别输入子模块,用于将特征筛选后的待识别用户样本输入异常用户识别模型,得到用户识别结果。
本实施例中,获取到待识别用户样本后,先根据预设的目标特征对样本进行特征筛选,得到特征维度符合模型的样本,保证了用户识别结果的准确性。
为解决上述技术问题,本申请实施例还提供计算机设备。具体请参阅图7,图7为本实施例计算机设备基本结构框图。
所述计算机设备4包括通过系统总线相互通信连接存储器41、处理器42、网络接口43。需要指出的是,图中仅示出了具有组件41-43的计算机设备4,但是应理解的是,并 不要求实施所有示出的组件,可以替代的实施更多或者更少的组件。其中,本技术领域技术人员可以理解,这里的计算机设备是一种能够按照事先设定或存储的指令,自动进行数值计算和/或信息处理的设备,其硬件包括但不限于微处理器、专用集成电路(Application Specific Integrated Circuit,ASIC)、可编程门阵列(Field-Programmable Gate Array,FPGA)、数字处理器(Digital Signal Processor,DSP)、嵌入式设备等。
所述计算机设备可以是桌上型计算机、笔记本、掌上电脑及云端服务器等计算设备。所述计算机设备可以与用户通过键盘、鼠标、遥控器、触摸板或声控设备等方式进行人机交互。
所述存储器41至少包括一种类型的计算机可读存储介质,所述计算机可读存储介质可以是非易失性,也可以是易失性,所述计算机可读存储介质包括闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等。在一些实施例中,所述存储器41可以是所述计算机设备4的内部存储单元,例如该计算机设备4的硬盘或内存。在另一些实施例中,所述存储器41也可以是所述计算机设备4的外部存储设备,例如该计算机设备4上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。当然,所述存储器41还可以既包括所述计算机设备4的内部存储单元也包括其外部存储设备。本实施例中,所述存储器41通常用于存储安装于所述计算机设备4的操作系统和各类应用软件,例如基于智能决策的异常用户识别方法的计算机可读指令等。此外,所述存储器41还可以用于暂时地存储已经输出或者将要输出的各类数据。
所述处理器42在一些实施例中可以是中央处理器(Central Processing Unit,CPU)、控制器、微控制器、微处理器、或其他数据处理芯片。该处理器42通常用于控制所述计算机设备4的总体操作。本实施例中,所述处理器42用于运行所述存储器41中存储的计算机可读指令或者处理数据,例如运行所述基于智能决策的异常用户识别方法的计算机可读指令。
所述网络接口43可包括无线网络接口或有线网络接口,该网络接口43通常用于在所述计算机设备4与其他电子设备之间建立通信连接。
本实施例中提供的计算机设备可以执行上述基于智能决策的异常用户识别方法的步骤。此处基于智能决策的异常用户识别方法的步骤可以是上述各个实施例的基于智能决策的异常用户识别方法中的步骤。
本实施例中,获取原始数据集后,通过数据比对进行数据重组得到有标签样本以及无标签样本;将有标签样本输入第一用户识别模型以进行第一训练,得到具有一定异常用户识别能力的第二用户识别模型;对无标签样本进行数据增强得到增强无标签样本集,由原本对一个无标签样本的预测改为对多个相似的无标签样本进行预测,以便提升第二用户识别模型的泛化能力;通过有标签样本和增强无标签样本集对第二用户识别模型进行综合训练,模型进一步从无标签样本中提取信息进行学习,最终得到异常用户识别模型,异常用户识别模型能够根据待识别用户样本准确输出用户识别结果,提高了异常用户识别的准确性。
本申请还提供了另一种实施方式,即提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,所述计算机可读指令可被至少一个处理器执行,以使所述至少一个处理器执行如上述的基于智能决策的异常用户识别方法的步骤。
本实施例中,获取原始数据集后,通过数据比对进行数据重组得到有标签样本以及无标签样本;将有标签样本输入第一用户识别模型以进行第一训练,得到具有一定异常用户识别能力的第二用户识别模型;对无标签样本进行数据增强得到增强无标签样本集,由原本对一个无标签样本的预测改为对多个相似的无标签样本进行预测,以便提升第二用户识 别模型的泛化能力;通过有标签样本和增强无标签样本集对第二用户识别模型进行综合训练,模型进一步从无标签样本中提取信息进行学习,最终得到异常用户识别模型,异常用户识别模型能够根据待识别用户样本准确输出用户识别结果,提高了异常用户识别的准确性。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,空调器,或者网络设备等)执行本申请各个实施例所述的方法。
显然,以上所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例,附图中给出了本申请的较佳实施例,但并不限制本申请的专利范围。本申请可以以许多不同的形式来实现,相反地,提供这些实施例的目的是使对本申请的公开内容的理解更加透彻全面。尽管参照前述实施例对本申请进行了详细的说明,对于本领域的技术人员来而言,其依然可以对前述各具体实施方式所记载的技术方案进行修改,或者对其中部分技术特征进行等效替换。凡是利用本申请说明书及附图内容所做的等效结构,直接或间接运用在其他相关的技术领域,均同理在本申请专利保护范围之内。
Claims (20)
- 一种基于智能决策的异常用户识别方法,包括下述步骤:获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 根据权利要求1所述的基于智能决策的异常用户识别方法,其中,所述对所述原始数据集进行数据重组,得到有标签样本以及无标签样本的步骤包括:将所述黑名单数据和所述验真用户数据分别与所述原始用户数据进行数据比对,以确定有标签用户列表及初始无标签样本;根据所述原始数据集对所述有标签用户列表进行数据填充,得到初始有标签样本;对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无标签样本。
- 根据权利要求1所述的基于智能决策的异常用户识别方法,其中,所述对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无标签样本的步骤具体包括:将所述初始有标签样本输入第一用户识别模型,以通过所述初始有标签样本对所述第一用户识别模型进行第三训练,得到第三用户识别模型;将所述初始无标签样本输入所述第三用户识别模型,得到所述初始无标签样本的伪标签;通过随机森林对所述初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
- 根据权利要求3所述的基于智能决策的异常用户识别方法,其中,所述通过随机森林对所述初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征的步骤包括:将所述初始有标签样本和带有伪标签的初始无标签样本作为待筛选样本进行若干次有放回随机采样,得到若干特征筛选训练集;基于所述若干特征筛选训练集,生成若干决策树以得到随机森林;根据袋外数据计算所述随机森林中各决策树的第一袋外数据误差,其中,所述袋外数据来自所述各决策树所对应的特征筛选训练集;随机改变所述袋外数据中的特征,并计算各决策树的第二袋外数据误差;根据计算得到的第二袋外数据误差和第一袋外数据误差计算各特征的特征贡献度;根据计算得到的特征贡献度对所述初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
- 根据权利要求1所述的基于智能决策的异常用户识别方法,其中,所述对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集的步骤包括:对于每个无标签样本,根据无标签样本间的欧氏距离确定无标签样本的临近样本集,其中,所述临近样本集包括预设数量的临近样本;对于每个临近样本,在临近样本与所述无标签样本的特征空间连线上,选取扩充样本点;根据选取的扩充样本点以及所述无标签样本,构建得到与所述无标签样本对应的增强无标签样本集。
- 根据权利要求1所述的基于智能决策的异常用户识别方法,其中,所述通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型的步骤包括:将所述有标签样本以及与所述无标签样本对应的增强无标签样本集输入所述第二用户识别模型,得到所述有标签样本的用户预测结果,以及所述增强无标签样本集中各增强无标签样本的用户预测结果;根据所述各增强无标签样本的用户预测结果,确定所述无标签样本的用户预测结果;将前轮第二训练中所述无标签样本的用户预测结果,作为当前第二训练中所述无标签样本的伪标签,以计算所述有标签样本和所述无标签样本的正则化交叉熵损失;根据所述正则化交叉熵损失对所述第二用户识别模型进行参数调整,直至模型收敛,得到异常用户识别模型。
- 根据权利要求3所述的基于智能决策的异常用户识别方法,其中,所述将待识别用户样本输入所述异常用户识别模型,得到用户识别结果的步骤包括:获取待识别用户样本;根据预设的目标特征对所述待识别用户样本进行特征筛选;将特征筛选后的待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 一种基于智能决策的异常用户识别装置,包括:数据集获取模块,用于获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;数据重组模块,用于对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;第一训练模块,用于将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;数据增强模块,用于对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;第二训练模块,用于通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;样本输入模块,用于将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 一种计算机设备,包括存储器和处理器,所述存储器中存储有计算机可读指令,所述处理器执行所述计算机可读指令时实现如下步骤:获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 根据权利要求9所述的计算机设备,其中,所述对所述原始数据集进行数据重组,得到有标签样本以及无标签样本的步骤包括:将所述黑名单数据和所述验真用户数据分别与所述原始用户数据进行数据比对,以确定有标签用户列表及初始无标签样本;根据所述原始数据集对所述有标签用户列表进行数据填充,得到初始有标签样本;对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无 标签样本。
- 根据权利要求9所述的计算机设备,其中,所述对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无标签样本的步骤具体包括:将所述初始有标签样本输入第一用户识别模型,以通过所述初始有标签样本对所述第一用户识别模型进行第三训练,得到第三用户识别模型;将所述初始无标签样本输入所述第三用户识别模型,得到所述初始无标签样本的伪标签;通过随机森林对所述初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
- 根据权利要求9所述的计算机设备,其中,所述对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集的步骤包括:对于每个无标签样本,根据无标签样本间的欧氏距离确定无标签样本的临近样本集,其中,所述临近样本集包括预设数量的临近样本;对于每个临近样本,在临近样本与所述无标签样本的特征空间连线上,选取扩充样本点;根据选取的扩充样本点以及所述无标签样本,构建得到与所述无标签样本对应的增强无标签样本集。
- 根据权利要求9所述的计算机设备,其中,所述通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型的步骤包括:将所述有标签样本以及与所述无标签样本对应的增强无标签样本集输入所述第二用户识别模型,得到所述有标签样本的用户预测结果,以及所述增强无标签样本集中各增强无标签样本的用户预测结果;根据所述各增强无标签样本的用户预测结果,确定所述无标签样本的用户预测结果;将前轮第二训练中所述无标签样本的用户预测结果,作为当前第二训练中所述无标签样本的伪标签,以计算所述有标签样本和所述无标签样本的正则化交叉熵损失;根据所述正则化交叉熵损失对所述第二用户识别模型进行参数调整,直至模型收敛,得到异常用户识别模型。
- 根据权利要求11所述的计算机设备,其中,所述将待识别用户样本输入所述异常用户识别模型,得到用户识别结果的步骤包括:获取待识别用户样本;根据预设的目标特征对所述待识别用户样本进行特征筛选;将特征筛选后的待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机可读指令;其中,所述计算机可读指令被处理器执行时实现如下步骤:获取原始数据集,其中,所述原始数据集包括黑名单数据、验真用户数据以及原始用户数据;对所述原始数据集进行数据重组,得到有标签样本以及无标签样本;将所述有标签样本输入第一用户识别模型,以通过所述有标签样本对所述第一用户识别模型进行第一训练,得到第二用户识别模型;对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集;通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型;将待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
- 根据权利要求15所述的计算机可读存储介质,其中,所述对所述原始数据集进行数据重组,得到有标签样本以及无标签样本的步骤包括:将所述黑名单数据和所述验真用户数据分别与所述原始用户数据进行数据比对,以确定有标签用户列表及初始无标签样本;根据所述原始数据集对所述有标签用户列表进行数据填充,得到初始有标签样本;对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无标签样本。
- 根据权利要求15所述的计算机可读存储介质,其中,所述对所述初始有标签样本和所述初始无标签样本进行特征筛选,得到有标签样本以及无标签样本的步骤具体包括:将所述初始有标签样本输入第一用户识别模型,以通过所述初始有标签样本对所述第一用户识别模型进行第三训练,得到第三用户识别模型;将所述初始无标签样本输入所述第三用户识别模型,得到所述初始无标签样本的伪标签;通过随机森林对所述初始有标签样本和带有伪标签的初始无标签样本进行特征筛选,得到有标签样本以及无标签样本,并将筛选到的特征确定为目标特征。
- 根据权利要求15所述的计算机可读存储介质,其中,所述对所述无标签样本进行数据增强,得到与所述无标签样本对应的增强无标签样本集的步骤包括:对于每个无标签样本,根据无标签样本间的欧氏距离确定无标签样本的临近样本集,其中,所述临近样本集包括预设数量的临近样本;对于每个临近样本,在临近样本与所述无标签样本的特征空间连线上,选取扩充样本点;根据选取的扩充样本点以及所述无标签样本,构建得到与所述无标签样本对应的增强无标签样本集。
- 根据权利要求15所述的计算机可读存储介质,其中,所述通过所述有标签样本以及与所述无标签样本对应的增强无标签样本集,对所述第二用户识别模型进行第二训练,得到异常用户识别模型的步骤包括:将所述有标签样本以及与所述无标签样本对应的增强无标签样本集输入所述第二用户识别模型,得到所述有标签样本的用户预测结果,以及所述增强无标签样本集中各增强无标签样本的用户预测结果;根据所述各增强无标签样本的用户预测结果,确定所述无标签样本的用户预测结果;将前轮第二训练中所述无标签样本的用户预测结果,作为当前第二训练中所述无标签样本的伪标签,以计算所述有标签样本和所述无标签样本的正则化交叉熵损失;根据所述正则化交叉熵损失对所述第二用户识别模型进行参数调整,直至模型收敛,得到异常用户识别模型。
- 根据权利要求17所述的计算机可读存储介质,其中,所述将待识别用户样本输入所述异常用户识别模型,得到用户识别结果的步骤包括:获取待识别用户样本;根据预设的目标特征对所述待识别用户样本进行特征筛选;将特征筛选后的待识别用户样本输入所述异常用户识别模型,得到用户识别结果。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202011211553.3 | 2020-11-03 | ||
| CN202011211553.3A CN112307472B (zh) | 2020-11-03 | 2020-11-03 | 基于智能决策的异常用户识别方法、装置及计算机设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022095352A1 true WO2022095352A1 (zh) | 2022-05-12 |
Family
ID=74332907
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/090422 Ceased WO2022095352A1 (zh) | 2020-11-03 | 2021-04-28 | 基于智能决策的异常用户识别方法、装置及计算机设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112307472B (zh) |
| WO (1) | WO2022095352A1 (zh) |
Cited By (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115146735A (zh) * | 2022-07-19 | 2022-10-04 | 广州伟宏智能科技有限公司 | 用户用电异常识别 |
| CN115714687A (zh) * | 2022-11-23 | 2023-02-24 | 武汉轻工大学 | 入侵流量检测方法、装置、设备及存储介质 |
| CN115795307A (zh) * | 2022-11-25 | 2023-03-14 | 天翼电子商务有限公司 | 基于共轭梯度的能源数据采集样本扩充方法及系统 |
| CN115905548A (zh) * | 2023-03-03 | 2023-04-04 | 美云智数科技有限公司 | 水军识别方法、装置、电子设备及存储介质 |
| CN116296333A (zh) * | 2023-03-15 | 2023-06-23 | 国网山东省电力公司淄博供电公司 | 有载调压分接开关运行状态检测方法及系统 |
| CN116776150A (zh) * | 2023-06-20 | 2023-09-19 | 平安科技(深圳)有限公司 | 接口异常访问识别方法、装置、计算机设备及存储介质 |
| CN116957082A (zh) * | 2022-11-04 | 2023-10-27 | 中国移动通信有限公司研究院 | 样本筛选方法、装置、电子设备及存储介质 |
| CN118940153A (zh) * | 2024-10-14 | 2024-11-12 | 浙江大华技术股份有限公司 | 一种异常账户检测方法以及电子设备、存储介质 |
Families Citing this family (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112307472B (zh) * | 2020-11-03 | 2024-06-18 | 平安科技(深圳)有限公司 | 基于智能决策的异常用户识别方法、装置及计算机设备 |
| CN113705072B (zh) * | 2021-04-13 | 2025-08-29 | 腾讯科技(深圳)有限公司 | 数据处理方法、装置、计算机设备和存储介质 |
| CN113344066B (zh) * | 2021-05-31 | 2024-08-23 | 中国工商银行股份有限公司 | 一种模型训练方法、业务分配方法、装置及设备 |
| CN113722197B (zh) * | 2021-08-31 | 2023-10-17 | 上海观安信息技术股份有限公司 | 移动终端异常识别方法、系统 |
| CN113658178B (zh) * | 2021-10-14 | 2022-01-25 | 北京字节跳动网络技术有限公司 | 组织图像的识别方法、装置、可读介质和电子设备 |
| CN113919357A (zh) * | 2021-10-29 | 2022-01-11 | 平安普惠企业管理有限公司 | 地址实体识别模型的训练方法、装置、设备及存储介质 |
| CN114417968B (zh) * | 2021-12-16 | 2025-10-03 | 深圳供电局有限公司 | 异常监测分类模型构建方法、异常监测方法以及装置 |
| CN114186646A (zh) * | 2022-02-15 | 2022-03-15 | 国网区块链科技(北京)有限公司 | 区块链异常交易识别方法及装置、存储介质及电子设备 |
| CN114897099A (zh) * | 2022-06-06 | 2022-08-12 | 上海淇玥信息技术有限公司 | 基于客群偏差平滑优化的用户分类方法、装置及电子设备 |
| CN115272983B (zh) * | 2022-09-29 | 2023-01-03 | 成都中轨轨道设备有限公司 | 基于图像识别的接触网悬挂状态监测方法及系统 |
| CN115600127A (zh) * | 2022-10-26 | 2023-01-13 | 中国农业银行股份有限公司(Cn) | 一种用户类型识别方法、装置、电子设备及介质 |
| CN115859099B (zh) * | 2022-11-22 | 2026-01-06 | 上海交通大学 | 样本生成方法、装置、电子设备和存储介质 |
| CN116129440B (zh) * | 2023-04-13 | 2023-07-04 | 新兴际华集团财务有限公司 | 异常用户端告警方法、装置、电子设备和介质 |
| CN116304932B (zh) * | 2023-05-19 | 2023-09-05 | 湖南工商大学 | 一种样本生成方法、装置、终端设备及介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103530373A (zh) * | 2013-10-15 | 2014-01-22 | 无锡清华信息科学与技术国家实验室物联网技术中心 | 不均衡感知数据下的移动应用分类方法 |
| CN108040073A (zh) * | 2018-01-23 | 2018-05-15 | 杭州电子科技大学 | 信息物理交通系统中基于深度学习的恶意攻击检测方法 |
| US20190243735A1 (en) * | 2018-02-05 | 2019-08-08 | Wuhan University | Deep belief network feature extraction-based analogue circuit fault diagnosis method |
| CN110732139A (zh) * | 2019-10-25 | 2020-01-31 | 腾讯科技(深圳)有限公司 | 检测模型的训练方法和用户数据的检测方法、装置 |
| CN111639540A (zh) * | 2020-04-30 | 2020-09-08 | 中国海洋大学 | 基于相机风格和人体姿态适应的半监督人物重识别方法 |
| CN112307472A (zh) * | 2020-11-03 | 2021-02-02 | 平安科技(深圳)有限公司 | 基于智能决策的异常用户识别方法、装置及计算机设备 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107133265B (zh) * | 2017-03-31 | 2021-07-09 | 咪咕动漫有限公司 | 一种识别行为异常用户的方法及装置 |
| CN108764281A (zh) * | 2018-04-18 | 2018-11-06 | 华南理工大学 | 一种基于半监督自步学习跨任务深度网络的图像分类方法 |
| CN109325844A (zh) * | 2018-06-25 | 2019-02-12 | 南京工业大学 | 多维数据下的网贷借款人信用评价方法 |
| CN111222648B (zh) * | 2020-01-15 | 2023-09-26 | 深圳前海微众银行股份有限公司 | 半监督机器学习优化方法、装置、设备及存储介质 |
| CN111783981A (zh) * | 2020-06-29 | 2020-10-16 | 百度在线网络技术(北京)有限公司 | 模型训练方法、装置、电子设备及可读存储介质 |
-
2020
- 2020-11-03 CN CN202011211553.3A patent/CN112307472B/zh active Active
-
2021
- 2021-04-28 WO PCT/CN2021/090422 patent/WO2022095352A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103530373A (zh) * | 2013-10-15 | 2014-01-22 | 无锡清华信息科学与技术国家实验室物联网技术中心 | 不均衡感知数据下的移动应用分类方法 |
| CN108040073A (zh) * | 2018-01-23 | 2018-05-15 | 杭州电子科技大学 | 信息物理交通系统中基于深度学习的恶意攻击检测方法 |
| US20190243735A1 (en) * | 2018-02-05 | 2019-08-08 | Wuhan University | Deep belief network feature extraction-based analogue circuit fault diagnosis method |
| CN110732139A (zh) * | 2019-10-25 | 2020-01-31 | 腾讯科技(深圳)有限公司 | 检测模型的训练方法和用户数据的检测方法、装置 |
| CN111639540A (zh) * | 2020-04-30 | 2020-09-08 | 中国海洋大学 | 基于相机风格和人体姿态适应的半监督人物重识别方法 |
| CN112307472A (zh) * | 2020-11-03 | 2021-02-02 | 平安科技(深圳)有限公司 | 基于智能决策的异常用户识别方法、装置及计算机设备 |
Cited By (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115146735A (zh) * | 2022-07-19 | 2022-10-04 | 广州伟宏智能科技有限公司 | 用户用电异常识别 |
| CN116957082A (zh) * | 2022-11-04 | 2023-10-27 | 中国移动通信有限公司研究院 | 样本筛选方法、装置、电子设备及存储介质 |
| CN115714687A (zh) * | 2022-11-23 | 2023-02-24 | 武汉轻工大学 | 入侵流量检测方法、装置、设备及存储介质 |
| CN115714687B (zh) * | 2022-11-23 | 2024-06-04 | 武汉轻工大学 | 入侵流量检测方法、装置、设备及存储介质 |
| CN115795307A (zh) * | 2022-11-25 | 2023-03-14 | 天翼电子商务有限公司 | 基于共轭梯度的能源数据采集样本扩充方法及系统 |
| CN115905548A (zh) * | 2023-03-03 | 2023-04-04 | 美云智数科技有限公司 | 水军识别方法、装置、电子设备及存储介质 |
| CN115905548B (zh) * | 2023-03-03 | 2024-05-10 | 美云智数科技有限公司 | 水军识别方法、装置、电子设备及存储介质 |
| CN116296333A (zh) * | 2023-03-15 | 2023-06-23 | 国网山东省电力公司淄博供电公司 | 有载调压分接开关运行状态检测方法及系统 |
| CN116776150A (zh) * | 2023-06-20 | 2023-09-19 | 平安科技(深圳)有限公司 | 接口异常访问识别方法、装置、计算机设备及存储介质 |
| CN118940153A (zh) * | 2024-10-14 | 2024-11-12 | 浙江大华技术股份有限公司 | 一种异常账户检测方法以及电子设备、存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112307472A (zh) | 2021-02-02 |
| CN112307472B (zh) | 2024-06-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022095352A1 (zh) | 基于智能决策的异常用户识别方法、装置及计算机设备 | |
| CN111784528B (zh) | 异常社群检测方法、装置、计算机设备及存储介质 | |
| CN112148987B (zh) | 基于目标对象活跃度的消息推送方法及相关设备 | |
| WO2022126971A1 (zh) | 基于密度的文本聚类方法、装置、设备及存储介质 | |
| WO2022142032A1 (zh) | 手写签名校验方法、装置、计算机设备及存储介质 | |
| WO2021155713A1 (zh) | 基于权重嫁接的模型融合的人脸识别方法及相关设备 | |
| WO2022007438A1 (zh) | 情感语音数据转换方法、装置、计算机设备及存储介质 | |
| WO2022174491A1 (zh) | 基于人工智能的病历质控方法、装置、计算机设备及存储介质 | |
| CN111222976B (zh) | 一种基于双方网络图数据的风险预测方法、装置和电子设备 | |
| CN111199474B (zh) | 一种基于双方网络图数据的风险预测方法、装置和电子设备 | |
| CN109325118B (zh) | 不平衡样本数据预处理方法、装置和计算机设备 | |
| CN112035549B (zh) | 数据挖掘方法、装置、计算机设备及存储介质 | |
| CN112288025B (zh) | 基于树结构的异常案件识别方法、装置、设备及存储介质 | |
| CN110855648A (zh) | 一种网络攻击的预警控制方法及装置 | |
| CN112995414B (zh) | 基于语音通话的行为质检方法、装置、设备及存储介质 | |
| CN114219664B (zh) | 产品推荐方法、装置、计算机设备及存储介质 | |
| CN111639360A (zh) | 智能数据脱敏方法、装置、计算机设备及存储介质 | |
| CN113887214B (zh) | 基于人工智能的意愿推测方法、及其相关设备 | |
| CN112417886A (zh) | 意图实体信息抽取方法、装置、计算机设备及存储介质 | |
| CN119603690A (zh) | 诈骗电话的识别方法、装置、计算机设备及存储介质 | |
| CN114186597A (zh) | 集群识别方法、装置、设备及存储介质 | |
| CN114238574B (zh) | 基于人工智能的意图识别方法及其相关设备 | |
| TW202018627A (zh) | 核身方法及裝置 | |
| CN117078332A (zh) | 异常行为检测方法、装置、计算机设备及存储介质 | |
| CN111353871B (zh) | 一种基于双方网络图数据的风险预测方法、装置和电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21888056 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21888056 Country of ref document: EP Kind code of ref document: A1 |
