WO2021135540A1 - 基于Neo4j的异常用户处理方法、装置、计算机设备和介质 - Google Patents
基于Neo4j的异常用户处理方法、装置、计算机设备和介质 Download PDFInfo
- Publication number
- WO2021135540A1 WO2021135540A1 PCT/CN2020/122828 CN2020122828W WO2021135540A1 WO 2021135540 A1 WO2021135540 A1 WO 2021135540A1 CN 2020122828 W CN2020122828 W CN 2020122828W WO 2021135540 A1 WO2021135540 A1 WO 2021135540A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- abnormal
- relationship
- user
- user group
- node
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/953—Querying, e.g. by the use of web search engines
- G06F16/9536—Search customisation based on social or collaborative filtering
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/36—Creation of semantic tools, e.g. ontology or thesauri
- G06F16/367—Ontology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/40—Business processes related to social networking or social networking services
Definitions
- This application relates to the field of big data, and in particular to a Neo4j-based abnormal user processing method, device, computer equipment and storage medium.
- this application provides a Neo4j-based processing method, device, computer equipment and storage medium for abnormal users to solve the technical problem of inaccurate processing of abnormal users in the prior art.
- a Neo4j-based abnormal user processing method includes:
- a corresponding target label is generated for the abnormal user group, and the detected abnormal user group with a specific operation is verified according to the target label.
- a Neo4j-based abnormal user processing device includes:
- the extraction module is used to extract user features from the acquired user data
- the detection module is used to input the extracted user characteristics into the Neo4j algorithm for prediction, and obtain the user group with the same attribute tag in the user data as an abnormal user group;
- the scoring module is configured to score the abnormal user group based on the preset weight ratio, according to the number of users in the abnormal user group, the calibration characteristics, and the common characteristics;
- the processing module is used to generate a corresponding target label for the abnormal user group according to the score, and perform verification processing on the detected abnormal user group with a specific operation according to the target label.
- a computer device includes a memory and a processor, and computer-readable instructions stored in the memory and capable of running on the processor.
- the processor executes the computer-readable instructions, the following is implemented based on Neo4j Steps of the abnormal user handling method:
- a corresponding target label is generated for the abnormal user group, and the detected abnormal user group with a specific operation is verified according to the target label.
- a computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the following steps of the Neo4j-based abnormal user processing method are implemented: User data for user feature extraction;
- a corresponding target label is generated for the abnormal user group, and the detected abnormal user group with a specific operation is verified according to the target label.
- Neo4j-based abnormal user processing method by rating abnormal users according to the number and characteristics of abnormal users in the abnormal user group, and then generating targeted tags based on the ratings, when abnormal users trigger specific operations Then, different processing mechanisms are triggered according to the targeted tags, and the abnormal users are subjected to targeted verification processing, which solves the technical problem of inaccurate verification processing for abnormal users in the prior art.
- Figure 1 is a schematic diagram of the application environment of the abnormal user handling method based on Neo4j;
- Figure 2 is a schematic flow diagram of a Neo4j-based abnormal user handling method
- FIG. 3 is a schematic flowchart of step 204 in FIG. 2;
- FIG. 4a is a schematic diagram of a relationship map in step 302;
- Figure 4b is another schematic diagram of the relationship map in step 302;
- FIG. 4c is a schematic diagram of another relationship map in step 302;
- FIG. 5 is a schematic flowchart of step 206 in FIG. 2;
- Figure 6 is a schematic diagram of an abnormal user processing device based on Neo4j
- Figure 7 is a schematic diagram of a computer device in an embodiment.
- the Neo4j-based abnormal user processing method provided in the embodiment of the application can be applied to the application environment as shown in FIG. 1.
- the application environment may include the terminal 102, the network, and the server 104.
- the network is used to provide a communication link medium between the terminal 102 and the server 104.
- the network may include various connection types, such as wired, wireless communication links or Fiber optic cable and so on.
- the user can use the terminal 102 to interact with the server 104 through the network to receive or send messages and so on.
- Various communication client applications such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc., may be installed on the terminal 102.
- the terminal 102 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, moving picture experts compress standard audio Level 3), MP4 (Moving Picture Experts Group Audio Layer IV, Motion Picture Experts compress standard audio level 4) Players, laptop portable computers and desktop computers, etc.
- MP3 players Moving Picture Experts Group Audio Layer III, moving picture experts compress standard audio Level 3
- MP4 Motion Picture Experts compress standard audio level 4
- laptop portable computers and desktop computers etc.
- the server 104 may be a server that provides various services, for example, a background server that provides support for pages displayed on the terminal 102.
- Neo4j-based abnormal user processing method provided in the embodiments of the present application is generally executed by the server/terminal. Accordingly, the Neo4j-based abnormal user processing apparatus is generally set in the server/terminal device.
- terminals, networks, and servers in FIG. 1 are merely illustrative. There can be any number of terminal devices, networks, and servers according to implementation needs.
- the terminal 102 communicates with the server 104 through the network.
- the server 104 obtains user data from the terminal 102 through the network, extracts the characteristics of the user data, inputs the user characteristics into the Neo4j algorithm to predict the abnormal user group, and then analyzes the abnormal user based on the number of users in the abnormal user group, the calibration characteristics and the common characteristics The group is scored, and the corresponding target label is generated according to the score, and finally the abnormal user in the abnormal user group is subjected to corresponding verification processing according to the target label.
- the terminal 102 and the server 104 are connected through a network.
- the network can be a wired network or a wireless network.
- the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, and portable wearable devices.
- the server 104 can be implemented by an independent server or a cluster of multiple servers.
- a Neo4j-based abnormal user processing method is provided. Taking the method applied to the server in FIG. 1 as an example, the method includes the following steps:
- Step 202 Perform user feature extraction on the acquired user data.
- User data includes the user's mobile phone number, IP address, device number, and the relationship between the mobile phone number, IP address, device number, and the relationship between the user characteristics of different users, and so on. In addition, it also includes the activities of different users. Types, such as daily check-in, newcomer registration, new registration invitation, etc. Among them, a user can use a mobile phone number, multiple IP addresses and device numbers.
- User feature extraction is the process of extracting feature values with the user's mobile phone number, IP address, and device number as dimensions, and the correlation between the three features of mobile phone number, IP address, and device number.
- the above user data can also be stored in the nodes of the blockchain in a distributed manner.
- Step 204 Input the extracted user characteristics into the Neo4j algorithm for prediction, and obtain user groups with the same attribute tags in the user data as abnormal user groups.
- Neo4j is a high-performance, NOSQL graph database that stores structured data on the network instead of in tables.
- Label Propagation is an algorithm in Neo4j that propagates labels through edges between nodes.
- the user’s mobile phone number, IP address, and device number all have a labeling feature to indicate whether the mobile phone number, IP address, or device number is in the blacklist, whitelist, or unconfirmed list of the server.
- a user’s mobile phone number, The device number corresponding to the mobile phone number can be in the whitelist, but the corresponding IP address can be in the blacklist; it is also possible that the user's mobile phone number, IP address, and device number are all in the blacklist.
- the data in the blacklist and whitelist is obtained based on historical experience. Some of the blacklist data is obtained from user reporting and blocking, and the whitelist is obtained based on the characteristics of the user's activity participation in a certain APP.
- the extracted user characteristics are input into the label propagation algorithm.
- the label propagation algorithm uses the three types of mobile phone number, IP and device number as relationship nodes, and builds a relationship network based on the connections between the three types of relationship nodes.
- the relationship between two relationship nodes is represented in the relationship network as an edge connecting the two relationship nodes.
- User data includes real-time user data. Real-time user data can be directly applied to user data collected in real time to ensure the timeliness of abnormal user prediction. Then predict the abnormal user group based on the constructed relationship network.
- the Neo4j algorithm is used to predict abnormal user groups, and there is no need to obtain a risk map according to a predetermined node template, which makes the prediction process more flexible and universal.
- Step 206 based on the preset weight ratio, score the abnormal user group according to the number of users in the abnormal user group, the calibration characteristics, and the common characteristics.
- the preset weight ratio is the ratio of the characteristics of the abnormal user group obtained based on historical experience and can be set according to specific application scenarios;
- the calibration features include but are not limited to users’ daily check-in, newcomer registration, new registration invitation, specific Activities and other dimensions;
- specific activities include but are not limited to double eleven activities, 618 activities and other activities for users that have virtual objects such as high-return cards, coupons, coupons, etc.;
- common features refer to abnormalities in abnormal user groups Users may have the same characteristic dimension characteristics, for example, more than 50% of abnormal users in an abnormal user group have the same device number, mobile phone attribution, registration time, or the same IP.
- Step 208 Generate a corresponding target label for the abnormal user group according to the score, and perform verification processing on the detected abnormal user group with a specific operation according to the target label.
- Specific operations generally include: an abnormal user's request to log in to an app account in an abnormal user group, an operation requesting to obtain a certain functional module in an app, and so on.
- Different range of scores correspond to different target tags, and different target tags will trigger different verification processing mechanisms. For example, when a score of 91 for a certain abnormal user group is obtained, a target label for the abnormal user group can be generated to refuse to log in and/or to refuse to participate in activities, to restrict the activities of the abnormal user group, such as those who refuse to log in.
- the verification processing can achieve the purpose of different verification processing for different abnormal user groups, and can also achieve the purpose of reducing the redundant data generated by the abnormal user groups, thereby improving the efficiency of server-side data processing.
- the abnormal user group is predicted by the Neo4j algorithm on the server, and then the abnormal user group is scored according to the three specified characteristics of the abnormal user group, and the corresponding target label is generated according to the scoring result.
- Different targeting tags deal with the login and activities of abnormal users in the abnormal user groups to different degrees, so as to achieve different verification processing purposes for different abnormal user groups, and also reduce the redundant data generated by the abnormal user groups, thereby reducing the redundant data generated by the abnormal user groups. The purpose of improving the efficiency of server data processing.
- step 204 includes:
- Step 302 Construct a relationship network based on the association relationship between the user characteristics through the Neo4j algorithm, where the relationship network is mainly composed of the relationship node, which is a mobile phone number, an IP address, and a device number.
- the triangle represents the IP address
- the circle represents the mobile phone number
- the pentagon represents the device number.
- An IP address can correspond to multiple mobile phone numbers
- a mobile phone number can also correspond to multiple device numbers.
- a device number can also correspond to multiple IP addresses.
- the device number and the IP address are associated with a mobile phone number to form a second-degree relationship
- the mobile phone number is directly associated with the device number and IP address to form a first-degree relationship.
- Figure 4a is a relationship network constructed based on normal user data (whitelisted users, user data of non-white and non-black unconfirmed nodes). It can be seen that the relationship network of normal users is relatively scattered in terms of single point characteristics; Figure 4b is The relationship network constructed based on the blacklisted user data, in some single-point features, presents the same scattered characteristics as normal users, making it difficult for single-point feature defense to be effective, but if the user characteristics are modeled and displayed in the form of a network, you will find In some special graphic features, abnormal behavior is obviously different from normal behavior.
- the single point feature refers to a single feature. For example, the number of device numbers used by the user and the number of IPs are two characteristics. From a single point of view, such as the number of device numbers, it is impossible to distinguish abnormal users from normal users.
- User characteristics refer to the user doing a certain thing at a certain time (period), for example, the user logs in to an APP 10 times from 1 am to 3 am; the user signs in to an APP 20 times within 30 days, these are all user characteristics .
- Neo4j building a relationship network based on user data through Neo4j is shown in the schematic diagram of the relationship map in FIG. 4c, where FIG. 4c is only exemplary.
- Step 304 Obtain the annotation features of the relationship node, and classify the relationship nodes according to the annotation feature to obtain blacklist nodes, whitelist nodes, and unconfirmed nodes, where the annotation feature is the initial category of the relationship node.
- the annotation feature indicates which of the blacklist nodes in the blacklist, the whitelist nodes in the whitelist, or the unconfirmed nodes in the unconfirmed list the user's mobile phone number, IP address, or device number belongs to.
- the relationship network just obtained will have an initial category.
- the initial category includes whether the IP address, mobile phone number, or device number belongs to the blacklist, whitelist, or unconfirmed list. Among them, the label that can be generated for the relationship node in the blacklist is 1. The label generated for the relationship node of the whitelist and the unconfirmed list is 0.
- Step 306 Generate an attribute label for the relationship node, where the attribute label includes abnormal and normal.
- the labeling characteristics of the mobile phone number, device number, and IP address in the blacklist are used to mark the initial attribute label of each relationship node as an exception, and there is also an attribute label: the black label; and the relationship node in the whitelist Annotation features.
- the initial attribute label used to label each relationship node is normal, and it also has an attribute label: white label.
- the black label and white label are to prevent the attribute labels of the relationship nodes in the black and white lists from being updated in subsequent node labels.
- the attribute label of the relationship node (mobile phone number, IP address, and device number) in the unconfirmed list is: unidentified (meaning unidentified label), which can be refreshed in subsequent label predictions, and its initial label The characteristics are normal.
- the attribute tags of the IP address, device number or mobile phone number in the blacklist or whitelist will not be updated or changed in the end.
- Step 308 Use the relationship node whose attribute tag is the abnormal mobile phone number and IP address as the seed node, and obtain the node path whose path length is not greater than the first preset value from the relationship network to obtain the node relationship map, where the relationship map includes An abnormal number relationship graph based on abnormal mobile phone number nodes and an abnormal IP relationship graph based on abnormal IP nodes.
- the abnormal IP address with the attribute tag as the seed node Take the abnormal IP address with the attribute tag as the seed node, and obtain the abnormal IP relationship graph based on the abnormal IP node composed of the node path whose path length is not greater than the first preset value, where the first preset value can be set to 10 based on experience . It is also necessary to use the cell phone number with the attribute label as the abnormal cell phone number as the seed node, and obtain the abnormal number relationship graph based on the abnormal cell phone number node composed of the node path whose path length is not greater than 10.
- Neo4j can be used to search for a node path composed of relationship nodes with a path length between 1 and 10.
- the expression can be (a)-[*1..10]->(b), cypher language is similar to SQL query statement. Among them, a and b represent relationship nodes, and [*1..10] represents a node path from 1 to 10 degrees.
- relationship nodes are used as seed nodes to construct the relationship graph in order to make the predicted graph more comprehensive and the prediction process more efficient; if a relationship node is used as the seed node to construct a large relationship graph, then It is necessary to obtain node paths not greater than N, where N is an integer; and N needs to be much greater than 10 to ensure that most of the node connections can be included in the relationship map, and this will lead to path cuts during the construction of the relationship map It takes too long to divide, which affects the efficiency of forecasting. Therefore, using two nodes to construct two relationship maps can ensure the integrity of the relationship maps and the efficiency of segmentation.
- Step 310 Perform label prediction on the unconfirmed nodes in the abnormal number relationship graph and the abnormal IP relationship graph respectively to obtain an abnormal user group.
- This step is also implemented in the label propagation algorithm.
- the labels of the relationship nodes in the unconfirmed list are normal by default, and the attribute labels of other relationship nodes are either determined to be abnormal or normal.
- the label update result of the relationship node B will be affected by the label of the relationship node A.
- the number of neighbor nodes whose attribute labels of unconfirmed nodes are abnormal and normal is compared to obtain the label comparison result.
- the label is updated according to the label of its neighbor node.
- the relationship map after the attribute label is updated is used as the map to be predicted.
- the preset correction condition is a condition for adjusting the label of the relationship node during the label update process. Specifically, after each round of label changes, it is necessary to modify the label results of a part of the relationship nodes: that is, the attribute labels of unconfirmed nodes that have a one-time relationship with the IP address in the whitelist are set to normal. For each unconfirmed node in the two relationship graphs, label refresh is performed in this manner in turn, until the label result of each relationship node remains unchanged or the number of refresh rounds reaches the threshold value of 10,000.
- the attribute tag of the mobile phone number that has a one-degree relationship with the IP address in the whitelist node is set to normal, where the one-degree relationship refers to an association relationship between two directly connected relationship nodes.
- the labels of all the relationship nodes are normal, so as to ensure that the attribute labels of the neighbor nodes of each mobile phone number node are also normal.
- the user corresponding to the mobile phone number node is marked as an abnormal user only when the tags of the mobile phone number node, IP node, and device number node of the same user are all abnormal.
- the user group corresponding to the relationship node in the relationship subgraph is regarded as the abnormal user group.
- the user is confirmed based on the mobile phone number/IP address.
- the number of fleece parties is gathered in crowds, using one mobile phone number to scan data under different IP addresses and different device numbers, and also using different mobile phone numbers and different device numbers under the same IP address; the types of user fraud can be divided.
- the obtained abnormal user groups can also be compared to perform deduplication processing on the abnormal user groups. Specifically, if the abnormal user groups based on the abnormal mobile phone number nodes are the same as the abnormal user groups based on the abnormal IP nodes, then the abnormal user groups are merged, Get the combined abnormal user group.
- the same user may be obtained, that is, the group A obtained by predicting the abnormal mobile phone number as the central node, and the group A obtained by the abnormal IP node as the center, so you need to find Whether the abnormal user group has the same label, merge the same users of the abnormal user group with the same label to save storage space, and generate the same targeted label for the merged abnormal user group, and subsequently when the abnormal user group is detected
- the specific operation of the abnormal user in, such as a login operation will be verified according to the target tag of the abnormal user.
- the relationship node directly uses the neighbor labels between the relationship nodes to update the label of the relationship node that needs to determine the label. It is not limited to updating the label for a specified relationship node, except for special labels (blacklist, whitelist). , The labels of all nodes are constantly being refreshed in the prediction, to ensure that the influence of neighbor labels can be fully utilized in the label prediction, and finally the relationship edge where the node with the label is normal is disconnected to generate a subgraph, making the prediction process flexible It is more flexible and universal, and in the process of tag iteration, iterates only for unconfirmed nodes, and the traversal time is short. And after iterating once, reset the label of the once node (ie mobile phone number) associated with the whitelisted IP address to normal to prevent mispredictions such as users who use corporate wifi from being blacklisted.
- blacklist whitelist
- step 206 includes:
- Step 502 Obtain the number of users of the abnormal user group, and set the first rating score of the abnormal user group according to the first preset rating score and the number of users.
- the number of users refers to the number of abnormal users in an abnormal user group, and the first rating score of the abnormal user group can be determined according to the number of users; specifically:
- the number of users of the abnormal user group and the first preset rating score are shown in Table 1:
- the corresponding first preset rating score is 10
- the first bottle machine score of the abnormal user group is 10.
- Step 504 Obtain the calibrated number of abnormal users with calibrated features in the abnormal user group, and determine the second rating score of the abnormal user group according to the operational priority of the calibrated features.
- the operational priority of the calibration feature indicates the priority of the influence of different calibration features on the second rating score. For example, 40 points for daily check-in, 60 points for newcomer registration, 80 points for new registration invitation, and 100 points for specific activities. Among them, the computing priority of specific activities is higher than daily check-in, newcomer registration and new registration invitation, and the computing priority of new registration invitation is also Greater than daily check-in and newcomer registration, and newcomer registration is greater than daily check-in.
- the score corresponding to the specific activity is used as the second rating score of the abnormal user group. More than 50% is the calibration quantity. For example, if the number of calibrations is less than 50%, the abnormal user group may not be scored, that is, the second rating score is 0.
- Step 506 Count the total number of abnormal users with common characteristics in the abnormal user group, and determine the third rating score of the abnormal user group according to the total number and the second preset rating score.
- the total number is the number of abnormal users with a certain common feature in the abnormal user group, which can be used to judge whether to determine the third rating score of the abnormal user group according to the score corresponding to the common feature.
- Common features are that abnormal users have the same IP, mobile phone attribution, the same registration time, and the same device number.
- Same IP-40 same mobile phone attribution -60, same registration time -80, same device number -100;
- the score of the abnormal user group is based on the score corresponding to the common feature with the highest score, that is, 100 points for the same device number.
- Step 508 Combine the first rating score, the second rating score, and the third rating score according to the preset weight ratio to obtain the score of the abnormal user group.
- the abnormal user group is scored.
- the preset weight ratio is obtained based on experience, and the accuracy of calculating the score based on this ratio is relatively high:
- abnormal user group A For example, abnormal user group A:
- the first grading score 20 points * 0.2
- Second grading score 80 points * 0.4
- the third grade score 100 points * 0.4
- Abnormal users with different targeted tags trigger different verification mechanisms on the server.
- the targeted tags for SMS verification at login will trigger the SMS verification mechanism when the user logs in, verify the user, etc., so that when the abnormal user logs in to the APP/webpage ,
- Adopt different verification methods increase the verification threshold for abnormal users, prevent machine batch operation, reduce the amount of server-side garbage data processing, and improve data processing efficiency.
- the target tag of the abnormal user group where the abnormal user is located is obtained; the target tag of the abnormal user is generated in advance before the specific operation of the user is detected.
- a specific operation of an abnormal user or some abnormal users in the abnormal user group, such as a login operation will obtain the target tag of the abnormal user, and call the corresponding verification processing interface according to the target tag, and the activities of the abnormal user Perform authorization verification processing; determine whether the number of verification processing times for the activity authorization of the abnormal user exceeds the authorization value, and set the abnormal user that exceeds the authorization value as a locked user.
- the SMS interface can be called according to their scores to send SMS verification codes. If the user logs in with the correct SMS within the valid time, the login is successful, then the next time the user logs in within the valid time, the SMS interface will not be called again. If the login fails, the SMS interface will be called again until the number of times the SMS interface is called, that is, the permission value is 3 consecutive times, which means that the user does not enter the correct verification code within the valid time, and the login fails, the account is locked, and Set the abnormal user as a locked user and restrict other operations.
- the verification threshold for abnormal users is increased, the batch operation of machines is prevented, the processing volume of server-side junk data is reduced, and the data processing efficiency is improved.
- a Neo4j-based abnormal user processing device is provided, and the Neo4j-based abnormal user processing device corresponds to the Neo4j-based abnormal user processing method in the foregoing embodiment in a one-to-one correspondence.
- the Neo4j-based abnormal user processing device includes:
- the extraction module 602 is configured to perform user feature extraction on the acquired user data.
- the detection module 604 is configured to input the extracted user characteristics into the Neo4j algorithm for prediction, and obtain the user group with the same attribute tag in the user data as an abnormal user group.
- the scoring module 606 is configured to score the abnormal user group according to the number of users in the abnormal user group, the calibration characteristics, and the common characteristics based on the preset weight ratio. as well as
- the processing module 608 is configured to generate a corresponding target label for the abnormal user group according to the score, and perform verification processing on the detected user group with a specific abnormal operation according to the target label.
- the above user data can also be stored in the nodes of the blockchain in a distributed manner.
- the detection module 604 includes:
- the construction sub-module is used to construct a relationship network based on the relationship between the characteristics of each user through the Neo4j algorithm.
- the relationship network is mainly composed of the relationship node, which is a mobile phone number, an IP address, and a device number.
- the classification sub-module is used to obtain the label features of the relationship nodes, and classify the relationship nodes according to the label features to obtain blacklist nodes, whitelist nodes, and unconfirmed nodes, where the label features are the initial categories of the relationship nodes.
- the label sub-module is used to generate attribute labels for the relationship nodes, where the attribute labels include abnormal and normal.
- the screening sub-module is used to respectively use the relationship node with the abnormal mobile phone number and IP address as the seed node to obtain the node path whose path length is not greater than the first preset value from the relationship network to obtain the node relationship graph, where:
- the relationship graph includes an abnormal number relationship graph based on abnormal mobile phone number nodes and an abnormal IP relationship graph based on abnormal IP nodes.
- the prediction sub-module is used to respectively perform label prediction on the unconfirmed nodes in the abnormal number relationship graph and the abnormal IP relationship graph to obtain an abnormal user group.
- the prediction sub-module includes:
- the comparison unit is used to compare the number of neighbor nodes whose attribute labels of unconfirmed nodes are abnormal and normal, and obtain a label comparison result.
- the update unit is used to update the attribute label of the unconfirmed node according to the label comparison result.
- the correction unit is used to correct the attribute label of the relationship node after the attribute label is updated according to the preset correction condition, and repeat the operation of comparing and updating the attribute label, until the attribute label of each unconfirmed node is no longer changed or the number of updates reaches Threshold, the relationship map after the last attribute tag update is used as the map to be predicted.
- the broken edge unit is used to disconnect the abnormal edges in the graph to be predicted to obtain multiple relationship subgraphs.
- the abnormal edge is an edge composed of any two relationship nodes directly connected, and at least one of the relationship nodes has an attribute label of a normal edge. .
- the statistical unit is used to count the attribute labels of the relationship nodes in the relationship sub-graph. If the number of mobile phone numbers or IP addresses with abnormal attribute labels is greater than the second preset value, the user group corresponding to the relationship node in the relationship sub-graph is taken as Abnormal user groups.
- prediction sub-module also includes:
- the comparison unit is used to compare the obtained abnormal user groups.
- the merging unit is configured to merge the abnormal user groups based on the abnormal mobile phone number node and the abnormal user group based on the abnormal IP node to obtain the combined abnormal user group.
- scoring module 606 includes:
- the first score submodule is used to obtain the number of users of the abnormal user group, and set the first rating score of the abnormal user group according to the first preset rating score and the number of users;
- the second score sub-module is used to obtain the calibrated number of abnormal users with calibrated features in the abnormal user group, and determine the second rating score of the abnormal user group according to the operational priority of the calibrated features;
- the third score sub-module is used to count the total number of abnormal users with common characteristics in the abnormal user group, and determine the third rating score of the abnormal user group according to the total number and the second preset rating score;
- the comprehensive score sub-module is used to combine the first rating score, the second rating score, and the third rating score according to the preset weight ratio to obtain the score of the abnormal user group.
- processing module 608 includes;
- the obtaining sub-module is used to obtain the target tag of the abnormal user group where the abnormal user is located if a specific operation of the abnormal user is detected;
- the calling sub-module is used to call the corresponding interface according to the interface calling type to process the activity permissions of the abnormal user;
- the locking sub-module is used to determine whether the processing times of the activity authority of the abnormal user exceeds the authority value, and set the abnormal user who exceeds the authority value as the locked user.
- Neo4j-based abnormal user processing device ranks abnormal users according to the number and characteristics of users in the abnormal user group, and then generates targeted tags based on the rating levels.
- targeted tags are triggered according to the targeted tags.
- the processing mechanism performs targeted verification processing on abnormal users, which solves the technical problem of inaccurate handling of abnormal users in the prior art.
- a computer device is provided.
- the computer device may be a server, and its internal structure diagram may be as shown in FIG. 7.
- the computer equipment includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide calculation and control capabilities.
- the memory of the computer device includes a non-volatile storage medium and an internal memory.
- the non-volatile storage medium stores an operating system, computer readable instructions, and a database.
- the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium.
- the database of the computer equipment is used to store user data.
- the network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instruction is executed by the processor, a Neo4j-based abnormal user processing method is realized.
- the computer device here is a device that can automatically perform numerical calculation and/or information processing in accordance with pre-set or stored instructions.
- Its hardware includes, but is not limited to, a microprocessor, a dedicated Integrated Circuit (Application Specific Integrated Circuit, ASIC), Programmable Gate Array (Field-Programmable Gate Array, FPGA), Digital Processor (Digital Signal Processor, DSP), embedded equipment, etc.
- a computer-readable storage medium is provided with computer-readable instructions stored thereon.
- the steps of the Neo4j-based abnormal user processing method in the above-mentioned embodiment are implemented, for example, Step 202 to step 208 shown in FIG. 2, or, when the processor executes computer-readable instructions, the function of each module/unit of the Neo4j-based abnormal user processing device in the above embodiment is realized, for example, modules 602 to modules shown in FIG. 6 608 function. To avoid repetition, I won’t repeat them here.
- Non-volatile computer readable storage medium when the computer readable instructions are executed, they may include the processes of the above-mentioned method embodiments.
- any reference to memory, storage, database, or other media used in the embodiments provided in this application may include non-volatile and/or volatile memory.
- Non-volatile memory may include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
- Volatile memory may include random access memory (RAM) or external cache memory.
- RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain Channel (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
- SRAM static RAM
- DRAM dynamic RAM
- SDRAM synchronous DRAM
- DDRSDRAM double data rate SDRAM
- ESDRAM enhanced SDRAM
- SLDRAM synchronous chain Channel
- memory bus Radbus direct RAM
- RDRAM direct memory bus dynamic RAM
- RDRAM memory bus dynamic RAM
- the blockchain referred to in this application is a new application mode of computer technology such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm.
- Blockchain essentially a decentralized database, is a series of data blocks associated with cryptographic methods. Each data block contains a batch of network transaction information for verification. The validity of the information (anti-counterfeiting) and the generation of the next block.
- the blockchain can include the underlying platform of the blockchain, the platform product service layer, and the application service layer.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Business, Economics & Management (AREA)
- Marketing (AREA)
- Animal Behavior & Ethology (AREA)
- Strategic Management (AREA)
- General Business, Economics & Management (AREA)
- Quality & Reliability (AREA)
- Operations Research (AREA)
- Life Sciences & Earth Sciences (AREA)
- Tourism & Hospitality (AREA)
- Computational Linguistics (AREA)
- Human Resources & Organizations (AREA)
- Entrepreneurship & Innovation (AREA)
- Economics (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Telephonic Communication Services (AREA)
Abstract
一种基于Neo4j的异常用户处理方法、装置、计算机设备及可读存储介质,涉及了大数据和数据挖掘领域。所述方法包括对获取到的用户数据进行用户特征提取(202);将提取到的用户特征输入到Neo4j算法中预测得到用户数据中具有相同属性标签的用户群体,作为异常用户群体(204);基于预设权重比例,根据异常用户群体中的用户数量、标定特征以及共同特征对异常用户群体进行评分(206);根据评分为异常用户群体生成对应的靶向标签,根据靶向标签对检测到的具有特定操作的异常用户群体进行处理(208)。解决了对异常用户的处理不准确的技术问题。
Description
本申请以2020年6月24日提交的申请号为202010591171.1,名称为“基于Neo4j的异常用户处理方法、装置、计算机设备和介质”的中国发明专利申请为基础,并要求其优先权。
本申请涉及大数据领域,特别是涉及一种基于Neo4j的异常用户处理方法、装置、计算机设备和存储介质。
目前,各大互联网平台为了争夺用户,往往会采用让利补贴的方式提高活跃度,但同时也滋养了一群特殊的用户——“羊毛党”。“羊毛党”逐渐从个体玩家发展为有组织、有规模的职业集体,这些异常用户不仅为企业带来了负担,还为后台服务器的数据处理带来了困难。现有技术中通常会将检测到的异常用户加入黑名单,但发明人意识到这种对异常用户验证处理的方式过于简单粗暴,这就会导致一类倾向于从异常用户转为正常用户的用户也无法享受正常的权限,给异常用户的验证处理造成极大的不准确性。
发明内容
基于此,有必要针对上述技术问题,本申请提供一种基于Neo4j的异常用户处理方法、装置、计算机设备及存储介质,以解决现有技术中对异常用户的处理不准确的技术问题。
一种基于Neo4j的异常用户处理方法,所述方法包括:
对获取到的用户数据进行用户特征提取;
将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;
基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;并
根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
一种基于Neo4j的异常用户处理装置,所述装置包括:
提取模块,用于对获取到的用户数据进行用户特征提取;
检测模块,用于将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;
评分模块,用于基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及
处理模块,用于根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
一种计算机设备,包括存储器和处理器,以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现下述基于Neo4j的异常用户处理方法的步骤:
对获取到的用户数据进行用户特征提取;
将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;
基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及
根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
一种计算机可读存储介质,所述计算机可读存储介质存储有计算机可读指令,所述计算机可读指令被处理器执行时实现下述基于Neo4j的异常用户处理方法的步骤:对获取到的用户数据进行用户特征提取;
将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;
基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及
根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
上述基于Neo4j的异常用户处理方法、装置、计算机设备和存储介质,通过根据异常用户群体中的用户数量、特征对异常用户进行评级分级,然后根据评级分级生成靶向标签,当异常用户触发特定操作后根据靶向标签触发不同的处理机制,对异常用户进行针对性的验证处理,解决了现有技术中对异常用户的验证处理不准确的技术问题。
为了更清楚地说明本申请实施例的技术方案,下面将对本申请实施例的描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为基于Neo4j的异常用户处理方法的应用环境示意图;
图2为基于Neo4j的异常用户处理方法的流程示意图;
图3为图2中步骤204的流程示意图;
图4a为步骤302中一个关系图谱示意图;
图4b为步骤302中另一个关系图谱示意图;
图4c为步骤302中的另一关系图谱示意图;
图5为图2中步骤206的流程示意图;
图6为基于Neo4j的异常用户处理装置的示意图;
图7为一个实施例中计算机设备的示意图。
除非另有定义,本文所使用的所有的技术和科学术语与属于本申请的技术领域的技术人员通常理解的含义相同;本文中在申请的说明书中所使用的术语只是为了描述具体的实施例的目的,不是旨在于限制本申请;本申请的说明书和权利要求书及上述附图说明中的术语“包括”和“具有”以及它们的任何变形,意图在于覆盖不排他的包含。本申请的说明书和权利要求书或上述附图中的术语“第一”、“第二”等是用于区别不同对象,而不是用于描述特定顺序。
在本文中提及“实施例”意味着,结合实施例描述的特定特征、结构或特性可以包含在本申请的至少一个实施例中。在说明书中的各个位置出现该短语并不一定均是指相同的实施例,也不是与其它实施例互斥的独立的或备选的实施例。本领域技术人员显式地和隐式地理解的是,本文所描述的实施例可以与其它实施例相结合。
为了使本申请的目的、技术方案及优点更加清楚明白,下面结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本申请实施例提供的基于Neo4j的异常用户处理方法,可以应用于如图1所示的应用环境中。其中,该应用环境可以包括终端102、网络以及服务端104,网络用于在终端102和服 务端104之间提供通信链路介质,网络可以包括各种连接类型,例如有线、无线通信链路或者光纤电缆等等。
用户可以使用终端102通过网络与服务端104交互,以接收或发送消息等。终端102上可以安装有各种通讯客户端应用,例如网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
终端102可以是具有显示屏并且支持网页浏览的各种电子设备,包括但不限于智能手机、平板电脑、电子书阅读器、MP3播放器(Moving Picture Experts Group Audio Layer III,动态影像专家压缩标准音频层面3)、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器、膝上型便携计算机和台式计算机等等。
服务端104可以是提供各种服务的服务器,例如对终端102上显示的页面提供支持的后台服务器。
需要说明的是,本申请实施例所提供的基于Neo4j的异常用户处理方法一般由服务端/终端执行,相应地,基于Neo4j的异常用户处理装置一般设置于服务端/终端设备中。
应该理解,图1中的终端、网络和服务端的数目仅仅是示意性的。根据实现需要,可以具有任意数目的终端设备、网络和服务器。
其中,终端102通过网络与服务端104进行通信。服务端104通过网络从终端102获取用户数据,并对用户数据进行特征提取,将用户特征输入Neo4j算法预测出异常用户群体,然后根据异常用户群体中的用户数量、标定特征以及共同特征对异常用户群体中进行分评分,并根据评分生成对应的靶向标签,最后根据靶向标签对异常用户群体中的异常用户进行对应的验证处理。其中,终端102和服务端104之间通过网络进行连接,该网络可以是有线网络或者无线网络,终端102可以但不限于是各种个人计算机、笔记本电脑、智能手机、平板电脑和便携式可穿戴设备,服务端104可以用独立的服务器或者是多个组成的服务器集群来实现。
在一个实施例中,如图2所示,提供了一种基于Neo4j的异常用户处理方法,以该方法应用于图1中的服务端为例进行说明,包括以下步骤:
步骤202,对获取到的用户数据进行用户特征提取。
用户数据包括用户的手机号码、IP地址、设备号以及手机号码、IP地址、设备号之间的联系、不同用户的用户特征之间的关系等等,另外,还包括不同的用户所参与的活动类型,比如每日签到、新人注册、邀新注册等等。其中,一个用户可以使用一个手机号、多个IP地址和设备号。
用户特征提取是为了提取出以用户的手机号码、IP地址、设备号为维度的特征值,以及手机号码、IP地址、设备号三种特征之间的关联关系的过程。
需要强调的是,为进一步保证上述用户数据的私密和安全性,上述用户数据还可以分布式存储于区块链的节点中。
步骤204,将提取到的用户特征输入到Neo4j算法中进行预测,得到用户数据中具有相同属性标签的用户群体,作为异常用户群体。
Neo4j是一个高性能的,NOSQL图形数据库,它将结构化数据存储在网络上而不是表中。标签传播算法(Label Propagation)是Neo4j中通过节点之间的边传播标签的算法。
用户的手机号码、IP地址以及设备号都有一个标注特征,标注该手机号码、IP地址或者设备号是否在服务端的黑名单、白名单或者未确认名单中中,其中,一个用户的手机号码、与该手机号码对应的设备号可以在白名单、但是其所对应的IP地址可以在黑名单中;也有可能是该用户的手机号码、IP地址以及设备号都在黑名单中。
黑名单和白名单中的数据是根据历史经验获取的,黑名单数据有些是用户举报、拉黑得到的,白名单是根据用户在某APP中的活动参与量等特征得到的。
将提取得到的用户特征输入到标签传播算法中,标签传播算法会以手机号码、IP以及设备号这三类作为关系节点,基于三类关系节点之间的联系构建关系网络。
其中,两关系节点之间的关系,在关系网络中表现为连接两关系节点的边。用户数据中包括实时用户数据,可以将实时用户数据直接应用实时采集到的用户数据,保证异常用户预测的及时性。然后根据构建的关系网络预测得到异常用户群体。
本实施例通过Neo4j算法进行异常用户群体的预测,无需根据既定的节点模板得到风险图,使得预测过程的灵活性更强,普适性更强。
步骤206,基于预设权重比例,根据异常用户群体中的用户数量、标定特征以及共同特征对异常用户群体进行评分。
预设权重比例是根据历史经验得到的用于整合异常用户群体各特征值的比例,可以根据具体应用场景设定;标定特征包括但不限于用户的每日签到、新人注册、邀新注册、特定活动等维度的特征;特定活动包括但不限于双十一活动、618活动等对用户来说具有获得高回报卡券、优惠券等虚拟对象的活动;共同特征指的是异常用户群体中的异常用户可能具有的相同特征维度的特征,比如一个异常用户群体中超过50%的异常用户具有相同的设备号、手机归属地、注册时间或相同的IP。
步骤208,根据评分为异常用户群体生成对应的靶向标签,并根据靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
特定操作一般包括:异常用户群体中异常用户的请求登录某app账号的操作、请求获取某app中某功能模块的操作等。不同范围的评分对应不同的靶向标签,不同的靶向标签会触发不同的验证处理机制。比如,当得到某一异常用户群体的评分为91时,可以为该异常用户群体生成一个靶向标签为拒绝登录和/或拒绝参与活动的标签,限制该异常用户群体的活动,比如拒绝登录的验证处理,以实现对不同异常用户群体进行不同验证处理的目的,还可以实现降低异常用户群体产生的冗余数据,从而提高服务端数据处理效率的目的。
上述基于Neo4j的异常用户处理方法中,通过在服务端通过Neo4j算法预测出异常用户群体,然后根据异常用户群体的三类指定的特征对其进行评分,根据评分结果生成对应的靶向标签,根据不同的靶向标签对异常用户群体中异常用户的登录、活动等行为进行不同程度的处理,实现对不同异常用户群体进行不同验证处理目的,还可以实现降低异常用户群体产生的冗余数据,从而提高服务端数据处理效率的目的。
在一个实施例中,如图3所示,步骤204,包括:
步骤302,通过Neo4j算法构建基于各用户特征之间关联关系的关系网络,其中,关系网络主要由关系节点为手机号码、IP地址以及设备号组成。
如图4a、4b关系图谱示意图所示,三角形代表IP地址、圆形代表手机号码、五边形代表设备号,一个IP地址可以对应多个手机号、一个手机号码也可以对应多个设备号、一个设备号也可以对应多个IP地址,通常,设备号与IP地址通过手机号码进行关联,形成二度关系,而手机号码分别与设备号、IP地址进行直接关联,形成一度关系。
其中,图4a是根据正常用户数据(白名单用户、非白非黑未确认节点的用户数据)构建出的关系网络,可以看出在单点特征上正常用户的关系网络较为分散;图4b是根据黑名单用户数据构建的关系网络,在一些单点特征上,同正常用户一样呈现出分散的特点,使得单点特征防御难以奏效,但是如果将用户特征用网络的形式建模展示,会发现在一些特殊的图形特征上,异常行为明显异于正常行为。其中,单点特征指的是单个特征。比如,用户使用的设备号数量、IP数量这是两个特征,单单从一个特征上看,比如设备号数量这一个单点特征看,无法区分异常用户和正常用户。
用户特征是指用户在某个时间(段)做了某一件事情,例如用户在凌晨1点到3点登陆了某APP10次;用户30天内在某APP中签到了20次,这些都属于用户特征。
进一步地,通过Neo4j构建基于用户数据的关系网络如图4c中关系图谱示意图所示,其中,图4c只是示例性的。
步骤304,获取关系节点的标注特征,并根据标注特征对关系节点进行分类,得到黑名单节点、白名单节点以及未确认节点,其中,标注特征为关系节点的初始类别。
标注特征表示用户的手机号码、IP地址或者设备号是属于黑名单中的黑名单节点、白名单中的白名单节点或者未确认名单中的未确认节点中的哪一种。
刚获取到的关系网络都会有一个初始类别,初始类别包括该IP地址、手机号码或设备号属于黑名单、白名单还是未确认名单,其中,可为黑名单的关系节点生成的标签为1,为白名单以及未确认名单的关系节点生成的标签为0。
步骤306,为关系节点生成属性标签,其中,属性标签包括异常、正常。
特别的,黑名单中的手机号码、设备号以及IP地址的标注特征,用于标注各关系节点的初始的属性标签属于异常,还有一个属性标签:black标签;而白名单中的关系节点的标注特征,用于标注各关系节点的初始的属性标签属于正常,且还具有一个属性标签:white标签,black标签和white标签是为了防止黑、白名单中关系节点的属性标签在后续节点标签更新时被刷新;而未确认名单中的关系节点(手机号码、IP地址以及设备号)的属性标签为:unidentified(意为未确认标签),在后续的标签预测中可以被刷新,其初始的标注特征为正常。
即,位于黑名单或者白名单中的IP地址、设备号或手机号码的属性标签最终不会被更新或改变。
步骤308,分别以属性标签为异常的手机号码、IP地址的关系节点为种子节点,从关系网络中获取路径长度不大于第一预设值的节点路径,得到节点关系图谱,其中,关系图谱包括基于异常手机号码节点的异常号码关系图谱和基于异常IP节点的异常IP关系图谱。
以属性标签为异常的IP地址为种子节点,获取路径长度不大于第一预设值的节点路径组成的基于异常IP节点的异常IP关系图谱,其中,第一预设值根据经验可以设置为10。还需要以属性标签为异常的手机号码为种子节点,获取路径长度不大于10的节点路径组成的基于异常手机号码节点的异常号码关系图谱。
具体地,可以通过Neo4j的cypher语言查找路径长度在1到10之间的关系节点组成的节点路径。表达方式可以为(a)-[*1..10]->(b),cypher语言类似SQL查询语句。其中,a、b表示的是关系节点,[*1..10]表示的是1到10度的节点路径。
本实施例以两个不同类型的关系节点作为种子节点进行关系图谱的构建是为了使得预测的图谱更加全面、预测的过程更加高效;若以一个关系节点作为种子节点构建一个大的关系图谱,则需要获取不大于N的节点路径,N为整数;而N需要远远大于10才能够保证大部分的节点连接情况能够被包含在关系图谱中,而这样就会导致在关系图谱构建过程中路径切分耗时过长,影响预测效率。所以使用两个节点构建两个关系图谱,即能够保证关系图谱的完整性,也能够保证切分的效率。
步骤310,分别对异常号码关系图谱与异常IP关系图谱中的未确认节点进行标签预测,得到异常用户群体。
这一步也是在标签传播算法里实现的,在首轮标签更新前,对于未确认名单中的关系节点的标签默认为正常,而其他的关系节点的属性标签要么确定是异常,要么确定是正常。
例如,关系节点A的属性标签已更新完,从正常变成了异常,而关系节点B与关系节点A互为邻居节点,关系节点B的标签更新结果就会受到关系节点A的标签影响。
具体地,对比未确认节点的属性标签为异常、正常的邻居节点的数量,得到标签对比结果。
标签刷新时,对于未确认节点(unidentified节点),根据其邻居节点的标签,对其标签进行更新。
根据标签对比结果更新未确认节点的属性标签。
对于未确认名单中的某个未确认节点,结合该关系节点所有的邻居节点的属性标签,若所有邻居节点的属性标签为正常的大于属性标签为异常的,则为该关系节点生成正常的属性标签,反之,若则生成异常的属性标签。如果,属性标签为正常的数量和为异常的数量的邻居节点一样多,则给该关系节点随机生成标签异常或者正常,这时生成异常或正常的标签对 异常用户群体的预测并无不利影响。
按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,并重复属性标签对比、更新的操作,直到每个未确认节点的属性标签不再改变或者更新次数达到阈值,将最后一次属性标签更新后关系图谱,作为待预测图谱。
其中,预设修正条件是在标签更新过程中对关系节点的标签进行调整的条件。具体地,在每一轮标签变更后,需要对一部分关系节点的标签结果进行修正:即与白名单中IP地址具有一度关系的未确认节点,其属性标签都置为正常。对两个关系图谱中的每个未确认节点,依次按这样的方式,进行标签刷新,直到每个关系节点的标签结果不变或刷新的轮数达到阈值10000。
标签迭代过程中,仅仅针对未确认名单中的关系节点进行迭代,遍历时间短。在对异常号码关系图谱中的属性标签进行预测时,在每次迭代后,对与白名单中IP地址关联的一度节点(即手机号码节点),重新置为正常,在对异常IP关系图谱中的属性标签进行预测时,在每次迭代后,将与白名单手机号关联的一度节点(IP)的标签重置为0,防止使用企业wifi的用户被置为黑名单。提高异常用户群体预测的准确率。
断开待预测图谱中的异常边,得到多个关系子图,其中,异常边为任意两关系节点直接相连组成的边中至少有一个关系节点的属性标签为正常的边。
具体地,将与白名单节点中的IP地址具有一度关系的手机号码的属性标签设为正常,其中,一度关系是指两个直接相连的关系节点的关联关系。
通过该方式生成的关系子图中,所有的关系节点的标签都为正常,以此保证每一个手机号节点的邻居节点的属性标签也都为正常。本实施例依照只有当同一用户的手机号节点、IP节点以及设备号节点的标签全部都是异常的时候,才会将该手机号结点所对应的用户标记为异常用户。
统计关系子图中关系节点的属性标签,若属性标签为异常的手机号码或IP地址的数量大于第二预设值,则将关系子图中的关系节点对应的用户群体作为异常用户群体。
用户是根据手机号码/IP地址来确认的。一般羊毛党的数量都是聚众出现,使用一个手机号码在不同IP地址下、不同设备号上刷数据,还有在同一IP地址下使用不同手机号、不同设备号;其用户欺诈类型又可以分为养号欺诈团伙、活动欺诈团伙,比如某个欺诈团伙,所有的手机号码的在某app上的签到次数为0,参加活动次数为0,则可以将该欺诈团伙定义为养号欺诈团伙;若某个欺诈团伙,所有的手机号参加活动的次数不为0,则该欺诈团伙定义为活动欺诈团伙。
进一步地,还可以比较得到的异常用户群体对异常用户群体进行去重处理,具体地,若基于异常手机号码节点的异常用户群体与基于异常IP节点的异常用户群体相同,则合并异常用户群体,得到合并后的异常用户群体。因为基于不同节点为中心进行群体预测,可能会得到同一用户,即在异常手机号码为中心节点进行预测得到的A群体中,也在异常IP节点为中心得到的A群体中,所以需要查找得到的异常用户群体是否有相同标签,将相同标签的异常用户群体的同一用户,进行合并,以节省存储空间,并给合并后的异常用户群体生成相同的靶向标签,后续当检测到该异常用户群体中的异常用户的特定操作,比如登录操作后,会根据该异常用户的靶向标签,对其进行对应的验证处理。如果没有检测到,则说明预测得到的是不同的异常用户群体,则为不同的异常用户群体生成不同的靶向标签,并当检测到这些异常用户群体中的异常用户的特定操作后,比如登录操作后,根据其对应的靶向标签,调用对应的验证接口对该异常用户进行验证处理。通过本实施例提高了数据存储的效率。
本实施例直接通过关系节点之间互为邻居标签,对需要确定标签的关系节点进行标签更新,不局限于为其中某一指定的关系节点更新标签,除特殊标签(黑名单、白名单)外,所有节点的标签在预测中都在不断被刷新,保证在标签预测时能够充分利用邻居标签的影响力,最后断开存在标签为正常的节点所在的关系边生成子图,使得预测过程的灵活性更强,普适性更强,而且在标签迭代过程中,仅仅针对未确认节点进行迭代,遍历时间短。且迭代一次 后,将与白名单的IP地址关联的一度节点(即手机号码)的标签重新置为正常,防止使用企业wifi的用户被置为黑名单这类情况的错误预测。
在一个实施例中,如图5所示,步骤206,包括:
步骤502,获取异常用户群体的用户数量,并根据第一预设评级分数和用户数量设定异常用户群体的第一评级分数。
用户数量指的是一个异常用户群体中异常用户的个数,可以根据该用户数量确定该异常用户群体的第一评级分数;具体地:
异常用户群体的用户数量与第一预设评级分数如表1:
| 用户数量/个 | 第一预设评级分数/分 |
| <=10 | 10 |
| (10,30] | 20 |
| (30,50] | 40 |
| (50,70] | 60 |
| (70,90] | 80 |
| >90 | 100 |
表1
例如,当用户数量低于10时,对应的第一预设评级分数为10分,则该异常用户群体的第一瓶机分数为10。
步骤504,获取异常用户群体中具有标定特征的异常用户的标定数量,根据标定特征的运算优先级确定异常用户群体的第二评级分数。
标定特征的运算优先级是指示不同标定特征对第二评级分数影响的优先程度。比如每日签到40分,新人注册60分,邀新注册80分,特定活动100分,其中,特定活动的运算优先级大于日签到,新人注册以及邀新注册,邀新注册的运算优先级又大于每日签到和新人注册,而新人注册的又大于每日签到。当一个异常用户群体中超过50%的异常用户参与了每日签到、邀新注册以及特定活动,则以特定活动对应的分数作为该异常用户群体的第二评级分数。其中超过50%即是标定数量。比如,若标定数量低于50%,则可以不为该异常用户群体进行评分,也就是第二评级分数为0。
步骤506,统计异常用户群体中具有共同特征的异常用户的共有数量,并根据共有数量以及第二预设评级分数确定异常用户群体的第三评级分数。
共有数量是异常用户群体中具有某共同特征的异常用户的人数,可以用于评断是否根据共有特征对应的分数确定异常用户群体的第三评级分数。
共同特征是指异常用户具有相同IP、手机归属地、相同注册时间以及相同设备号。
相同IP-40、同手机归属地-60、同注册时间-80、同设备号-100;
若一个异常用户群体有50%异常用户使用相同IP、同注册时间、同设备号,则该异常用户群体的得分依据分数最高的共有特征对应的分数,即同设备号100分。
步骤508,按照预设权重比例联合第一评级分数、第二评级分数以及第三评级分数,得到异常用户群体的评分。
依据以上3种分级分数条件,并按照预设权重比例1:2:2,给异常用户群体打分,其中,预设权重比例是根据经验得到的,以这种比例计算评分的准确率比较高:
比如,异常用户群体A:
第一分级分数:20分*0.2
第二分级分数:80分*0.4
第三分级分数:100分*0.4
则,分数为:
4+32+40=76分
最后根据最终分数为该异常用户群体生成不同的靶向标签;
具体地:
<60 登录时短信验证
[60,80] 语音验证
[81,90] 人脸验证
>91 拒绝登录
不同的靶向标签的异常用户触发服务端不同的验证机制,登录时短信验证的靶向标签会在用户登录时触发短信验证机制,对用户进行验证等等,从而当异常用户登录APP/网页时,采取不同的验证方式,增加异常用户的验证门槛,防止机器批量操作,降低服务端垃圾数据的处理量,提高数据处理效率。
进一步地,若检测到异常用户的特定操作,则获取异常用户所在异常用户群体的靶向标签;该异常用户的靶向标签是在检测到还用户的特定操作之前预先生成的,当检测到该异常用户群体中的某个或某些异常用户的特定操作,比如登录操作后,则会获取该异常用户的靶向标签,并根据靶向标签调用对应的验证处理接口,对该异常用户的活动权限进行验证处理;判断对异常用户的活动权限的验证处理次数是否超过权限值,并将超过权限值的异常用户设为锁定用户。
具体地,针对异常用户,根据他们的得分可以调用短信接口,发送短信验证码。如果在有效时间内,用户使用正确短信登录则登录成功,那么在有效时间内用户下一次登录时,则不会再次调用该短信接口。如果登录失败,则会再一次调用短信接口,直到调用短信接口的次数,即权限值连续3次,则意味着在有效时间内用户都没有输入正确的验证码,则登录失败,锁定账号,并将该异常用户设为锁定用户,并限制其其他操作。
通过本实施例采取不同的验证方式,增加异常用户的验证门槛,防止机器批量操作,降低服务端垃圾数据的处理量,提高了数据处理效率。
应该理解的是,虽然图2-图3、图5的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,这些步骤可以以其它的顺序执行。而且,图2-图3、图5中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,这些子步骤或者阶段的执行顺序也不必是依次进行,而是可以与其它步骤或者其它步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
在一个实施例中,如图6所示,提供了一种基于Neo4j的异常用户处理装置,该基于Neo4j的异常用户处理装置与上述实施例中基于Neo4j的异常用户处理方法一一对应。该基于Neo4j的异常用户处理装置包括:
提取模块602,用于对获取到的用户数据进行用户特征提取。
检测模块604,用于将提取到的用户特征输入到Neo4j算法中进行预测,得到用户数据中具有相同属性标签的用户群体,作为异常用户群体。
评分模块606,用于基于预设权重比例,根据异常用户群体中的用户数量、标定特征以及共同特征对异常用户群体进行评分。以及
处理模块608,用于根据评分为异常用户群体生成对应的靶向标签,并根据靶向标签对检测到的具有特定操作异常的用户群体进行验证处理。
需要强调的是,为进一步保证上述用户数据的私密和安全性,上述用户数据还可以分布式存储于区块链的节点中。
进一步地,检测模块604,包括:
构建子模块,用于通过Neo4j算法构建基于各用户特征之间关联关系的关系网络,其中,关系网络主要由关系节点为手机号码、IP地址以及设备号组成。
分类子模块,用于获取关系节点的标注特征,并根据标注特征对关系节点进行分类,得到黑名单节点、白名单节点以及未确认节点,其中,标注特征为关系节点的初始类别。
标签子模块,用于为关系节点生成属性标签,其中,属性标签包括异常、正常。
筛选子模块,用于分别以属性标签为异常的手机号码、IP地址的关系节点为种子节点,从关系网络中获取路径长度不大于第一预设值的节点路径,得到节点关系图谱,其中,关系图谱包括基于异常手机号码节点的异常号码关系图谱和基于异常IP节点的异常IP关系图谱。
预测子模块,用于分别对异常号码关系图谱与异常IP关系图谱中的未确认节点进行标签预测,得到异常用户群体。
进一步地,预测子模块,包括:
对比单元,用于对比未确认节点的属性标签为异常、正常的邻居节点的数量,得到标签对比结果。
更新单元,用于根据标签对比结果更新未确认节点的属性标签。
修正单元,用于按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,并重复属性标签对比、更新的操作,直到每个未确认节点的属性标签不再改变或者更新次数达到阈值,将最后一次属性标签更新后关系图谱,作为待预测图谱。
断边单元,用于断开待预测图谱中的异常边,得到多个关系子图,其中,异常边为任意两关系节点直接相连组成的边中至少有一个关系节点的属性标签为正常的边。
统计单元,用于统计关系子图中关系节点的属性标签,若属性标签为异常的手机号码或IP地址的数量大于第二预设值,则将关系子图中的关系节点对应的用户群体作为异常用户群体。
进一步地,预测子模块,还包括:
比较单元,用于比较得到的异常用户群体。
合并单元,用于若基于异常手机号码节点的异常用户群体与基于异常IP节点的异常用户群体相同,则合并异常用户群体,得到合并后的异常用户群体。
进一步地,评分模块606,包括:
第一分数子模块,用于获取异常用户群体的用户数量,并根据第一预设评级分数和用户数量设定异常用户群体的第一评级分数;
第二分数子模块,用于获取异常用户群体中具有标定特征的异常用户的标定数量,根据标定特征的运算优先级确定异常用户群体的第二评级分数;
第三分数子模块,用于统计异常用户群体中具有共同特征的异常用户的共有数量,并根据共有数量以及第二预设评级分数确定异常用户群体的第三评级分数;
综合分数子模块,用于按照预设权重比例联合第一评级分数、第二评级分数以及第三评级分数,得到异常用户群体的评分。
进一步地,处理模块608,包括;
获取子模块,用于若检测到异常用户的特定操作,则获取异常用户所在异常用户群体的靶向标签;以及
调用子模块,用于根据接口调用类型调用对应的接口对异常用户的活动权限进行处理;以及
锁定子模块,用于判断对异常用户的活动权限的处理次数是否超过权限值,并将超过权限值的异常用户设为锁定用户。
上述基于Neo4j的异常用户处理装置,通过根据异常用户群体中的用户数量、特征对异常用户进行评级分级,然后根据评级分级生成靶向标签,当异常用户触发特定操作后根据靶向标签触发不同的处理机制,对异常用户进行针对性的验证处理,解决了现有技术中对异常用户的处理不准确的技术问题。
在一个实施例中,提供了一种计算机设备,该计算机设备可以是服务器,其内部结构图可以如图7所示。该计算机设备包括通过系统总线连接的处理器、存储器、网络接口和数据库。其中,该计算机设备的处理器用于提供计算和控制能力。该计算机设备的存储器包括非易失性存储介质、内存储器。该非易失性存储介质存储有操作系统、计算机可读指令和数据 库。该内存储器为非易失性存储介质中的操作系统和计算机可读指令的运行提供环境。该计算机设备的数据库用于存储用户数据。该计算机设备的网络接口用于与外部的终端通过网络连接通信。该计算机可读指令被处理器执行时以实现一种基于Neo4j的异常用户处理方法。
其中,本技术领域技术人员可以理解,这里的计算机设备是一种能够按照事先设定或存储的指令,自动进行数值计算和/或信息处理的设备,其硬件包括但不限于微处理器、专用集成电路(Application Specific Integrated Circuit,ASIC)、可编程门阵列(Field-Programmable Gate Array,FPGA)、数字处理器(Digital Signal Processor,DSP)、嵌入式设备等。通过根据异常用户群体中的用户数量、特征对异常用户进行评级分级,然后根据评级分级生成靶向标签,当异常用户触发特定操作后根据靶向标签触发不同的处理机制,对异常用户进行针对性的验证处理,解决了现有技术中对异常用户的验证处理不准确的技术问题。
在一个实施例中,提供了一种计算机可读存储介质,其上存储有计算机可读指令,计算机可读指令被处理器执行时实现上述实施例中基于Neo4j的异常用户处理方法的步骤,例如图2所示的步骤202至步骤208,或者,处理器执行计算机可读指令时实现上述实施例中基于Neo4j的异常用户处理装置的各模块/单元的功能,例如图6所示模块602至模块608的功能。为避免重复,此处不再赘述。通过根据异常用户群体中的用户数量、特征对异常用户进行评级分级,然后根据评级分级生成靶向标签,当异常用户触发特定操作后根据靶向标签触发不同的处理机制,对异常用户进行针对性的验证处理,解决了现有技术中对异常用户的验证处理不准确的技术问题。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,所述的计算机可读指令可存储于
一非易失性计算机可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
本申请所指区块链是分布式数据存储、点对点传输、共识机制、加密算法等计算机技术的新型应用模式。区块链(Blockchain),本质上是一个去中心化的数据库,是一串使用密码学方法相关联产生的数据块,每一个数据块中包含了一批次网络交易的信息,用于验证其信息的有效性(防伪)和生成下一个区块。区块链可以包括区块链底层平台、平台产品服务层以及应用服务层等。
所属领域的技术人员可以清楚地了解到,为了描述的方便和简洁,仅以上述各功能单元、模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能单元、模块完成,即将所述装置的内部结构划分成不同的功能单元或模块,以完成以上描述的全部或者部分功能。
以上实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上所述实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对发明专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形、改进或者对部分技术特征进行等同替换,而这些修改或者替换,并不使相同技术方案的本质脱离本申请个实施例技术方案地精神和范 畴,都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。
Claims (22)
- 一种基于Neo4j的异常用户处理方法,所述方法包括:对获取到的用户数据进行用户特征提取;将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
- 根据权利要求1所述的方法,其中,所述用户特征包括用户的手机号码、IP地址、设备号以及各所述用户特征之间的关联关系,所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体,包括:通过Neo4j算法构建基于各所述用户特征之间关联关系的关系网络,其中,所述关系网络主要由关系节点为所述手机号码、所述IP地址以及所述设备号组成;获取所述关系节点的标注特征,并根据所述标注特征对所述关系节点进行分类,得到黑名单节点、白名单节点以及未确认节点,其中,所述标注特征为所述关系节点的初始类别;为所述关系节点生成属性标签,其中,所述属性标签包括异常、正常;分别以所述属性标签为异常的手机号码、IP地址的关系节点为种子节点,从所述关系网络中获取路径长度不大于第一预设值的节点路径,得到节点关系图谱,其中,所述关系图谱包括基于异常手机号码节点的异常号码关系图谱和基于异常IP节点的异常IP关系图谱;分别对所述异常号码关系图谱与所述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体。
- 根据权利要求2所述的方法,其中,所述分别对所述异常号码关系图谱与所述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体,包括:对比所述未确认节点的属性标签为异常、正常的邻居节点的数量,得到标签对比结果;根据所述标签对比结果更新所述未确认节点的属性标签;按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,并重复属性标签对比、更新的操作,直到每个所述未确认节点的属性标签不再改变或者更新次数达到阈值,将最后一次属性标签更新后关系图谱,作为待预测图谱;断开所述待预测图谱中的异常边,得到多个关系子图,其中,所述异常边为任意两关系节点直接相连组成的边中至少有一个关系节点的属性标签为正常的边;统计所述关系子图中关系节点的属性标签,若属性标签为异常的手机号码或IP地址的数量大于第二预设值,则将所述关系子图中的关系节点对应的用户群体作为所述异常用户群体。
- 根据权利要求3所述的方法,其中,所述按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,包括:将与所述白名单节点中的IP地址具有一度关系的手机号码的属性标签设为正常,其中,所述一度关系是指两个直接相连的关系节点的关联关系。
- 根据权利要求2所述的方法,其中,在所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体之后,还包括:比较得到的异常用户群体;若基于异常手机号码节点的异常用户群体与基于异常IP节点的异常用户群体相同,则合并所述异常用户群体,得到合并后的异常用户群体。
- 根据权利要求1所述的方法,其中,所述标定特征包括至少一个用户行为特征,不同所述用户行为特征具有不同的运算优先级,所述基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分,包括:获取所述异常用户群体的用户数量,并根据第一预设评级分数和所述用户数量设定所述 异常用户群体的第一评级分数;获取所述异常用户群体中具有所述标定特征的异常用户的标定数量,根据所述标定特征的运算优先级确定所述异常用户群体的第二评级分数;统计所述异常用户群体中具有所述共同特征的异常用户的共有数量,并根据所述共有数量以及第二预设评级分数确定所述异常用户群体的第三评级分数;按照所述预设权重比例联合所述第一评级分数、所述第二评级分数以及所述第三评级分数,得到所述异常用户群体的评分。
- 根据权利要求1所述的方法,其中,所述靶向标签包括接口调用类型、权限值,所述根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理,包括:若检测到异常用户的所述特定操作,则获取所述异常用户所在异常用户群体的靶向标签;并根据所述接口调用类型调用对应的接口对所述异常用户的活动权限进行处理;并判断对所述异常用户的活动权限的处理次数是否超过所述权限值,并将超过所述权限值的异常用户设为锁定用户。
- 一种基于Neo4j的异常用户处理装置,包括:提取模块,用于对获取到的用户数据进行用户特征提取;检测模块,用于将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;评分模块,用于基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及处理模块,用于根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
- 一种计算机设备,包括存储器和处理器,所述存储器存储有计算机可读指令,所述处理器执行所述计算机可读指令时实现如下基于Neo4j的异常用户处理方法的步骤:对获取到的用户数据进行用户特征提取;将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理。
- 根据权利要求9所述的计算机设备,其中,所述用户特征包括用户的手机号码、IP地址、设备号以及各所述用户特征之间的关联关系,所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体,包括:通过Neo4j算法构建基于各所述用户特征之间关联关系的关系网络,其中,所述关系网络主要由关系节点为所述手机号码、所述IP地址以及所述设备号组成;获取所述关系节点的标注特征,并根据所述标注特征对所述关系节点进行分类,得到黑名单节点、白名单节点以及未确认节点,其中,所述标注特征为所述关系节点的初始类别;为所述关系节点生成属性标签,其中,所述属性标签包括异常、正常;分别以所述属性标签为异常的手机号码、IP地址的关系节点为种子节点,从所述关系网络中获取路径长度不大于第一预设值的节点路径,得到节点关系图谱,其中,所述关系图谱包括基于异常手机号码节点的异常号码关系图谱和基于异常IP节点的异常IP关系图谱;分别对所述异常号码关系图谱与所述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体。
- 根据权利要求10所述的计算机设备,其中,所述分别对所述异常号码关系图谱与所 述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体,包括:对比所述未确认节点的属性标签为异常、正常的邻居节点的数量,得到标签对比结果;根据所述标签对比结果更新所述未确认节点的属性标签;按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,并重复属性标签对比、更新的操作,直到每个所述未确认节点的属性标签不再改变或者更新次数达到阈值,将最后一次属性标签更新后关系图谱,作为待预测图谱;断开所述待预测图谱中的异常边,得到多个关系子图,其中,所述异常边为任意两关系节点直接相连组成的边中至少有一个关系节点的属性标签为正常的边;统计所述关系子图中关系节点的属性标签,若属性标签为异常的手机号码或IP地址的数量大于第二预设值,则将所述关系子图中的关系节点对应的用户群体作为所述异常用户群体。
- 根据权利要求11所述的计算机设备,其中,所述按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,包括:将与所述白名单节点中的IP地址具有一度关系的手机号码的属性标签设为正常,其中,所述一度关系是指两个直接相连的关系节点的关联关系。
- 根据权利要求10所述的计算机设备,其中,在所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体之后,还包括:比较得到的异常用户群体;若基于异常手机号码节点的异常用户群体与基于异常IP节点的异常用户群体相同,则合并所述异常用户群体,得到合并后的异常用户群体。
- 根据权利要求9所述的计算机设备,其中,所述标定特征包括至少一个用户行为特征,不同所述用户行为特征具有不同的运算优先级,所述基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分,包括:获取所述异常用户群体的用户数量,并根据第一预设评级分数和所述用户数量设定所述异常用户群体的第一评级分数;获取所述异常用户群体中具有所述标定特征的异常用户的标定数量,根据所述标定特征的运算优先级确定所述异常用户群体的第二评级分数;统计所述异常用户群体中具有所述共同特征的异常用户的共有数量,并根据所述共有数量以及第二预设评级分数确定所述异常用户群体的第三评级分数;按照所述预设权重比例联合所述第一评级分数、所述第二评级分数以及所述第三评级分数,得到所述异常用户群体的评分。
- 根据权利要求9所述的计算机设备,其中,所述靶向标签包括接口调用类型、权限值,所述根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理,包括:若检测到异常用户的所述特定操作,则获取所述异常用户所在异常用户群体的靶向标签;并根据所述接口调用类型调用对应的接口对所述异常用户的活动权限进行处理;并判断对所述异常用户的活动权限的处理次数是否超过所述权限值,并将超过所述权限值的异常用户设为锁定用户。
- 一种计算机可读存储介质,其上存储有计算机可读指令,所述计算机可读指令被处理器执行时实现如下基于Neo4j的异常用户处理方法的步骤:对获取到的用户数据进行用户特征提取;将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体;基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分;以及根据评分为异常用户群体生成对应的靶向标签,并根据所述靶向标签对检测到的具有特 定操作的异常用户群体进行验证处理。
- 根据权利要求16所述的计算机可读存储介质,其中,所述用户特征包括用户的手机号码、IP地址、设备号以及各所述用户特征之间的关联关系,所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体,包括:通过Neo4j算法构建基于各所述用户特征之间关联关系的关系网络,其中,所述关系网络主要由关系节点为所述手机号码、所述IP地址以及所述设备号组成;获取所述关系节点的标注特征,并根据所述标注特征对所述关系节点进行分类,得到黑名单节点、白名单节点以及未确认节点,其中,所述标注特征为所述关系节点的初始类别;为所述关系节点生成属性标签,其中,所述属性标签包括异常、正常;分别以所述属性标签为异常的手机号码、IP地址的关系节点为种子节点,从所述关系网络中获取路径长度不大于第一预设值的节点路径,得到节点关系图谱,其中,所述关系图谱包括基于异常手机号码节点的异常号码关系图谱和基于异常IP节点的异常IP关系图谱;分别对所述异常号码关系图谱与所述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体。
- 根据权利要求17所述的计算机可读存储介质,其中,所述分别对所述异常号码关系图谱与所述异常IP关系图谱中的未确认节点进行标签预测,得到所述异常用户群体,包括:对比所述未确认节点的属性标签为异常、正常的邻居节点的数量,得到标签对比结果;根据所述标签对比结果更新所述未确认节点的属性标签;按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,并重复属性标签对比、更新的操作,直到每个所述未确认节点的属性标签不再改变或者更新次数达到阈值,将最后一次属性标签更新后关系图谱,作为待预测图谱;断开所述待预测图谱中的异常边,得到多个关系子图,其中,所述异常边为任意两关系节点直接相连组成的边中至少有一个关系节点的属性标签为正常的边;统计所述关系子图中关系节点的属性标签,若属性标签为异常的手机号码或IP地址的数量大于第二预设值,则将所述关系子图中的关系节点对应的用户群体作为所述异常用户群体。
- 根据权利要求18所述的计算机可读存储介质,其中,所述按照预设修正条件对属性标签更新后的关系节点进行属性标签的修正,包括:将与所述白名单节点中的IP地址具有一度关系的手机号码的属性标签设为正常,其中,所述一度关系是指两个直接相连的关系节点的关联关系。
- 根据权利要求17所述的计算机可读存储介质,其中,在所述将提取到的用户特征输入到Neo4j算法中进行预测,得到所述用户数据中具有相同属性标签的用户群体,作为异常用户群体之后,还包括:比较得到的异常用户群体;若基于异常手机号码节点的异常用户群体与基于异常IP节点的异常用户群体相同,则合并所述异常用户群体,得到合并后的异常用户群体。
- 根据权利要求16所述的计算机可读存储介质,其中,所述标定特征包括至少一个用户行为特征,不同所述用户行为特征具有不同的运算优先级,所述基于预设权重比例,根据所述异常用户群体中的用户数量、标定特征以及共同特征对所述异常用户群体进行评分,包括:获取所述异常用户群体的用户数量,并根据第一预设评级分数和所述用户数量设定所述异常用户群体的第一评级分数;获取所述异常用户群体中具有所述标定特征的异常用户的标定数量,根据所述标定特征的运算优先级确定所述异常用户群体的第二评级分数;统计所述异常用户群体中具有所述共同特征的异常用户的共有数量,并根据所述共有数量以及第二预设评级分数确定所述异常用户群体的第三评级分数;按照所述预设权重比例联合所述第一评级分数、所述第二评级分数以及所述第三评级分数,得到所述异常用户群体的评分。
- 根据权利要求16所述的计算机可读存储介质,其中,所述靶向标签包括接口调用类型、权限值,所述根据所述靶向标签对检测到的具有特定操作的异常用户群体进行验证处理,包括:若检测到异常用户的所述特定操作,则获取所述异常用户所在异常用户群体的靶向标签;并根据所述接口调用类型调用对应的接口对所述异常用户的活动权限进行处理;并判断对所述异常用户的活动权限的处理次数是否超过所述权限值,并将超过所述权限值的异常用户设为锁定用户。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010591171.1 | 2020-06-24 | ||
| CN202010591171.1A CN111814064B (zh) | 2020-06-24 | 2020-06-24 | 基于Neo4j的异常用户处理方法、装置、计算机设备和介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021135540A1 true WO2021135540A1 (zh) | 2021-07-08 |
Family
ID=72855059
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/122828 Ceased WO2021135540A1 (zh) | 2020-06-24 | 2020-10-22 | 基于Neo4j的异常用户处理方法、装置、计算机设备和介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN111814064B (zh) |
| WO (1) | WO2021135540A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113902060A (zh) * | 2021-11-11 | 2022-01-07 | 上海识装信息科技有限公司 | 一种团体用户识别方法、装置、设备及存储介质 |
| CN114265835A (zh) * | 2021-10-28 | 2022-04-01 | 深圳永安在线科技有限公司 | 基于图挖掘的数据分析方法、装置及相关设备 |
| CN119728313A (zh) * | 2025-03-03 | 2025-03-28 | 深圳市悦道科技有限公司 | 一种基于通信数据处理的网络安全管理方法 |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112261484B (zh) * | 2020-12-21 | 2021-04-27 | 武汉斗鱼鱼乐网络科技有限公司 | 一种目标用户识别方法、装置、电子设备和存储介质 |
| CN112699217B (zh) * | 2020-12-29 | 2023-04-18 | 西安九索数据技术股份有限公司 | 一种基于用户文本数据和通讯数据的行为异常用户识别方法 |
| CN113760674A (zh) * | 2021-01-15 | 2021-12-07 | 北京京东拓先科技有限公司 | 信息生成方法、装置、电子设备和计算机可读介质 |
| CN113469696B (zh) * | 2021-06-29 | 2024-08-13 | 中国银联股份有限公司 | 一种用户异常度评估方法、装置及计算机可读存储介质 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106301978A (zh) * | 2015-05-26 | 2017-01-04 | 阿里巴巴集团控股有限公司 | 团伙成员账号的识别方法、装置及设备 |
| CN107104973A (zh) * | 2017-05-09 | 2017-08-29 | 北京潘达互娱科技有限公司 | 用户行为的校验方法及装置 |
| CN108446821A (zh) * | 2018-02-07 | 2018-08-24 | 中国平安人寿保险股份有限公司 | 风险监控的方法、装置、存储介质及终端 |
| CN109284380A (zh) * | 2018-09-25 | 2019-01-29 | 平安科技(深圳)有限公司 | 基于大数据分析的非法用户识别方法及装置、电子设备 |
| CN109831459A (zh) * | 2019-03-22 | 2019-05-31 | 百度在线网络技术(北京)有限公司 | 安全访问的方法、装置、存储介质和终端设备 |
| CN110297912A (zh) * | 2019-05-20 | 2019-10-01 | 平安科技(深圳)有限公司 | 欺诈识别方法、装置、设备及计算机可读存储介质 |
| US20200195693A1 (en) * | 2018-12-14 | 2020-06-18 | Forcepoint, LLC | Security System Configured to Assign a Group Security Policy to a User Based on Security Risk Posed by the User |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106708844A (zh) * | 2015-11-12 | 2017-05-24 | 阿里巴巴集团控股有限公司 | 一种用户群体的划分方法和装置 |
| CN110019547B (zh) * | 2017-11-10 | 2021-07-16 | 平安普惠企业管理有限公司 | 获取客户间的关联关系的方法、装置、设备及介质 |
| CN107943879A (zh) * | 2017-11-14 | 2018-04-20 | 上海维信荟智金融科技有限公司 | 基于社交网络的欺诈团体检测方法及系统 |
| CN108764917A (zh) * | 2018-05-04 | 2018-11-06 | 阿里巴巴集团控股有限公司 | 一种欺诈团伙的识别方法和装置 |
| CN109902486A (zh) * | 2019-01-24 | 2019-06-18 | 平安科技(深圳)有限公司 | 电子装置、异常用户处理策略智能决策方法及存储介质 |
| CN110223168B (zh) * | 2019-06-24 | 2022-06-28 | 浪潮卓数大数据产业发展有限公司 | 一种基于企业关系图谱的标签传播反欺诈检测方法及系统 |
| CN110493181B (zh) * | 2019-07-05 | 2023-04-07 | 中国平安财产保险股份有限公司 | 用户行为检测方法、装置、计算机设备及存储介质 |
-
2020
- 2020-06-24 CN CN202010591171.1A patent/CN111814064B/zh active Active
- 2020-10-22 WO PCT/CN2020/122828 patent/WO2021135540A1/zh not_active Ceased
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106301978A (zh) * | 2015-05-26 | 2017-01-04 | 阿里巴巴集团控股有限公司 | 团伙成员账号的识别方法、装置及设备 |
| CN107104973A (zh) * | 2017-05-09 | 2017-08-29 | 北京潘达互娱科技有限公司 | 用户行为的校验方法及装置 |
| CN108446821A (zh) * | 2018-02-07 | 2018-08-24 | 中国平安人寿保险股份有限公司 | 风险监控的方法、装置、存储介质及终端 |
| CN109284380A (zh) * | 2018-09-25 | 2019-01-29 | 平安科技(深圳)有限公司 | 基于大数据分析的非法用户识别方法及装置、电子设备 |
| US20200195693A1 (en) * | 2018-12-14 | 2020-06-18 | Forcepoint, LLC | Security System Configured to Assign a Group Security Policy to a User Based on Security Risk Posed by the User |
| CN109831459A (zh) * | 2019-03-22 | 2019-05-31 | 百度在线网络技术(北京)有限公司 | 安全访问的方法、装置、存储介质和终端设备 |
| CN110297912A (zh) * | 2019-05-20 | 2019-10-01 | 平安科技(深圳)有限公司 | 欺诈识别方法、装置、设备及计算机可读存储介质 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114265835A (zh) * | 2021-10-28 | 2022-04-01 | 深圳永安在线科技有限公司 | 基于图挖掘的数据分析方法、装置及相关设备 |
| CN113902060A (zh) * | 2021-11-11 | 2022-01-07 | 上海识装信息科技有限公司 | 一种团体用户识别方法、装置、设备及存储介质 |
| CN119728313A (zh) * | 2025-03-03 | 2025-03-28 | 深圳市悦道科技有限公司 | 一种基于通信数据处理的网络安全管理方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN111814064A (zh) | 2020-10-23 |
| CN111814064B (zh) | 2024-09-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2021135540A1 (zh) | 基于Neo4j的异常用户处理方法、装置、计算机设备和介质 | |
| US12381798B2 (en) | Systems and methods for conducting more reliable assessments with connectivity statistics | |
| US11968105B2 (en) | Systems and methods for social graph data analytics to determine connectivity within a community | |
| US11886555B2 (en) | Online identity reputation | |
| WO2021120676A1 (zh) | 联邦学习网络下的模型训练方法及其相关设备 | |
| CN109165683B (zh) | 基于联邦训练的样本预测方法、装置及存储介质 | |
| US10333964B1 (en) | Fake account identification | |
| CN112699382B (zh) | 物联网网络安全风险的评估方法、装置及计算机存储介质 | |
| US20120072982A1 (en) | Detecting potential fraudulent online user activity | |
| CN104954234B (zh) | 一种微博数据获取方法、装置及舆情分析方法 | |
| CN110148053B (zh) | 用户信贷额度评估方法、装置、电子设备和可读介质 | |
| US9418119B2 (en) | Method and system to determine a category score of a social network member | |
| US11128479B2 (en) | Method and apparatus for verification of social media information | |
| US12166795B2 (en) | Cyber security system and method | |
| WO2023184831A1 (zh) | 确定目标对象的方法及标识关联图的构建方法、装置 | |
| CN108229964B (zh) | 交易行为轮廓构建与认证方法、系统、介质及设备 | |
| CN119835323B (zh) | 用户画像技术的个性化定制推送系统及方法 | |
| CN115238813A (zh) | 一种共享账号的风险评估方法、装置、设备及存储介质 | |
| US11120115B2 (en) | Identification method and apparatus | |
| CN114066535A (zh) | 风险识别方法、装置、设备及可读存储介质 | |
| US9886726B1 (en) | Analyzing social networking groups for detecting social networking spam | |
| CN110457600A (zh) | 查找目标群体的方法、装置、存储介质和计算机设备 | |
| KR102471731B1 (ko) | 사용자를 위한 네트워크 보안 관리 방법 | |
| CN118520163A (zh) | 项目推荐方法及装置 | |
| CN111782967B (zh) | 信息处理方法、装置、电子设备和计算机可读存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20910790 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20910790 Country of ref document: EP Kind code of ref document: A1 |