WO2018103456A1 - 一种基于特征匹配网络的社团划分方法、装置及电子设备 - Google Patents
一种基于特征匹配网络的社团划分方法、装置及电子设备 Download PDFInfo
- Publication number
- WO2018103456A1 WO2018103456A1 PCT/CN2017/105985 CN2017105985W WO2018103456A1 WO 2018103456 A1 WO2018103456 A1 WO 2018103456A1 CN 2017105985 W CN2017105985 W CN 2017105985W WO 2018103456 A1 WO2018103456 A1 WO 2018103456A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- account information
- similarity
- hash
- community
- matching network
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q40/00—Finance; Insurance; Tax strategies; Processing of corporate or income taxes
- G06Q40/03—Credit; Loans; Processing thereof
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/22—Matching criteria, e.g. proximity measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
Definitions
- Embodiments of the present invention relate to the field of data processing, and in particular, to a community partitioning method, apparatus, and electronic device based on a feature matching network.
- credit card cashing refers to cardholders obtaining cash after fraudulent consumer transactions or conspiring with merchants to swipe their cards. After the refund or purchase, it is easy to realize the goods and then sell and obtain cash.
- the fake card fraud refers to the fraudulent behavior of writing magnetic, embossed or lithographically forged a real and valid bank card according to the magnetic stripe information format of the bank card; Fraud refers to the fraudster getting some or all of the information of the real cardholder and impersonating the actual cardholder's change of the account's information for fraudulent purposes.
- Credit card crimes are constantly moving toward high-tech, group, and professional development. The implementation of the case is more concealed and the methods are constantly being refurbished. This poses a threat to the financial security of banks and cardholders and has become an important factor restricting the long-term healthy development of the credit card industry.
- clustering is usually adopted to deal with it.
- there are various defects in adopting this method For example, on the one hand, if data is added to the anti-fraud model later, The anti-fraud model makes it difficult to update the data.
- the nodes can be divided into several classes, the structure within the group and the relationship between the structures are still difficult to describe.
- the embodiment of the invention provides a method, a device and an electronic device for classifying a community based on a feature matching network, which are used to solve the problem in the prior art that if the data is added to the anti-fraud model, the anti-fraud model is difficult to update data and is clustered. After that, the structure within the group and the relationship between the structures are still difficult to describe.
- An embodiment of the present invention provides a community division method based on a feature matching network, including:
- the same account information of the sub-hash vector is divided into the same group
- calculating the similarity between each account information in the same group including:
- the n/m is used as the similarity between the i-th account information and the j-th account information; the i-th account information and the The j-th account information is any one of the account information.
- calculating the similarity between each account information in the same group including:
- the i-th account information and the j-th account information are in the same group, the number of the hash vector of the i-th account information and the hash vector of the j-th account information are the same and the hash vector value is the same. h; the i-th account information and the j-th account information are any one of the account information;
- determining a K-bit hash vector corresponding to each account information according to the preset K hash functions including:
- a feature vector indicating account information wherein c 1 , c 2 ..., c d represent the characteristic attributes of the account information, Represents a non-zero vector randomly selected,
- performing community division on each account information including:
- the calculating the similarity strength of each account information according to the similarity between the account information including:
- w ai,z is The sum of the weights of the edges between any account information ai and the z-th account information.
- the embodiment of the invention further provides a community division device based on a feature matching network, comprising:
- a determining unit configured to determine a K-bit hash vector corresponding to each account information according to a preset K hash function
- a second dividing unit configured to divide account information of the same sub-hash vector into the same group for each class
- a calculating unit configured to calculate a similarity between each account information in the same group
- Forming a network unit if the similarity between the account information is greater than a threshold, establishing an interconnection edge between the account information to form a feature matching network;
- the third dividing unit is configured to perform community division on the account information according to the feature matching network.
- the calculating unit is specifically configured to: if the i-th account information and the j-th account information are in the same group of n, use n/m as the similarity between the i-th account information and the j-th account information. And the i-th account information and the j-th account information are any one of the account information.
- the calculating unit is further configured to: if the i-th account information and the j-th account information are in the same group, the hash vector of the i-th account information and the hash vector of the j-th account information are located a number h of the same bit and a hash vector value; the i-th account information and the j-th account information are any one of the account information;
- the determining unit is configured to determine, according to formula (3), a K-bit hash vector corresponding to each account information.
- a feature vector indicating account information wherein c 1 , c 2 ..., c d represent the characteristic attributes of the account information, Represents a non-zero vector randomly selected,
- the third dividing unit is specifically configured to: (1) divide each account information into different communities in the feature matching network;
- the calculating unit is further configured to calculate, according to formula (4), a similar strength s i,j between the i-th account information and the j-th account information;
- w ai,z is The sum of the weights of the edges between any account information ai and the z-th account information.
- An embodiment of the present invention further provides an electronic device, including:
- At least one processor and,
- the information is divided into the same group; the similarity between the account information in the same group is calculated; if the similarity between the account information is greater than the threshold, the information is established between the account information. Forming a feature matching network by establishing an interconnection edge; and performing community division on the account information according to the feature matching network.
- a community division method, device, and electronic device based on a feature matching network are provided, and a K-bit hash vector corresponding to each account information is determined according to a preset K hash function;
- the similarity degree if the similarity between the account information is greater than the threshold, an interconnection edge is established between each account information to form a feature matching network; and the account information is grouped according to the feature matching network.
- the K-bit hash vector corresponding to each account information is first determined according to the preset K hash functions. For a large number of account information in the network, only two hash values are generated. The Greek function is not enough, so it is determined that the K-bit hash vector corresponding to each account information can cope with complex network account information. Then, for each class, the same account information of the sub-hash vector is divided into a group, and the similarity between any account information in the same group is calculated, which can avoid the calculation of similarity between any account information in the entire network.
- the technical solution of the present invention can effectively reduce the calculation of the similarity between the account information, and only calculate the similarity between the account information in the same group.
- an interconnection edge is established between the account information to form a feature matching network; according to the feature matching network, the account information is divided into groups, which can more accurately target each account.
- the information is divided into associations, which not only makes the association relationship between the associations clear, but also analyzes the classified associations, finds out the abnormal associations, and then performs abnormal account checking on the accounts in the abnormal associations, and more specifically looks for them. Raise fraudulent accounts and improve the efficiency of responding to fraudulent accounts.
- you need to add account information to the classified community you only need to repeat the above simple steps for the added account information, and update the added account information to the corresponding location, and it will not cause update difficulties. problem.
- FIG. 1 is a schematic flowchart of a community partitioning method based on a feature matching network according to an embodiment of the present invention
- FIG. 2 is a flowchart of an overall schematic diagram of the present invention according to an embodiment of the present invention.
- FIG. 3 is a schematic structural diagram of a community division apparatus based on a feature matching network according to an embodiment of the present disclosure
- FIG. 4 is a schematic structural diagram of an electronic device according to an embodiment of the present invention.
- the technical solution of the embodiments of the present invention can be applied to various network fraud scenarios of various banks, such as fraud of credit card products, fraud of bank card products, fraudulent card fraud, fake card fraud, cash fraud, and the like.
- the application scenario of the technical solution of the embodiment of the present invention may also be the discovery of the abnormal account information community, the commonality of discovering specific types of fraud, the discovery of other fraudulent account information according to the fraudulent account information sample, and the help of discovering unknown fraud types.
- FIG. 1 is a schematic flowchart showing a method for community partitioning based on a feature matching network according to an embodiment of the present invention. As shown in FIG. 1 , the method includes the following steps:
- Step S101 Determine a K-bit hash vector corresponding to each account information according to the preset K hash functions
- Step S103 For each class, divide the account information with the same sub-hash vector into the same group;
- Step S104 Calculate the similarity between each account information in the same group
- Step S105 If the similarity between the account information is greater than the threshold, an interconnection edge is established between each account information to form a feature matching network.
- Step S106 Perform community division on each account information according to the feature matching network.
- a K-bit hash vector corresponding to each account information is determined according to a preset K hash function. Specifically, a hash vector can be obtained by processing each preset hash function. Then, according to the preset K hash functions, a K-bit hash vector can be generated, and each account information corresponds to a K-bit hash vector.
- each account information includes multiple feature attributes. If only one hash function is used in the prior art, only one hash function is used, there is a disadvantage that it is insufficient to express a plurality of feature attributes of an account information. Therefore, this step can effectively avoid this disadvantage.
- the value of K can be set according to the specific situation of each account information in the specific implementation. For example, if K can be set to 4, the account information can be represented as a 4-bit hash vector.
- the 4-bit hash vector is divided into two types of sub-hash vectors. The advantage of the division is to reduce the computational complexity for the subsequent calculation of the similarity between the accounts, so as to avoid the hash vector of the account information is not divided in the prior art. However, there is a disadvantage that the calculation of the similarity is directly performed on any two of the account information directly.
- Step S104 Calculate the similarity between the account information in the same group.
- the ratio of the number of bits of the hash vector of each account information in the same group to the size of the bit may be counted, for example, account information.
- the hash vector of 1 is 0010
- the similarity of the information then, the number of bits of the hash vector of the two account information is 3, and the size of the bit is 4 bits, so the similarity between the two account information is 3/4, also
- the similarity between any two account information in the same group can be calculated according to the calculation formula for the similarity.
- the calculation formula of the similarity can be the Euclidean distance, the cosine distance, the Jaccard distance formula, and the like.
- only calculating the similarity between any two account information in the same group can greatly reduce the amount of calculation.
- N account information samples are taken, then N account information samples are grouped into 2 k groups, and the number of account information samples in each group is N/2 k , and any two account information is performed in each group.
- the number of similarity calculations is The number of similarities calculated by 2 k groups for any two account information is Therefore, the number of times that all classes need to be similarly calculated is among them, Is the number of divided classes, this value is a constant that can be controlled according to the actual situation, and the traditional method calculates any two account information in all accounts for similarity calculation needs to be performed.
- the calculation amount of the similarity between any two account information in the same group using the present invention is calculated by the conventional method to calculate the similarity between any two account information in all accounts. Reduce the multiple of the 2 k level.
- the similarity of the account information in each group is relatively large, so the similarity calculation of the account information in the same group can also improve the efficiency and accuracy of network establishment.
- Step S105 If the similarity between the account information is greater than the threshold, an interconnection edge is established between each account information to form a feature matching network. Specifically, if the similarity between any two account information is greater than a threshold, An interconnection edge is established between any two account information, and the weight of the edge is the similarity value between the two account information, and finally forms a feature matching network.
- the threshold value can be selected to select a higher value.
- the feature matching network is easy to perform subsequent calculations.
- the value of the threshold can be adjusted according to actual conditions.
- Step S106 Perform community division on each account information according to the feature matching network. Specifically, according to the similarity value between the calculated account information, the closer the similarity value is, the easier it is to be divided into the same community. After dividing the community, it is easier to check the fraudulent accounts in the network, and the proportion of fraudulent account samples in each community can be calculated. If the proportion is larger, the possibility that the community is an abnormal community is greater, and it can be based on business needs. Conduct relevant investigations, and then calculate the accounts in the abnormal community according to some indicators, find out representative accounts, and then conduct related cases investigation on these representative accounts. Some indicators may be account information in the community.
- Method 1 Optionally, calculating the similarity between each account information in the same group, including: if the i-th account information and the j-th account information are in the same group of n, the n/m is used as the i-th account information.
- the similarity degree with the j-th account information; the i-th account information and the j-th account information are any one of the account information, specifically, two account information, such as account information, are randomly selected in all the account information.
- Method 2 Optionally, calculating the similarity between each account information in the same group, including: if the i-th account information and the j-th account information are in the same group, counting the hash vector of the i-th account information and the j-th The number H of the hash information of the account information is the same and the hash vector value is the same; the i-th account information and the j-th account information are any one of the account information; the i-th account information is similar to the j-th account information.
- Degree s h/K, specifically, if any two account information in all account information, account information 1 and account information 2 are in the same group, and account information 1 and account information 2 are both 4 digits, that is, K is 4, and the first three digits of the account information 1 and the account information 2 are identical, and the fourth digit is different. Then, the similarity s of the account information 1 and the account information 2 is 3/4.
- the above two methods for calculating the similarity between the account information in the same group can be concluded that the first method is the similarity between the calculated two account information in each class, and the second method is the calculation.
- the similarity between the two account information in the same group in each category can be seen.
- the first method is to roughly calculate two accounts.
- the similarity between the class and the class to which the information belongs, and the similarity between the two account information calculated in the second category is more accurate.
- both methods calculate the similarity between any two account information in all account information in the network by using the Euclidean distance formula or the like in the prior art. Significant improvements have been made to further accelerate the establishment of the network.
- determining a K-bit hash vector corresponding to each account information according to the preset K hash functions including: determining a K-bit hash vector corresponding to each account information according to formula (1)
- a feature vector indicating account information wherein c 1 , c 2 ..., c d represent the characteristic attributes of the account information, Represents a non-zero vector randomly selected,
- the default hash function is Is any one of the preset K hash functions, the hash function The value is represented by 0 or 1. That is to say, such a hash function can only generate two hash values, which is obviously insufficient for a large amount of account information, so it is determined according to such a hash function.
- K-bit hash vector for each account Is a K-bit binary number, for example, can be a 6-bit binary number, specifically 010110, then, among them, a feature vector indicating account information, c 1 , c 2 ..., c d represent characteristic attributes of the account information, and the specific account information characteristic attributes may be transaction amount, transaction time, transaction place, number of transaction places, transfer place, transfer amount, transfer number, and the like.
- the feature vector of each account information may be screened to obtain a set of theoretically best feature vectors in a specific implementation. Specifically, the fraud account information sample and the normal account information sample are extracted in a certain period of time, and the extracted The fraudulent account information sample and the normal account information sample are combined into one overall account information sample.
- Feature vector According to the preset K hash functions, the K-bit hash vector corresponding to each account information is determined, and the feature attribute of each account information can be fully extracted and represented by the feature vector, which can cope with the huge amount of account information in the complex network. happensing.
- K-bit hash vector corresponding to each account information The determination is actually obtained through a hash random mapping process.
- hash mapping The main purpose of using hash random mapping here is to enable the feature vector of the account information to be mapped to a uniform representation of 0 or 1, for subsequent processing, rather than simple dimensionality reduction; second, the original feature vector Mapping to the new hash space will make the data with similar eigenvectors similar in the new hash space.
- the probability is: A monotonically increasing mapping relationship from the similarity s to the probability p.
- Table 1 Relationship between account information samples and classes
- the relationship between the account information sample and the class can be expressed as a matrix of K rows and N columns, N represents the number of account information samples taken, c 1 to c N represents N account information samples, and N account information is obtained.
- the community information is divided into account information, including:
- the line of the account information in the node similarity strength matrix is similarly intensityd.
- the order of large to small attempts to transfer the account information to other communities; if the module information is positive from the p-th community to the q-th community, the account information is divided into the q-th community and ends;
- calculating the similarity strength of each account information according to the similarity between the accounts including:
- ⁇ (i) represents the neighbor set of the i-th account information
- ⁇ (i) ⁇ (j) represents the common neighbor set of the i-th account information and the j-th account information
- w ai,z is any account information ai and the first The weight of the side between the z account information.
- step (1) the feature matching network is initialized, and each account information is divided into different communities, and the division in this step may be randomly divided; in step (2), according to formula (2)
- step (2) according to formula (2)
- To calculate the similarity strength of each account information specifically, if the common neighbor of the account information 1 and the account information 2 is the account information 3, the account information 1 and the account information 2 are combined with the weight of the side of the account information 3 is 5, then, The weight of any account information ai and the account information 3 is 5, and thus, the similarity between the account information 1 and the account information 2 is 1/5.
- other account information is also calculated by this method. If four account information samples are taken, after calculation, a 4*4 matrix is formed.
- the matrix is It can be seen from this matrix that the similarity between the account information 1 and the account information 2 is 0.25, the similarity between the account information 1 and the account information 3 is 0.7, and the similarity between the account information 2 and the account information 3 is 0.4; (3) Steps, from the row of the account information in the similarity strength matrix, try to transfer the account information to other communities in order of similar strength, for example, from the first line of the similarity matrix, you want to put the account When the information 1 is divided into other communities, the community information of the account information 3 with the similarity (the largest in the first row) is preferentially selected. If ⁇ Q ⁇ 0, the account information 1 is attempted to be divided into the community in which the account information 4 (0.4 times in the first line) is located.
- the account information 1 is attempted to be divided into the community in which the account information 2 is located. If ⁇ Q ⁇ 0 is still present, the account information 1 is reserved as an independent community, the matrix is not updated, and the calculation of the second line is performed. If ⁇ Q>0 is found during the above-mentioned attempt, for example, the account 1 is preferentially divided into the community where the account information 3 with the similarity is large (the largest in the first row) is located, and ⁇ Q>0 is found, then The attempt is successful and the first line of calculation ends.
- the calculation formula of the module degree difference ⁇ Q To verify whether the above-mentioned attempt to divide the account information is correct, where n represents all the weights in the network, k i represents the weight of the edge connected to the vertex i, and k i, in represents the weight of the account information i within the community.
- ⁇ in indicates the edge weight of the community
- ⁇ tot indicates the weight of the edge connected to the account information inside the community, including the side inside the community and the side outside the community. If ⁇ Q is a positive number, then the division is accepted. If it is not a positive number, give up this division.
- the account information is preferentially divided into the community of the neighbor account information that is most similar to it, which greatly saves the number of attempts of the community division, further improves the speed of the algorithm, and further attempts on the account information. Whether the division is reasonable or not is verified by the modularity difference formula, which more effectively ensures the rationality and accuracy of the attempted division.
- FIG. 2 exemplarily shows a whole schematic flow chart of the present invention, as shown in FIG. 2:
- Step S201 mapping the feature attribute of each account information to a multi-bit hash map vector by using a hash mapping method
- Step S202 classify the hash map vector of each account information.
- Step S203 For each class, divide the same account information of the hash mapping vector into a group;
- Step S204 Perform similarity calculation on any two account information in each group
- Step S205 If the similarity of any two account information in each group is greater than a threshold, the interconnection edge between the two account information is established, and the weight of the edge is similarity, thereby forming a feature matching network, wherein the formed
- the feature matching network is a sparse feature matching network
- Step S206 Perform community division on the feature matching network according to the similarity strength matrix of each account information in the feature matching network.
- the feature attribute of each account information is mapped into a new hash space by a random hash mapping method to form a hash mapping vector of each account information.
- the hash map vector of each account information is classified, and an edge can be established between the account information of high similarity, which effectively avoids the calculation of the similarity between a large number of any two account information, and efficiently establishes for each edge.
- the credible weight value can improve the accuracy and speed of subsequent community division.
- the feature matching network is established according to the similarity of each account information, and then The similarity strength matrix of each account information in the network divides the feature matching network into associations, which not only can effectively detect abnormal communities and carry out targeted measures, but also can detect unknown fraud types, and match the feature matching network through similar strength matrix.
- the account information is preferentially divided into the community with the neighbor account information that is most similar to it, which greatly saves the number of community division attempts and further improves the speed of the algorithm.
- feature matching network through the formation of feature matching network, related accounts The similarity between the information is permanently stored as the weight of the edge. Even if more new account information comes in, it will not affect the original interconnection edge in the network. It only needs to insert the new account information into the original feature matching. In the network.
- the random hash mapping method is first used to classify each account information, and then the similarity calculation is performed with the account information in the class. If the similarity is greater than the threshold, then Add a new edge. Subsequent only need to perform a smaller but more accurate community partitioning algorithm to achieve the function. At the same time, the structure of the feature matching network can more clearly display the association structure within the community and between the communities, which cannot be achieved by the traditional clustering method.
- the device includes a determining unit 301, a first dividing unit 302, a second dividing unit 303, and a calculating unit 304.
- a network unit 305 and a third dividing unit 306 are formed. among them:
- a determining unit 301 configured to determine a K-bit hash vector corresponding to each account information according to the preset K hash functions;
- a second dividing unit 303 configured to divide, for each class, account information with the same sub-hash vector into the same group;
- the calculating unit 304 is configured to calculate a similarity between each account information in the same group;
- Forming a network unit 305 configured to establish an interconnection edge between each account information to form a feature matching network if the similarity between the account information is greater than a threshold;
- the third dividing unit 306 is configured to perform community division on each account information according to the feature matching network.
- the calculating unit 304 is specifically configured to:
- n/m is used as the similarity between the i-th account information and the j-th account information; the i-th account information and the j-th account information are accounts. Any of the information.
- the calculating unit 304 is further specifically configured to:
- the hash number of the i-th account information and the hash vector of the j-th account information are the same number and the hash vector value is the same number h;
- the account information and the j-th account information are any one of the account information;
- the determining unit 301 is configured to:
- a feature vector indicating account information wherein c 1 , c 2 ..., c d represent the characteristic attributes of the account information, Represents a non-zero vector randomly selected,
- the third dividing unit 306 is specifically configured to:
- the calculating unit 304 is further specifically configured to:
- ⁇ (i) represents the neighbor set of the i-th account information
- ⁇ (i) ⁇ (j) represents the common neighbor set of the i-th account information and the j-th account information
- w ai,z is any account information ai and the first The weight of the side between the z account information.
- an electronic device according to an embodiment of the present invention is applicable to the above embodiment of the present invention.
- the electronic device may include one or more processors 410 and a memory 420, and one processor 410 is exemplified in FIG.
- the apparatus for performing the community matching method based on the feature matching network may further include: an input device 430 and an output device 440.
- the processor 410, the memory 420, the input device 430, and the output device 440 may be through a bus or other means Connection, as shown in Figure 4 by bus connection.
- the memory 420 is a non-volatile computer readable storage medium, and can be used for storing a non-volatile software program, a non-volatile computer executable program, and a module, such as a feature matching network-based community partitioning in the embodiment of the present application.
- the corresponding program instruction/module The processor 410 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 420, that is, implementing the community partitioning method based on the feature matching network of the above method embodiments.
- the memory 420 may include a storage program area and a storage data area, wherein the storage program area may store an operating system, an application required for at least one function; the storage data area may store the creation according to the use of the feature matching network based community division device Data, etc.
- memory 420 can include high speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid state storage device.
- the memory 420 can optionally include memory remotely disposed relative to the processor 410, which can be connected to the feature matching network based community partitioning device over a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
- the input device 430 can receive the input digital or character information and generate key signal inputs related to user settings and function control of the feature matching network based community partitioning device.
- Output device 440 can include a display device such as a display screen.
- the one or more modules are stored in the memory 420, and when executed by the one or more processors 410, perform a feature matching network based community partitioning method in any of the above method embodiments.
- the processor 410 is specifically configured to:
- the n/m is used as the similarity between the i-th account information and the j-th account information; the i-th account information and the The j-th account information is any one of the account information.
- the processor 410 is specifically configured to:
- the i-th account information and the j-th account information are in the same group, the number of the hash vector of the i-th account information and the hash vector of the j-th account information are the same and the hash vector value is the same. h; the i-th account information and The jth account information is any one of the account information;
- the processor 410 is specifically configured to:
- a feature vector indicating account information wherein c 1 , c 2 ..., c d represent the characteristic attributes of the account information, Represents a non-zero vector randomly selected,
- the processor 410 is specifically configured to:
- the processor 410 is specifically configured to:
- w ai,z is The sum of the weights of the edges between any account information ai and the z-th account information.
- a community division device based on a feature matching network is provided, and a K-bit hash vector corresponding to each account information is determined according to a preset K hash function;
- the hash vector corresponding to the account information is sequentially divided into class sub-hash vectors; for each class, the account information with the same sub-hash vector is divided into the same group; and the similarity between the account information in the same group is calculated; If the similarity between each account information is greater than the threshold For the value, an interconnection edge is established between each account information to form a feature matching network.
- the feature matching network the community information of each account information is divided according to the similarity between the account information, and the account information is divided into associations.
- the K-bit hash vector corresponding to each account information is first determined according to the preset K hash functions. For a large number of account information in the network, only two hash values are generated. The Greek function is not enough, so it is determined that the K-bit hash vector corresponding to each account information can cope with complex network account information. Then, for each class, the same account information of the sub-hash vector is divided into a group, and the similarity between any account information in the same group is calculated, which can avoid the calculation of similarity between any account information in the entire network.
- the technical solution of the present invention can effectively reduce the calculation of the similarity between the account information, and only calculate the similarity between the account information in the same group.
- an interconnection edge is established between the account information to form a feature matching network; according to the feature matching network, the account information is divided into groups, which can more accurately target each account.
- the information is divided into associations, which not only makes the association relationship between the associations clear, but also analyzes the classified associations, finds out the abnormal associations, and then performs abnormal account checking on the accounts in the abnormal associations, and more specifically looks for them. Raise fraudulent accounts and improve the efficiency of responding to fraudulent accounts.
- you need to add account information to the classified community you only need to repeat the above simple steps for the added account information, and update the added account information to the corresponding location, and it will not cause update difficulties. problem.
- embodiments of the present invention can be provided as a method, or a computer program product. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or a combination of software and hardware. Moreover, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) including computer usable program code.
- a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) including computer usable program code.
- the computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture comprising the instruction device.
- the apparatus implements the functions specified in one or more blocks of a flow or a flow and/or block diagram of the flowchart.
- These computer program instructions can also be loaded onto a computer or other programmable data processing device such that a series of operational steps are performed on a computer or other programmable device to produce computer-implemented processing for execution on a computer or other programmable device.
- the instructions provide steps for implementing the functions specified in one or more of the flow or in a block or blocks of a flow diagram.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Business, Economics & Management (AREA)
- General Engineering & Computer Science (AREA)
- Evolutionary Biology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Finance (AREA)
- Accounting & Taxation (AREA)
- Life Sciences & Earth Sciences (AREA)
- Artificial Intelligence (AREA)
- Economics (AREA)
- General Business, Economics & Management (AREA)
- Technology Law (AREA)
- Strategic Management (AREA)
- Marketing (AREA)
- Development Economics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种基于特征匹配网络的社团划分方法、装置及电子设备,所述方法包括,根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量(S101);将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量(S102);针对每个类,将子哈希向量相同的账号信息划分为同一组(S103);计算同一组内的各账号信息之间的相似度(S104);若各账号信息之间的相似度大于阈值,则在各账号信息之间建立互连边,形成特征匹配网络(S105);根据特征匹配网络,对各账号信息进行社团划分(S106)。该方法可以根据划分后的社团进行社团分析,发现异常社团。
Description
本申请要求在2016年12月6日提交中国专利局、申请号为201611110731.7、发明名称为“一种基于特征匹配网络的社团划分方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本发明实施例涉及数据处理领域,尤其涉及一种基于特征匹配网络的社团划分方法、装置及电子设备。
目前,国内信用卡市场面临的风险形势日益严峻,信用卡套现、伪卡欺诈、盗卡欺诈等案件日益增加,具体的,信用卡套现是指持卡人通过虚假消费交易或与商户合谋刷卡后获取现金,之后退款或购买容易变现商品后变卖获取现金等行为、伪卡欺诈是指按照银行卡的磁条信息格式写磁,凸印或平印伪造真实有效的银行卡进行交易的欺诈行为;盗卡欺诈是指欺诈者获得真实持卡人的部分或者全部信息并假冒真实持卡人对账户的信息进行变更以达到欺诈目的的行为。信用卡犯罪手段不断向着高科技、集团化、专业化发展,案件实施过程更为隐蔽,手法不断翻新,这对银行和持卡人的资金安全构成威胁,成为制约信用卡产业长期健康发展的重要因素。
面对各种各样的欺诈手段,现有技术中,通常采用聚类的方法来应对,然而采用这种方法存在多种缺陷,例如,一方面,如果后续对反欺诈模型添加数据,会对反欺诈模型更新数据造成困难,另一方面,经过聚类之后,虽然能将节点划分为若干类,但群体内的结构以及结构之间的关联仍然难以描述。
综上所述,现有技术中存在着如果后续对反欺诈模型添加数据,造成反欺诈模型更新数据困难;经过聚类之后,群体内的结构以及结构之间的关联仍然难以描述的问题,因此,需要采取有效的措施来解决以上问题。
发明内容
本发明实施例提供一种基于特征匹配网络的社团划分方法、装置及电子设备,用以解决现有技术中存在着如果后续对反欺诈模型添加数据,造成反欺诈模型更新数据困难、经过聚类之后,群体内的结构以及结构之间的关联仍然难以描述的问题。
本发明实施例提供一种基于特征匹配网络的社团划分方法,包括:
根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;
将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;
针对每个类,将子哈希向量相同的账号信息划分为同一组;
计算同一组内的各账号信息之间的相似度;
若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建立互连边,形成特征匹配网络;
根据所述特征匹配网络,对所述各账号信息进行社团划分。
可选的,计算同一组内的各账号信息之间的相似度,包括:
若第i账号信息与第j账号信息位于n类同组中,则将n/m作为所述第i帐号信息与所述第j账号信息之间的相似度;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个。
可选的,计算同一组内的各账号信息之间的相似度,包括:
若第i账号信息与第j账号信息位于同一组中,统计所述第i账号信息的哈希向量与所述第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个;
所述第i账号信息与所述第j账号信息的相似度s=h/K。
可选的,根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量,包括:
可选的,根据所述特征匹配网络,对所述各账号信息进行社团划分,包括:
(1)将各账号信息划分在所述特征匹配网络中不同的社区中;
(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;
(3)针对每个账号信息,从所述节点相似强度矩阵中所述账号信息所在的行,按相似强度从大到小的的顺序尝试将所述账号信息划至其他社区中;若所述账号信息自第p社区划分至第q社区后的模块度差为正数,则将所述账号信息划分至第q社区后结束;
(4)重复执行,直到社区结构不再改变为止。
可选的,所述根据各账号信息之间的相似度,计算每个账号信息的相似强度,包括:
根据公式(2)计算所述第i账号信息与所述第j账号信息之间的相似强度si,j;
其中,Γ(i)表示所述第i账号信息的邻居集合,Γ(i)∩Γ(j)表示所述第i账号信息与所述第j账号信息的共同邻居集合,wai,z为任意账号信息ai与第z账号信息之间的边的权重和。
本发明实施例还提供一种基于特征匹配网络的社团划分装置,包括:
确定单元,用于根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;
第一划分单元,用于将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;
第二划分单元,用于针对每个类,将子哈希向量相同的账号信息划分为同一组;
计算单元,用于计算同一组内的各账号信息之间的相似度;
形成网络单元,用于若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建立互连边,形成特征匹配网络;
第三划分单元,用于根据所述特征匹配网络,对所述各账号信息进行社团划分。
可选的,计算单元,具体用于若第i账号信息与第j账号信息位于n类同组中,则将n/m作为所述第i帐号信息与所述第j账号信息之间的相似度;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个。
可选的,计算单元,具体还用于若第i账号信息与第j账号信息位于同一组中,统计所述第i账号信息的哈希向量与所述第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个;
所述第i账号信息与所述第j账号信息的相似度s=h/K。
可选的,第三划分单元,具体用于(1)将各账号信息划分在所述特征匹配网络中不同的社区中;
(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;
(3)针对每个账号信息,从所述节点相似强度矩阵中所述账号信息所在的行,按相似强度从大到小的的顺序尝试将所述账号信息划至其他社区中;若所述账号信息自第p社区划分至第q社区后的模块度差为正数,则将所述账号信息划分至第q社区后结束;
(4)重复执行,直到社区结构不再改变为止。
可选的,计算单元,具体还用于根据公式(4)计算所述第i账号信息与所述第j账号信息之间的相似强度si,j;
其中,Γ(i)表示所述第i账号信息的邻居集合,Γ(i)∩Γ(j)表示所述第i账号信息与所述第j账号信息的共同邻居集合,wai,z为任意账号信息ai与第z账号信息之间的边的权重和。
本发明实施例还提供一种电子设备,包括:
至少一个处理器;以及,
与所述至少一个处理器通信连接的存储器;其中,
所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够:根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;针对每个类,将子哈希向量相同的账号信息划分为同一组;计算同一组内的各账号信息之间的相似度;若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建
立互连边,形成特征匹配网络;根据所述特征匹配网络,对所述各账号信息进行社团划分。
本发明实施例中提供了一种基于特征匹配网络的社团划分方法、装置及电子设备,根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;针对每个类,将子哈希向量相同的账号信息划分为同一组;计算同一组内的各账号信息之间的相似度;若各账号信息之间的相似度大于阈值,则在各账号信息之间建立互连边,形成特征匹配网络;根据特征匹配网络,对各账号信息进行社团划分。本发明实施例中首先通过根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量,对于网络中数量巨大的账号信息来说,仅仅产生两个哈希值的哈希函数是不够的,因此确定每个账号信息对应的K位哈希向量能够应对复杂的网络账号信息。然后针对每个类,将子哈希向量相同的账号信息划分为一组,计算同一组内任意账号信息之间的相似度,能够避免针对整个网络中任意账号信息之间计算相似度而带来的计算量非常大的缺点;本发明技术方案能够有效减少账号信息之间相似度的计算量,仅仅计算同一组内的账号信息之间的相似度。最后根据确定各账号信息之间的相似度大于阈值,在各账号信息之间建立互连边,形成特征匹配网络;根据特征匹配网络,对各账号信息进行社团划分,能够更精准的对各账号信息进行社团划分,这样不仅能够使社团之间的关联关系很清楚,而且能够对划分的社团进行分析,找出异常社团,进而对异常社团内的账号进行异常账号排查,更加有针对性地找出欺诈账号,提高应对欺诈账号的效率。此外,如果需要对划分出的社团添加账号信息,只需要对该添加的账号信息重复以上简单的几个步骤,将所添加的账号信息更新到相应的位置即可,并不会产生更新困难的问题。
为了更清楚地说明本发明实施例中的技术方案,下面将对实施例描述中所需要使用的附图作简要介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域的普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本发明实施例提供了一种基于特征匹配网络的社团划分方法流程示意图;
图2为本发明实施例提供了本发明的整体思路流程图;
图3为本发明实施例提供的一种基于特征匹配网络的社团划分装置结构示意图;
图4为本发明实施例提供的电子设备的结构示意图。
为了使本发明的目的、技术方案及有益效果更加清楚明白,以下结合附图及实施例,对本发明进行进一步详细说明。应当理解,此处所描述的具体实施例仅仅用以解释本发明,并不用于限定本发明。
应理解,本发明实施例的技术方案可以应用于各种银行出现的网络欺诈手段的场景,比如可以是信用卡产品的欺诈、银行卡产品的欺诈、盗卡欺诈、伪卡欺诈、套现欺诈等等。本发明实施例的技术方案的应用场景也可以是对异常账号信息社团的发现、发现特定种类欺诈的共性、根据欺诈账号信息样本发现其它欺诈账号信息、帮助发现未知欺诈类型等。
图1示例性示出了本发明实施例提供的一种基于特征匹配网络的社团划分方法流程示意图,如图1所示,包括以下步骤:
步骤S101:根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;
步骤S102:将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;
步骤S103:针对每个类,将子哈希向量相同的账号信息划分为同一组;
步骤S104:计算同一组内的各账号信息之间的相似度;
步骤S105:若各账号信息之间的相似度大于阈值,则在各账号信息之间建立互连边,形成特征匹配网络;
步骤S106:根据特征匹配网络,对各账号信息进行社团划分。
步骤S101中,根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量,具体来说,经过每个预设的哈希函数的处理都能得到一位哈希向量,那么,根据预设的K个哈希函数,就可以产生K位哈希向量,而每个账号信息对应K位哈希向量,具体实施中,每个账号信息是包含多个特征属性的,如果仅仅使用现有技术中一个账号信息只用一个哈希函数来表示的话,会存在不足以表达一个账号信息的多个特征属性的缺点,所以,本步骤可以有效避免这个缺点。其中,K的取值可以根据具体实施中各账号信息的具体情况来设定,比如,K可以设定为4,那么账号信息就可以表示为一个4位的哈希向量。
步骤S102:将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量,具体来说,比如,K=4,k=2,那么,就将每个账号信息为4位的哈希向量划分为2类子哈希向量,划分的好处是为后续计算账号间的相似度减少计算量,避免出现像现有技术中并没有对账号信息的哈希向量进行划分而出现直接对所有账号信息中的任意两个账号来进行相似度计算而造成的计算量特别大的缺点。
步骤S103:针对每个类,将子哈希向量相同的账号信息划分为同一组,具体来说,对
每个账号信息划分为各类之后,针对划分的每个类,将子哈希向量相同的账号信息划分为同一组,比如,K=4,k=2的话,在第1类中,所有账号信息中4位哈希向量中前两位相同的为一组,同样,在第2类中,所有账号信息中4位哈希向量中后两位相同的账号信息为一组。这样划分的目的也是为了后面减少计算相似度的计算量,只计算各类之间子哈希向量相同的账号信息之间的相似度。
步骤S104:计算同一组内的各账号信息之间的相似度,具体实施中,可以统计同一组内各账号信息的哈希向量的位相同的个数与位的大小的比值,比如,账号信息1的哈希向量为0010,账号信息2的哈希向量为0011,按照K=4,k=2,那么两个账号信息在第一类中位于同一组,则确定位于同一组的两个账号信息的相似度;那么,两个账号信息的哈希向量的位相同的个数是3,位的大小是4位的,所以,这两个账号信息之间的相似度为3/4,也可以根据关于相似度的计算公式来计算同一组内的任意两个账号信息之间的相似度,比如相似度的计算公式可以是欧式距离、余弦距离、杰卡德距离公式等。一方面,相比于计算所有账号信息中的任意两个账号信息的相似度,只计算同一组内的任意两个账号信息之间的相似度能够大大减少计算量。比如,取N个账号信息样本,那么N个账号信息样本就被分到了2k个组内,每个组内的账号信息样本数为N/2k,每组内进行任意两个账号信息进行相似度计算的次数为2k个组进行任意两个账号信息进行相似度计算的次数为因此,所有类需要进行相似度计算的次数就为其中,是划分的类的个数,这个值是一个根据实际情况可以进行控制的常数,而传统的方法计算所有账号中任意两个账号信息进行相似度计算需要进行次,综上可以看出,采用本发明的计算同一组内的任意两个账号信息之间的相似度的计算量比传统的方法计算所有账号中任意两个账号信息的相似度的计算量大约缩减2k级别的倍数。另一方面,每一组内的账号信息的相似度是较大的,所以对同一组内的账号信息进行相似度计算,也能够提高网络建立的效率和准确率。
步骤S105:若各账号信息之间的相似度大于阈值,则在各账号信息之间建立互连边,形成特征匹配网络,具体来说,如果任意两个账号信息之间的相似度大于阈值,就在任意两个账号信息之间建立一条互连边,边的权重就是两个账号信息之间的相似度值,最终形成特征匹配网络。具体实施中,阈值的选取可以选择较高的值没这样最终可以生成较为稀
疏的特征匹配网络,便于后续的计算,另外,阈值的取值可以根据实际情况进行调整。
步骤S106:根据特征匹配网络,对各账号信息进行社团划分,具体来说,根据计算出来的各账号信息之间的相似度值,相似度值越接近的越容易被划分到同一个社团中。划分社团之后,对于网络中的欺诈账号更容易去排查,可以计算欺诈账号样本在每个社团中的比例,比例较大的,则该社团为异常社团的可能性就越大,可以根据业务需要进行相关调查,再对异常社团内的账号根据一些指标来进行计算,找出具有代表性的账号,对这些具有代表性的账号再进行相关案件排查,其中,一些指标可以是社团内账号信息的度中心性、紧密中心性、特征向量中心性等;或者也可以对社团内的账号信息进行特征再分析,以期发现该社团的一些共同行为的特征,进行有针对性地欺诈预防。此外,如果新加入的账号信息形成新的社团,则可以根据前面查出来的异常社团进行比对,这对于未知欺诈的侦测与预防是大有裨益的。
计算同一组内的各账号信息之间的相似度,可以以下面两种方法来计算:
方式1:可选地,计算同一组内的各账号信息之间的相似度,包括:若第i账号信息与第j账号信息位于n类同组中,则将n/m作为第i帐号信息与第j账号信息之间的相似度;第i账号信息与第j账号信息为各账号信息中的任一个,具体来说,在所有账号信息中任意取两个账号信息,比如称为账号信息1与账号信息2,m取3,也就是账号信息1与账号信息2分在了3类中,这3类分别称为第1类、第2类、第3类,假设这两个账号信息在第1类与第3类中同组,那么,这两个账号信息在这3类中的相似度为2/3。
方式2:可选地,计算同一组内的各账号信息之间的相似度,包括:若第i账号信息与第j账号信息位于同一组中,统计第i账号信息的哈希向量与第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;第i账号信息与第j账号信息为各账号信息中的任一个;第i账号信息与第j账号信息的相似度s=h/K,具体来说,如果所有账号信息中任意的两个账号信息,账号信息1与账号信息2位于同一组,并且账号信息1与账号信息2都是4位的,也就是K为4,账号信息1与账号信息2的4位哈希向量中,前3位是完全相同的,第4位不同,那么,账号信息1与账号信息2的相似度s为3/4。
以上两种计算同一组内各个账号信息之间的相似度的计算方法,可以得出,第1中方法是计算的两个账号信息在各个类中的相似度,而第2种方法是计算的被分到了各类中同一组中的两个账号信息之间的相似度,可以看出,这两种方法中,相比于第2种方法,第1种方法是比较粗略的计算两个账号信息所属的类与类之间的相似度,而第2种计算的两个账号信息在同一组之间的相似度则更精准。不过,这两种方法都相比于现有技术中利用欧式距离公式等来计算网络中所有账号信息中任意两个账号信息之间的相似度的计算量
上得到了明显的改善,进一步加速了网络的建立。
表示账号信息的特征向量,其中,c1,c2…,cd表示账号信息的特征属性,表示随机选取的一个非零向量,具体来说,预设的哈希函数是是预设的K个哈希函数中的任一个,哈希函数的值用0或1来表示,也就是说这样的一个哈希函数只能产生两个哈希值,对于数量巨大的账号信息来说明显是不够的,所以根据这样的哈希函数,来确定每个账号的K位哈希向量是一个K位的二进制数,比如,可以是6位的二进制数,具体可以为010110,那么,
其中,表示账号信息的特征向量,c1,c2…,cd表示账号信息的特征属性,具体的账号信息特征属性可以是交易金额、交易时间、交易地点、交易地点数、转账地点、转账金额、转账次数等。其中,各账号信息的特征向量在具体实施中可以经过筛选来得到一批理论上效果最好的特征向量,具体地,在一定时间段内抽取欺诈账号信息样本以及正常账号信息样本,将抽取的欺诈账号信息样本以及正常账号信息样本组合为一个整体账号信息样本,根据业务经验进行整体账号信息的数据预处理、特征筛选及属性相关性分析等步骤之后,筛选出一批理论上效果最好的特征向量。根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量,能够充分提取每个账号信息的特征属性并用特征向量表示出来,能够应对复杂的网络中账号信息数量巨大的情况。此外,需要说明的是,第一,每个账号信息对应的K位哈希向量的确定实际上是经过一个哈希随机映射的过程得来的,是由经过哈希映射得到这里使用哈希随机映射的主要目的是使得使得账号信息的特征向量能映射为0或1的统一表示,以便后续处理,而并非简单的降维;第二,原来的特征向量映射到新的哈希空间中,会使得在原来的特征向量相似的数据在新的哈希空间中数据也相似的概率很大,这个概率为:符合相似度s到概率p
的单调递增映射关系。
以上实施方式中,对于每个账号信息对应的K位哈希向量以及将每个账号信息对应的哈希向量顺序划分为m=K/k类子哈希向量的关系,下面以一个表格的方式将其展示出来,表1示例性地示出了账号信息样本与类之间的关系,如表1所示:
表1:账号信息样本与类之间的关系
表1中,账号信息样本与类之间的关系可以表示成一个K行N列的矩阵,N表示取的账号信息样本数,c1到cN代表N个账号信息样本,将N个账号信息样本分到m=K/k个类,其中,表格中除第一行之外下面的每一行代表一个类,N个账号信息样本被分到了2k个组内。
可选地,根据特征匹配网络,对各账号信息进行社团划分,包括:
(1)将各账号信息划分在特征匹配网络中不同的社区中;
(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;
(3)针对每个账号信息,从节点相似强度矩阵中账号信息所在的行,按相似强度从
大到小的的顺序尝试将账号信息划至其他社区中;若账号信息自第p社区划分至第q社区后的模块度差为正数,则将账号信息划分至第q社区后结束;
(4)重复执行,直到社区结构不再改变为止。
可选地,根据各账号之间的相似度,计算每个账号信息的相似强度,包括:
根据公式(2)计算第i账号信息与第j账号信息之间的相似强度si,j;
其中,Γ(i)表示第i账号信息的邻居集合,Γ(i)∩Γ(j)表示第i账号信息与第j账号信息的共同邻居集合,wai,z为任意账号信息ai与第z账号信息之间的边的权重和。
具体实施中,第(1)步骤,初始化特征匹配网络,将每个账号信息划分到不同的社区中,这一步骤中的划分可以是随机划分的;第(2)步骤,根据公式(2)来计算各账号信息的相似强度,具体地,假如账号信息1与账号信息2的共同邻居是账号信息3,账号信息1与账号信息2合起来与账号信息3的边的权重是5,那么,任意账号信息ai与账号信息3相连边的权重为5,因而,账号信息1与账号信息2的相似强度是1/5,类似的,其它账号信息之间也是用此方法来计算。假如,取4个账号信息样本,经过计算之后,形成一个4*4的矩阵,假如,这个矩阵为从这个矩阵可以看出,账号信息1与账号信息2的相似度为0.25,账号信息1与账号信息3的相似度为0.7,账号信息2与账号信息3的相似度为0.4;第(3)步骤,从这个相似强度矩阵中账号信息所在的行,按相似强度从大到小的的顺序尝试将账号信息划至其他社区中,例如从这个相似矩阵第一行可以看出,想要把账号信息1划分到其它某一社团中时,优先选择相似度较大的账号信息3(第一行中0.6最大)所在的社区中去。如果ΔQ<0,再将账号信息1尝试划分到账号信息4(第一行中0.4次大)所在的社团中去。如果ΔQ<0,则再将账号信息1尝试划分到账号信息2所在的社团中去。如果仍然ΔQ<0,则账号信息1作为一个独立的社团进行保留,矩阵不做更新,再进行第2行的计算。如果上述尝试过程中只要发现ΔQ>0,比如优先尝试的将账号1划分到相似度较大的账号信息3(第一行中0.6最大)所在的社区中去以后,发现ΔQ>0,那么表示尝试成功,第一行计算结束。由于此时账号1的状态已经发生改变,因此将矩阵中第一行第一列所有数据删除,表示后续账号信息不再与账号
信息1进行比较,也就是,变成然后以同样的过程开始新一轮的尝试计算,即对账号信息2进行社团划分。其中,模块度差ΔQ的计算公式:来验证上面对账号信息的尝试划分社区是否正确,其中,n表示网络中所有的权重,ki表示与顶点i连接的边的权重,ki,in表示账号信息i在社区内部的权重之和,Σin表示社区内部的边权重和,Σtot表示与社区内部账号信息连接的边的权重和,包括社区内部的边以及社区外部的边,若ΔQ为正数,则接受本次的划分,若不为正数,则放弃本次的划分。通过账号信息的相似强度矩阵的计算,优先将账号信息划分到与其最相似的邻居账号信息的社团中去,大大节省了社团划分的尝试次数,进一步提高了算法的速度,另外,对账号信息尝试的划分是否合理通过模块度差公式来验证,更加有效保证了尝试划分的合理性与准确性。
为了更好的理解本发明技术方案,图2示例性地示出了本发明的整体思路流程图,如图2所示:
步骤S201:将各账号信息的特征属性通过哈希映射的方法映射为一个多位的哈希映射向量;
步骤S202:将各账号信息的哈希映射向量进行分类;
步骤S203:对于每个类,将哈希映射向量相同的账号信息划分为一组;
步骤S204:对每组中的任意两个账号信息进行相似度计算;
步骤S205:若每组中的任意两个账号信息的相似度大于阈值,则建立这两个账号信息之间的互连边,边的权重为相似度,从而形成特征匹配网络,其中,形成的特征匹配网络是稀疏的特征匹配网络;
步骤S206:根据特征匹配网络中各账号信息的相似强度矩阵对特征匹配网络进行社团划分。
与现有技术相比,本发明实施例中,第一,通过随机哈希映射的方法将各账号信息的特征属性映射到一个新的哈希空间中,形成各账号信息的哈希映射向量,对各账号信息的哈希映射向量进行分类,能够在高相似度的账号信息之间建立边,有效避免了大量的任意两个账号信息之间的相似度计算,且高效地为每条边建立了可信的权重值,能够提高后续社团划分的精度与速度;第二,根据各账号信息的相似度建立了特征匹配网络,然后根据
网络中各账号信息的相似强度矩阵对特征匹配网络进行社团划分,不仅可以有效发现异常社团并进行有针对性地措施,同时可以侦测未知的欺诈类型,而且通过相似强度矩阵对对特征匹配网络进行社团划分,即优先将账号信息划分到与其最相似的邻居账号信息的社团中去,大大节省了社团划分尝试的次数,进一步提高了算法的速度;第三,通过形成特征匹配网络,相关账号信息间的相似度作为边的权重被永久存储,即使有较多的新的账号信息进来,也不会对网络中原来的互连边产生影响,仅仅需要将新的账号信息插入到原特征匹配网络中。在向原特征匹配网络图添加新数账号信息的时候,仍然先采用随机哈希映射方法及对各账号信息进行分类,然后与类内的账号信息进行相似度计算,如果该相似度大于阈值,则添加新的边。后续只需要进行计算量较小但是更加精准的社团划分算法即可实现功能。同时,特征匹配网络的结构能更加清晰地展示社团内部及社团间的关联结构,这是传统聚类方法所不能实现的。
基于相同构思,本发明实施例提供的一种基于特征匹配网络的社团划分装置,如图3所示,该装置包括确定单元301、第一划分单元302、第二划分单元303、计算单元304、形成网络单元305和第三划分单元306。其中:
确定单元301:用于根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;
第一划分单元302:用于将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;
第二划分单元303:用于针对每个类,将子哈希向量相同的账号信息划分为同一组;
计算单元304:用于计算同一组内的各账号信息之间的相似度;
形成网络单元305:用于若各账号信息之间的相似度大于阈值,则在各账号信息之间建立互连边,形成特征匹配网络;
第三划分单元306:用于根据特征匹配网络,对各账号信息进行社团划分。
可选地,计算单元304具体用于:
若第i账号信息与第j账号信息位于n类同组中,则将n/m作为第i帐号信息与第j账号信息之间的相似度;第i账号信息与第j账号信息为各账号信息中的任一个。
可选地,计算单元304具体还用于:
若第i账号信息与第j账号信息位于同一组中,统计第i账号信息的哈希向量与第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;第i账号信息与第j账号信息为各账号信息中的任一个;
第i账号信息与第j账号信息的相似度s=h/K。
可选地,确定单元301用于:
可选地,第三划分单元306具体用于:
(1)将各账号信息划分在特征匹配网络中不同的社区中;
(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;
(3)针对每个账号信息,从节点相似强度矩阵中账号信息所在的行,按相似强度从大到小的的顺序尝试将账号信息划至其他社区中;若账号信息自第p社区划分至第q社区后的模块度差为正数,则将账号信息划分至第q社区后结束;
(4)重复执行,直到社区结构不再改变为止。
可选地,计算单元304具体还用于:
根据公式(4)计算第i账号信息与第j账号信息之间的相似强度si,j;
其中,Γ(i)表示第i账号信息的邻居集合,Γ(i)∩Γ(j)表示第i账号信息与第j账号信息的共同邻居集合,wai,z为任意账号信息ai与第z账号信息之间的边的权重和。
参见图4,为本发明实施例提供的一种电子设备,该电子设备可应用于本发明的上述实施例。
该电子设备可包括一个或多个处理器410以及存储器420,图4中以一个处理器410为例。
执行基于特征匹配网络的社团划分方法的设备还可以包括:输入装置430和输出装置440。
处理器410、存储器420、输入装置430和输出装置440可以通过总线或者其他方式
连接,图4中以通过总线连接为例。
存储器420作为一种非易失性计算机可读存储介质,可用于存储非易失性软件程序、非易失性计算机可执行程序以及模块,如本申请实施例中的基于特征匹配网络的社团划分方法对应的程序指令/模块。处理器410通过运行存储在存储器420中的非易失性软件程序、指令以及模块,从而执行服务器的各种功能应用以及数据处理,即实现上述方法实施例基于特征匹配网络的社团划分方法。
存储器420可以包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需要的应用程序;存储数据区可存储根据基于特征匹配网络的社团划分装置的使用所创建的数据等。此外,存储器420可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他非易失性固态存储器件。在一些实施例中,存储器420可选包括相对于处理器410远程设置的存储器,这些远程存储器可以通过网络连接至基于特征匹配网络的社团划分装置。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
输入装置430可接收输入的数字或字符信息,以及产生与基于特征匹配网络的社团划分装置的用户设置以及功能控制有关的键信号输入。输出装置440可包括显示屏等显示设备。
所述一个或者多个模块存储在所述存储器420中,当被所述一个或者多个处理器410执行时,执行上述任意方法实施例中的基于特征匹配网络的社团划分方法。
处理器410,被配置了一个或多个可执行程序,所述一个或多个可执行程序用于执行以下过程:根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;针对每个类,将子哈希向量相同的账号信息划分为同一组;计算同一组内的各账号信息之间的相似度;若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建立互连边,形成特征匹配网络;根据所述特征匹配网络,对所述各账号信息进行社团划分。
较佳地,处理器410具体用于:
若第i账号信息与第j账号信息位于n类同组中,则将n/m作为所述第i帐号信息与所述第j账号信息之间的相似度;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个。
较佳地,处理器410具体用于:
若第i账号信息与第j账号信息位于同一组中,统计所述第i账号信息的哈希向量与所述第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;所述第i账号信息与
所述第j账号信息为所述各账号信息中的任一个;
所述第i账号信息与所述第j账号信息的相似度s=h/K。
较佳地,处理器410具体用于:
较佳地,处理器410具体用于:
(1)将各账号信息划分在所述特征匹配网络中不同的社区中;
(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;
(3)针对每个账号信息,从所述节点相似强度矩阵中所述账号信息所在的行,按相似强度从大到小的的顺序尝试将所述账号信息划至其他社区中;若所述账号信息自第p社区划分至第q社区后的模块度差为正数,则将所述账号信息划分至第q社区后结束;
(4)重复执行,直到社区结构不再改变为止。
较佳地,处理器410具体用于:
根据公式(2)计算所述第i账号信息与所述第j账号信息之间的相似强度si,j;
其中,Γ(i)表示所述第i账号信息的邻居集合,Γ(i)∩Γ(j)表示所述第i账号信息与所述第j账号信息的共同邻居集合,wai,z为任意账号信息ai与第z账号信息之间的边的权重和。
从上述内容可看出:本发明实施例中提供一种基于特征匹配网络的社团划分装置,根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;将每个账号信息对应的哈希向量,顺序划分为类子哈希向量;针对每个类,将子哈希向量相同的账号信息划分为同一组;计算同一组内的各账号信息之间的相似度;若各账号信息之间的相似度大于阈
值,则在各账号信息之间建立互连边,形成特征匹配网络;根据特征匹配网络,对各账号信息进行社团划分根据各账号信息之间的相似度,对各账号信息进行社团划分。本发明实施例中首先通过根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量,对于网络中数量巨大的账号信息来说,仅仅产生两个哈希值的哈希函数是不够的,因此确定每个账号信息对应的K位哈希向量能够应对复杂的网络账号信息。然后针对每个类,将子哈希向量相同的账号信息划分为一组,计算同一组内任意账号信息之间的相似度,能够避免针对整个网络中任意账号信息之间计算相似度而带来的计算量非常大的缺点;本发明技术方案能够有效减少账号信息之间相似度的计算量,仅仅计算同一组内的账号信息之间的相似度。最后根据确定各账号信息之间的相似度大于阈值,在各账号信息之间建立互连边,形成特征匹配网络;根据特征匹配网络,对各账号信息进行社团划分,能够更精准的对各账号信息进行社团划分,这样不仅能够使社团之间的关联关系很清楚,而且能够对划分的社团进行分析,找出异常社团,进而对异常社团内的账号进行异常账号排查,更加有针对性地找出欺诈账号,提高应对欺诈账号的效率。此外,如果需要对划分出的社团添加账号信息,只需要对该添加的账号信息重复以上简单的几个步骤,将所添加的账号信息更新到相应的位置即可,并不会产生更新困难的问题。
本领域内的技术人员应明白,本发明的实施例可提供为方法、或计算机程序产品。因此,本发明可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本发明可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
本发明是参照根据本发明实施例的方法、设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。
这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。
这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。
尽管已描述了本发明的优选实施例,但本领域内的技术人员一旦得知了基本创造性概念,则可对这些实施例作出另外的变更和修改。所以,所附权利要求意欲解释为包括优选实施例以及落入本发明范围的所有变更和修改。
显然,本领域的技术人员可以对本发明进行各种改动和变型而不脱离本发明的精神和范围。这样,倘若本发明的这些修改和变型属于本发明权利要求及其等同技术的范围之内,则本发明也意图包含这些改动和变型在内。
Claims (15)
- 一种基于特征匹配网络的社团划分方法,其特征在于,包括:根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;针对每个类,将子哈希向量相同的账号信息划分为同一组;计算同一组内的各账号信息之间的相似度;若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建立互连边,形成特征匹配网络;根据所述特征匹配网络,对所述各账号信息进行社团划分。
- 如权利要求1所述的方法,其特征在于,计算同一组内的各账号信息之间的相似度,包括:若第i账号信息与第j账号信息位于n类同组中,则将n/m作为所述第i帐号信息与所述第j账号信息之间的相似度;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个。
- 如权利要求1所述的方法,其特征在于,计算同一组内的各账号信息之间的相似度,包括:若第i账号信息与第j账号信息位于同一组中,统计所述第i账号信息的哈希向量与所述第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个;所述第i账号信息与所述第j账号信息的相似度s=h/K。
- 如权利要求1至4任一项所述的方法,其特征在于,根据所述特征匹配网络,对所述各账号信息进行社团划分,包括:(1)将各账号信息划分在所述特征匹配网络中不同的社区中;(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;(3)针对每个账号信息,从所述节点相似强度矩阵中所述账号信息所在的行,按相似强度从大到小的的顺序尝试将所述账号信息划至其他社区中;若所述账号信息自第p社区划分至第q社区后的模块度差为正数,则将所述账号信息划分至第q社区后结束;(4)重复执行,直到社区结构不再改变为止。
- 一种基于特征匹配网络的社团划分装置,其特征在于,包括:确定单元,用于根据预设的K个哈希函数,确定每个账号信息对应的K位哈希向量;第一划分单元,用于将每个账号信息对应的哈希向量,顺序划分为m=K/k类子哈希向量;第二划分单元,用于针对每个类,将子哈希向量相同的账号信息划分为同一组;计算单元,用于计算同一组内的各账号信息之间的相似度;形成网络单元,用于若所述各账号信息之间的相似度大于阈值,则在所述各账号信息之间建立互连边,形成特征匹配网络;第三划分单元,用于根据所述特征匹配网络,对所述各账号信息进行社团划分。
- 如权利要求7所述的装置,其特征在于,所述计算单元,具体用于若第i账号信息与第j账号信息位于n类同组中,则将n/m作为所述第i帐号信息与所述第j账号信息之间的相似度;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个。
- 如权利要求7所述的装置,其特征在于,所述计算单元,具体还用于若第i账号信息与第j账号信息位于同一组中,统计所述第i账号信息的哈希向量与所述第j账号信息的哈希向量中位于同一位且哈希向量值相同的个数h;所述第i账号信息与所述第j账号信息为所述各账号信息中的任一个;所述第i账号信息与所述第j账号信息的相似度s=h/K。
- 如权利要求7至10任一项所述的装置,其特征在于,所述第三划分单元,具体用于(1)将各账号信息划分在所述特征匹配网络中不同的社区中;(2)根据各账号信息之间的相似度,计算每个账号信息的相似强度,从而生成节点相似强度矩阵;(3)针对每个账号信息,从所述节点相似强度矩阵中所述账号信息所在的行,按相似强度从大到小的的顺序尝试将所述账号信息划至其他社区中;若所述账号信息自第p社区划分至第q社区后的模块度差为正数,则将所述账号信息划分至第q社区后结束;(4)重复执行,直到社区结构不再改变为止。
- 一种电子设备,其特征在于,包括:至少一个处理器;以及,与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1至6任一项所述的方法。
- 一种非易失性计算机存储介质,所述非易失性计算机存储介质用于存储有计算机可执行指令,所述计算机可执行指令用于使所述计算机执行权利要求1至6任一项所述的方法。
- 一种计算机程序产品,所述计算机程序产品包括存储在非易失性计算机存储介质上计算机程序,所述计算机程序包括程序指令,当所述程序指令被计算机执行时,使所述计算机执行权利要求1至6任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201611110731.7 | 2016-12-06 | ||
| CN201611110731.7A CN106709800B (zh) | 2016-12-06 | 2016-12-06 | 一种基于特征匹配网络的社团划分方法和装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018103456A1 true WO2018103456A1 (zh) | 2018-06-14 |
Family
ID=58937536
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/105985 Ceased WO2018103456A1 (zh) | 2016-12-06 | 2017-10-13 | 一种基于特征匹配网络的社团划分方法、装置及电子设备 |
Country Status (3)
| Country | Link |
|---|---|
| CN (1) | CN106709800B (zh) |
| TW (1) | TWI662421B (zh) |
| WO (1) | WO2018103456A1 (zh) |
Cited By (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109598509A (zh) * | 2018-10-17 | 2019-04-09 | 阿里巴巴集团控股有限公司 | 风险团伙的识别方法和装置 |
| CN109816535A (zh) * | 2018-12-13 | 2019-05-28 | 中国平安财产保险股份有限公司 | 欺诈识别方法、装置、计算机设备及存储介质 |
| CN109859054A (zh) * | 2018-12-13 | 2019-06-07 | 平安科技(深圳)有限公司 | 网络社团挖掘方法、装置、计算机设备及存储介质 |
| CN110046929A (zh) * | 2019-03-12 | 2019-07-23 | 平安科技(深圳)有限公司 | 一种欺诈团伙识别方法、装置、可读存储介质及终端设备 |
| CN110297948A (zh) * | 2019-05-28 | 2019-10-01 | 阿里巴巴集团控股有限公司 | 关系网络构建方法以及装置 |
| CN111292171A (zh) * | 2020-02-28 | 2020-06-16 | 中国工商银行股份有限公司 | 金融理财产品推送方法及装置 |
| CN111343012A (zh) * | 2020-02-17 | 2020-06-26 | 平安科技(深圳)有限公司 | 云平台的缓存服务器部署方法、装置和计算机设备 |
| CN111666501A (zh) * | 2020-06-30 | 2020-09-15 | 腾讯科技(深圳)有限公司 | 异常社团识别方法、装置、计算机设备和存储介质 |
| CN111681090A (zh) * | 2020-05-28 | 2020-09-18 | 平安银行股份有限公司 | 业务系统的账号分组方法、装置、终端设备及存储介质 |
| CN112926991A (zh) * | 2021-03-30 | 2021-06-08 | 顶象科技有限公司 | 一种套现团伙严重等级划分方法及系统 |
| CN114881783A (zh) * | 2022-05-16 | 2022-08-09 | 中国银联股份有限公司 | 一种异常卡识别方法、装置、电子设备及存储介质 |
| CN115001971A (zh) * | 2022-04-14 | 2022-09-02 | 西安交通大学 | 天地一体化信息网络下改进社团发现的虚拟网络映射方法 |
| CN116450886A (zh) * | 2023-02-24 | 2023-07-18 | 支付宝(杭州)信息技术有限公司 | 边数据增加方法及装置、介质、设备 |
Families Citing this family (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106709800B (zh) * | 2016-12-06 | 2020-08-11 | 中国银联股份有限公司 | 一种基于特征匹配网络的社团划分方法和装置 |
| CN107194623B (zh) * | 2017-07-20 | 2021-01-05 | 深圳市分期乐网络科技有限公司 | 一种团伙欺诈的发现方法及装置 |
| CN107871277B (zh) * | 2017-07-25 | 2021-04-13 | 平安普惠企业管理有限公司 | 服务器、客户关系挖掘的方法及计算机可读存储介质 |
| CN110019193B (zh) * | 2017-09-25 | 2022-10-14 | 腾讯科技(深圳)有限公司 | 相似帐号识别方法、装置、设备、系统及可读介质 |
| CN108295476B (zh) * | 2018-03-06 | 2021-12-28 | 网易(杭州)网络有限公司 | 确定异常交互账户的方法和装置 |
| CN110227268B (zh) * | 2018-03-06 | 2022-06-07 | 腾讯科技(深圳)有限公司 | 一种检测违规游戏帐号的方法及装置 |
| CN108829769B (zh) * | 2018-05-29 | 2021-08-06 | 创新先进技术有限公司 | 一种可疑群组发现方法和装置 |
| CN109191107A (zh) * | 2018-06-29 | 2019-01-11 | 阿里巴巴集团控股有限公司 | 交易异常识别方法、装置以及设备 |
| CN109559218A (zh) * | 2018-11-07 | 2019-04-02 | 北京先进数通信息技术股份公司 | 一种异常交易的确定方法、装置及存储介质 |
| CN110688540B (zh) * | 2019-10-08 | 2022-06-10 | 腾讯科技(深圳)有限公司 | 一种作弊账户筛选方法、装置、设备及介质 |
| CN113034296B (zh) * | 2019-12-24 | 2023-09-22 | 腾讯科技(深圳)有限公司 | 用户账号的选择方法、装置、计算机设备及存储介质 |
| CN111444454B (zh) * | 2020-03-24 | 2023-05-05 | 哈尔滨工程大学 | 一种基于谱方法的动态社团划分方法 |
| CN111552842A (zh) * | 2020-03-30 | 2020-08-18 | 贝壳技术有限公司 | 一种数据处理的方法、装置和存储介质 |
| CN111784528B (zh) * | 2020-05-27 | 2024-07-02 | 平安科技(深圳)有限公司 | 异常社群检测方法、装置、计算机设备及存储介质 |
| CN112149000B (zh) * | 2020-09-09 | 2021-12-17 | 浙江工业大学 | 一种基于网络嵌入的在线社交网络用户社区发现方法 |
| CN113761080B (zh) * | 2021-04-01 | 2024-07-19 | 京东城市(北京)数字科技有限公司 | 社区划分方法、装置、设备及存储介质 |
| CN113326178A (zh) * | 2021-06-22 | 2021-08-31 | 北京奇艺世纪科技有限公司 | 一种异常账号传播方法、装置、电子设备和存储介质 |
| CN115641119A (zh) * | 2021-07-20 | 2023-01-24 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、设备以及计算机可读存储介质 |
| CN114781517B (zh) * | 2022-04-22 | 2025-07-18 | 京东科技控股股份有限公司 | 风险识别的方法、装置及终端设备 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090281971A1 (en) * | 2008-05-09 | 2009-11-12 | International Business Machines Corporation | System and method for classifying data streams with very large cardinality |
| EP2611101A1 (en) * | 2011-12-29 | 2013-07-03 | Verisign, Inc. | Systems and methods for detecting similarities in network traffic |
| CN106095813A (zh) * | 2016-05-31 | 2016-11-09 | 北京奇艺世纪科技有限公司 | 一种用户标识识别方法和装置 |
| CN106709800A (zh) * | 2016-12-06 | 2017-05-24 | 中国银联股份有限公司 | 一种基于特征匹配网络的社团划分方法和装置 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6996273B2 (en) * | 2001-04-24 | 2006-02-07 | Microsoft Corporation | Robust recognizer of perceptually similar content |
| WO2007002820A2 (en) * | 2005-06-28 | 2007-01-04 | Yahoo! Inc. | Search engine with augmented relevance ranking by community participation |
| WO2013094361A1 (ja) * | 2011-12-19 | 2013-06-27 | インターナショナル・ビジネス・マシーンズ・コーポレーション | ソーシャル・メデイアにおけるコミュニティを検出する方法、コンピュータ・プログラム、コンピュータ |
| EP2742439A1 (en) * | 2012-06-01 | 2014-06-18 | Qatar Foundation | A method for processing a large-scale data set, and associated apparatus |
| US20150120583A1 (en) * | 2013-10-25 | 2015-04-30 | The Mitre Corporation | Process and mechanism for identifying large scale misuse of social media networks |
-
2016
- 2016-12-06 CN CN201611110731.7A patent/CN106709800B/zh active Active
-
2017
- 2017-10-13 WO PCT/CN2017/105985 patent/WO2018103456A1/zh not_active Ceased
- 2017-12-06 TW TW106142677A patent/TWI662421B/zh active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090281971A1 (en) * | 2008-05-09 | 2009-11-12 | International Business Machines Corporation | System and method for classifying data streams with very large cardinality |
| EP2611101A1 (en) * | 2011-12-29 | 2013-07-03 | Verisign, Inc. | Systems and methods for detecting similarities in network traffic |
| CN106095813A (zh) * | 2016-05-31 | 2016-11-09 | 北京奇艺世纪科技有限公司 | 一种用户标识识别方法和装置 |
| CN106709800A (zh) * | 2016-12-06 | 2017-05-24 | 中国银联股份有限公司 | 一种基于特征匹配网络的社团划分方法和装置 |
Cited By (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109598509A (zh) * | 2018-10-17 | 2019-04-09 | 阿里巴巴集团控股有限公司 | 风险团伙的识别方法和装置 |
| CN109598509B (zh) * | 2018-10-17 | 2023-09-01 | 创新先进技术有限公司 | 风险团伙的识别方法和装置 |
| CN109816535A (zh) * | 2018-12-13 | 2019-05-28 | 中国平安财产保险股份有限公司 | 欺诈识别方法、装置、计算机设备及存储介质 |
| CN109859054A (zh) * | 2018-12-13 | 2019-06-07 | 平安科技(深圳)有限公司 | 网络社团挖掘方法、装置、计算机设备及存储介质 |
| CN109859054B (zh) * | 2018-12-13 | 2024-03-05 | 平安科技(深圳)有限公司 | 网络社团挖掘方法、装置、计算机设备及存储介质 |
| CN110046929B (zh) * | 2019-03-12 | 2023-06-20 | 平安科技(深圳)有限公司 | 一种欺诈团伙识别方法、装置、可读存储介质及终端设备 |
| CN110046929A (zh) * | 2019-03-12 | 2019-07-23 | 平安科技(深圳)有限公司 | 一种欺诈团伙识别方法、装置、可读存储介质及终端设备 |
| CN110297948A (zh) * | 2019-05-28 | 2019-10-01 | 阿里巴巴集团控股有限公司 | 关系网络构建方法以及装置 |
| CN110297948B (zh) * | 2019-05-28 | 2023-08-22 | 创新先进技术有限公司 | 关系网络构建方法以及装置 |
| CN111343012A (zh) * | 2020-02-17 | 2020-06-26 | 平安科技(深圳)有限公司 | 云平台的缓存服务器部署方法、装置和计算机设备 |
| CN111292171A (zh) * | 2020-02-28 | 2020-06-16 | 中国工商银行股份有限公司 | 金融理财产品推送方法及装置 |
| CN111292171B (zh) * | 2020-02-28 | 2023-06-27 | 中国工商银行股份有限公司 | 金融理财产品推送方法及装置 |
| CN111681090A (zh) * | 2020-05-28 | 2020-09-18 | 平安银行股份有限公司 | 业务系统的账号分组方法、装置、终端设备及存储介质 |
| CN111666501A (zh) * | 2020-06-30 | 2020-09-15 | 腾讯科技(深圳)有限公司 | 异常社团识别方法、装置、计算机设备和存储介质 |
| CN111666501B (zh) * | 2020-06-30 | 2024-04-12 | 腾讯科技(深圳)有限公司 | 异常社团识别方法、装置、计算机设备和存储介质 |
| CN112926991A (zh) * | 2021-03-30 | 2021-06-08 | 顶象科技有限公司 | 一种套现团伙严重等级划分方法及系统 |
| CN112926991B (zh) * | 2021-03-30 | 2024-04-30 | 中国银联股份有限公司 | 一种套现团伙严重等级划分方法及系统 |
| CN115001971A (zh) * | 2022-04-14 | 2022-09-02 | 西安交通大学 | 天地一体化信息网络下改进社团发现的虚拟网络映射方法 |
| CN114881783A (zh) * | 2022-05-16 | 2022-08-09 | 中国银联股份有限公司 | 一种异常卡识别方法、装置、电子设备及存储介质 |
| CN116450886A (zh) * | 2023-02-24 | 2023-07-18 | 支付宝(杭州)信息技术有限公司 | 边数据增加方法及装置、介质、设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106709800B (zh) | 2020-08-11 |
| TWI662421B (zh) | 2019-06-11 |
| TW201822022A (zh) | 2018-06-16 |
| CN106709800A (zh) | 2017-05-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2018103456A1 (zh) | 一种基于特征匹配网络的社团划分方法、装置及电子设备 | |
| US20230281629A1 (en) | Utilizing a check-return prediction machine-learning model to intelligently generate check-return predictions for network transactions | |
| US11594053B2 (en) | Deep-learning-based identification card authenticity verification apparatus and method | |
| US11403643B2 (en) | Utilizing a time-dependent graph convolutional neural network for fraudulent transaction identification | |
| US20210081798A1 (en) | Neural network method and apparatus | |
| CN113657896A (zh) | 一种基于图神经网络的区块链交易拓扑图分析方法和装置 | |
| CN108734380A (zh) | 风险账户判定方法、装置及计算设备 | |
| CN112750038B (zh) | 交易风险的确定方法、装置和服务器 | |
| CN112966728B (zh) | 一种交易监测的方法及装置 | |
| CN114140246A (zh) | 模型训练方法、欺诈交易识别方法、装置和计算机设备 | |
| Stephe et al. | Blockchain-based private AI model with RPOA based sampling method for credit card fraud detection | |
| CN114358101B (zh) | 基于交易对手匹配的虚拟货币交易所名称识别方法和装置 | |
| Xiao et al. | Explainable fraud detection for few labeled time series data | |
| CN116307671A (zh) | 风险预警方法、装置、计算机设备、存储介质 | |
| CN113344581B (zh) | 业务数据处理方法及装置 | |
| CN114926162A (zh) | 银行转账交易的风险控制方法及装置 | |
| CN115115372B (zh) | 对象的风险评估方法、装置、电子设备及可读存储介质 | |
| CN106779723A (zh) | 一种移动终端风险评估方法及装置 | |
| Jose et al. | Detection of credit card fraud using resampling and boosting technique | |
| CN113159937A (zh) | 识别风险的方法、装置和电子设备 | |
| CN115567224B (zh) | 一种用于检测区块链交易异常的方法及相关产品 | |
| Gupta et al. | Integrating Explainable AI in Financial Fraud Detection Systems for Enhanced Decision Transparency | |
| Kumar et al. | Banking fraud detection using optimized enhanced stacked autoencoder approach | |
| CN111582873A (zh) | 评估交互事件的方法及装置、电子设备、存储介质 | |
| CN116861226A (zh) | 一种数据处理的方法以及相关装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17878551 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17878551 Country of ref document: EP Kind code of ref document: A1 |




















