WO2020181910A1 - 识别风险商家的方法及装置 - Google Patents
识别风险商家的方法及装置 Download PDFInfo
- Publication number
- WO2020181910A1 WO2020181910A1 PCT/CN2020/070672 CN2020070672W WO2020181910A1 WO 2020181910 A1 WO2020181910 A1 WO 2020181910A1 CN 2020070672 W CN2020070672 W CN 2020070672W WO 2020181910 A1 WO2020181910 A1 WO 2020181910A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- merchant
- marked
- identified
- cluster
- similarity
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q20/00—Payment architectures, schemes or protocols
- G06Q20/38—Payment protocols; Details thereof
- G06Q20/40—Authorisation, e.g. identification of payer or payee, verification of customer or shop credentials; Review and approval of payers, e.g. check credit lines or negative lists
- G06Q20/401—Transaction verification
- G06Q20/4016—Transaction verification involving fraud or risk level assessment in transaction processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/22—Matching criteria, e.g. proximity measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/23—Clustering techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/06—Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
- G06Q10/063—Operations research, analysis or management
- G06Q10/0635—Risk analysis of enterprise or organisation activities
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q20/00—Payment architectures, schemes or protocols
- G06Q20/08—Payment architectures
- G06Q20/12—Payment architectures specially adapted for electronic shopping systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/018—Certifying business or products
- G06Q30/0185—Product, service or business identity fraud
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/06—Buying, selling or leasing transactions
- G06Q30/0601—Electronic shopping [e-shopping]
- G06Q30/0609—Qualifying participants for shopping transactions
Definitions
- the present disclosure relates to the field of computer technology, in particular, to methods and devices for identifying risky merchants.
- the anti-risk team With the development of e-commerce technology, more and more transactions are conducted through the Internet. In the process of Internet transactions, the anti-risk team generally monitors suspicious transaction behaviors and identifies risky behaviors in a timely manner.
- the present disclosure provides a method and device for identifying risky merchants.
- the risk in the merchants to be identified is identified based on the similarity between the labeled merchant clusters marked as risky merchants and the merchants to be identified Merchants can not only accurately identify risky merchants, but also reduce the amount of calculation in the identification process and improve identification efficiency.
- a method for identifying risky merchants includes: for each labeled merchant cluster in at least one labeled merchant cluster, determining the metric value of the merchant to be identified and the user associated with the labeled merchant cluster.
- the similarity between the business to be identified and the marked business cluster, each marked business in the marked business clusters is marked as a risk business; and based on the similarity between the business to be identified and the marked business cluster, the identification Whether the business to be identified is a risk business.
- the user correlation metric value between the merchant to be identified and the marked merchant cluster may include the user correlation coefficient between the merchant to be identified and the marked merchant cluster, and the user correlation coefficient between the merchant to be identified and the merchant cluster.
- the number of identical users in the marked merchant cluster is described.
- the user correlation coefficient may include: among the users of the merchants to be identified, the proportion of users who are the same as those of the marked merchant cluster; and/or the merchants to be identified and the users of the marked merchant cluster Feature similarity.
- each of the labeled merchant clusters has at least one representative labeled merchant
- the merchant to be identified is determined based on the metric value associated with the user of the labeled merchant cluster in the at least one merchant to be identified
- the similarity with the marked merchant cluster includes: based on the vector representation of the to-be-identified merchant for the marked merchant cluster and the vector representation of at least one representative of the marked merchant in the marked merchant cluster, determining that the to-be-identified merchant and the at least one Representing the similarity of the marked merchant; and determining the similarity between the to-be-identified merchant and the marked merchant cluster based on the similarity between the to-be-identified merchant and the at least one representative marked merchant.
- the vector representation for the merchants of the marked merchant cluster is a vector representation based on the user correlation coefficient of the merchant and the marked merchant cluster and the same number of users of the merchant and the marked merchant cluster.
- the similarity between the merchant to be identified and the at least one representative marked merchant is determined Previously, based on the user correlation metric value between the merchant to be identified and the marked merchant cluster, determining the similarity between the merchant to be identified and the marked merchant cluster further includes: the same user number dimension and user correlation coefficient dimension in the vector representation Treated as having the same value range.
- the at least one representative marked merchant in each marked merchant cluster is determined based on the number of users of each marked merchant in the marked merchant cluster.
- the similarity between the to-be-identified business and at least one representative marked business in each marked business cluster is characterized by any one of the following: Euclidean distance, Manhattan distance, and cosine of the angle distance.
- identifying the at least one at-risk merchant in the at least one to-be-identified merchant cluster includes: for each of the marked merchant clusters, at least Among the similarities between the two to-be-identified merchants and the marked merchant cluster, the previously predetermined number of merchants with the largest similarity value greater than the predetermined threshold are identified as risky merchants of the category corresponding to the marked merchant cluster.
- an apparatus for identifying risky merchants including: a similarity determination unit configured to target each labeled merchant cluster in at least one labeled merchant cluster, based on the business to be identified and the labeled merchant cluster To determine the similarity between the merchant to be identified and the marked merchant cluster, each marked merchant in each marked merchant cluster is marked as a risk merchant of a corresponding risk category; and a risk merchant identification unit is configured To identify whether the merchant to be identified is a risk merchant based on the similarity between the merchant to be identified and the clusters of each marked merchant.
- the user correlation metric value of the merchant to be identified and the marked merchant cluster includes the user correlation coefficient of the merchant to be identified and the marked merchant cluster, and the user correlation coefficient between the merchant to be identified and the merchant cluster. Mark the number of identical users in the merchant cluster.
- the user correlation coefficient may include: among the users of the merchants to be identified, the proportion of users who are the same as those of the marked merchant cluster; and/or the merchants to be identified and the users of the marked merchant cluster Feature similarity.
- each of the marked merchant clusters has at least one representative marked merchant
- the similarity determination unit includes: a first similarity determination module configured to be based on the pending mark for the marked merchant cluster. Identify the vector representation of the merchant and the vector representation of at least one representative of the marked merchant in the cluster of marked merchants, determine the similarity between the to-be-identified merchant and the at least one representative marked merchant; and the second similarity determination module is configured to be based on all The similarity between the business to be identified and the at least one representative marked business is determined, and the similarity between the business to be identified and the marked business cluster is determined.
- the vector representation for the merchants of the marked merchant cluster is a vector representation based on the user correlation coefficient of the merchant and the marked merchant cluster and the same number of users of the merchant and the marked merchant cluster.
- the similarity determination unit further includes: a normalization processing module configured to determine the relationship between the merchant to be identified and the user associated metric value of the cluster of marked merchants Before marking the similarity of the merchant cluster, the same user quantity dimension and the user correlation coefficient dimension in the vector representation are processed as having the same value range.
- a normalization processing module configured to determine the relationship between the merchant to be identified and the user associated metric value of the cluster of marked merchants Before marking the similarity of the merchant cluster, the same user quantity dimension and the user correlation coefficient dimension in the vector representation are processed as having the same value range.
- the risky merchant identification unit is configured to: for each of the marked merchant clusters, determine the similarity between at least two merchants to be identified and the marked merchant cluster that is greater than a predetermined threshold The previously predetermined number of businesses to be identified with the largest value are identified as risk businesses in the category corresponding to the marked business cluster.
- a computing device including: at least one processor; and a memory that stores instructions, and when the instructions are executed by the at least one processor, the at least one The processor executes the method of identifying risk merchants as described above.
- a non-transitory machine-readable storage medium that stores executable instructions that, when executed, cause the machine to execute the method for identifying risky merchants as described above.
- the similarity between the merchant to be identified and the marker merchant cluster is calculated based on the user correlation metric value of the merchant to be identified and the risky labeled merchant cluster, and then the merchant to be identified is identified based on the similarity
- the risk merchants in the cluster can be comprehensively identified based on the characteristics of multiple marked merchants in the marked merchant cluster. This not only can accurately identify risky merchants, but also has a small amount of calculation in the identification process, which can also improve Recognition efficiency.
- the similarity between the merchant to be identified and the cluster of marked merchants is calculated based on the user correlation coefficient between the merchant to be identified and the cluster of marked merchants and the number of identical users of the merchant to be identified and the cluster of marked merchants Therefore, the similarity between the to-be-identified business and the marked-business cluster can be determined based on the relative and absolute correlation attributes of the business to be identified and the marked business cluster, so as to improve the accuracy of risk identification.
- each labeled merchant cluster is represented by a labeled merchant, so that the merchant to be identified and each labeled merchant cluster are determined based on the vector representation of the merchant to be identified and the representative labeled merchant of each labeled merchant cluster
- the similarity can reduce the complexity of the identification process of risky businesses and further improve the identification efficiency.
- the same user quantity dimension and user correlation coefficient dimension in the vector representation are processed as having the same value range, This avoids ignoring the characteristics of a certain dimension due to the difference in the value range of each dimension, thereby further improving the accuracy of identifying risky merchants.
- the representative of each labeled merchant cluster is determined based on the number of users of each labeled merchant in each labeled merchant cluster, so that the representative that can represent the labeled merchant cluster can be selected according to the actual situation Mark merchants to increase the flexibility of the identification process.
- Fig. 1 is a flowchart of a method for identifying risky merchants according to an embodiment of the present disclosure
- FIG. 2 is a flowchart of an example of a similarity determination process in the method for identifying risky merchants according to an embodiment of the present disclosure
- FIG. 3 is a flowchart of an example of a risky merchant identification process in a method for identifying a risky merchant according to an embodiment of the present disclosure
- Fig. 4 is a structural block diagram of an apparatus for identifying risky merchants according to an embodiment of the present disclosure
- Fig. 5 is a structural block diagram of an example of a similarity determination unit in an apparatus for identifying risky merchants according to an embodiment of the present disclosure
- Fig. 6 is a structural block diagram of a computing device for implementing a method for recognizing user intentions according to another embodiment of the present disclosure.
- the term “including” and its variants means open terms, meaning “including but not limited to.”
- the term “based on” means “based at least in part on.”
- the terms “one embodiment” and “an embodiment” mean “at least one embodiment.”
- the term “another embodiment” means “at least one other embodiment.”
- the terms “first”, “second”, etc. may refer to different or the same objects. Other definitions can be included below, either explicit or implicit. Unless clearly indicated in the context, the definition of a term is consistent throughout the specification.
- Fig. 1 is a flowchart of a method for identifying risky merchants according to an embodiment of the present disclosure.
- the similarity between the to-be-identified merchant and the marked merchant cluster is determined based on the user correlation metric value between the merchant to be identified and the marked merchant cluster .
- Each marked merchant in each marked merchant cluster is marked as a risk merchant of the corresponding risk category.
- Risky merchants refer to merchants with risks of abnormal transactions (for example, illegal transactions such as money laundering criminal transactions).
- the collected risk merchants can be manually marked, and risk merchants marked as the same category are aggregated into the same marked merchant cluster.
- a trained classification model can also be used to cluster known risky merchants into clusters of labeled merchants.
- the classification model can be trained using the collected user data of the merchant.
- the risk category corresponding to each marked merchant cluster may be, for example, drug risk, smuggling risk, gambling risk, and future fraud risk.
- the user groups of each marked merchant in the marked merchant cluster can be merged to serve as the user group of the marked merchant cluster. In this way, the user characteristics of each marked merchant can be integrated into the user characteristics of the corresponding marked merchant cluster.
- the user of each merchant can be determined based on the fund exchange relationship between the user and the merchant.
- the risky merchants among the merchants to be identified can be identified by determining the similarity between the merchant to be identified and each marked merchant cluster, instead of calculating the merchants to be identified and each identified merchant cluster. Knowingly mark the similarity of merchants, thereby greatly reducing the amount of calculation and improving the efficiency of calculation.
- the similarity is calculated based on the user groups generated by combining the user groups of each marked merchant in the marked merchant cluster, which can enrich the characteristics of each marked merchant cluster, thereby improving the accuracy of identifying risky merchants.
- the user correlation metric value between the merchant to be identified and the marked merchant cluster may include the user correlation coefficient between the merchant to be identified and the marked merchant cluster and the number of identical users of the merchant to be identified and the marked merchant cluster.
- the user correlation coefficient may be the proportion of the users of the merchants to be identified that are the same as the users of the marked merchant cluster. That is, the user correlation coefficient can be calculated by using mathematical formula 1.
- A represents the number of users of the business to be identified
- G represents the number of users of the marked business cluster
- M(A, G) represents the user correlation coefficient between the business to be identified and the marked business cluster
- a ⁇ G represents the business to be identified The same number of users as the marked merchant cluster.
- the user correlation coefficient may also include the similarity of user characteristics between the merchant to be identified and the cluster of marked merchants.
- a trained classification model can be used to determine the similarity of user characteristics between the to-be-identified merchant and each labeled merchant cluster based on the user data of the merchant to be identified, and then the merchant to be identified can be determined based on the determined user-feature similarity Similarity with each marked merchant cluster.
- a feature extraction model (such as a logistic regression model, etc.) can also be used to extract user features from the user data (such as user basic information, user behavior data, etc.) of the merchant to be identified and the merchant cluster marked, and then based on the The extracted user characteristics of the merchant to be identified and the user characteristics of each marked merchant cluster are used to determine the similarity of the user characteristics between the merchant to be identified and each marked merchant cluster.
- the user correlation coefficient can be regarded as the relative user correlation metric value between the merchant to be identified and the marked merchant cluster, and the same number of users can be regarded as the absolute user correlation measure between the merchant to be identified and the marked merchant cluster. Therefore, this example can determine the similarity based on the relative and absolute associated attributes of the merchant to be identified and the cluster of marked merchants, thereby making the calculated similarity more accurate.
- the maximum value of the similarity between the merchant to be identified and each cluster of marked merchants may be determined as the risk coefficient of the merchant to be identified. Then, based on the risk coefficient of the merchant to be identified and a predetermined risk threshold, it can be determined whether the merchant to be identified is a risk merchant. For example, it can be set to determine that the merchant to be identified is a risk merchant when the risk coefficient of the merchant is higher than a certain risk threshold.
- each cluster of marked merchants may have at least one representative marked merchant.
- the representative marked merchant may be determined based on the number of users of each marked merchant in the marked merchant cluster. For example, the marked merchants in each marked merchant cluster can be sorted according to the number of users.
- the representative marked merchant may be, for example, the marked merchant whose ordinal number is the median in the ranking result, or two or more marked merchants in the middle position in the ranking result. In addition, it can also be two or more marked merchants separated by a predetermined ordinal number in the sorting.
- the representative labeled merchants can be selected from the first to 300th labeled merchants every 50 ordinal numbers.
- the similarity between the merchant to be identified and the corresponding marked merchant cluster can be determined by determining the similarity between the merchant to be identified and the representative marked merchant.
- FIG. 2 is a flowchart of an example of determining similarity based on a representative labeled merchant in a method for identifying risky merchants according to an embodiment of the present disclosure.
- the vector representation of the to-be-identified merchant for the labeled merchant cluster and the vector representation of at least one representative of the labeled merchant in the labeled merchant cluster determine the relationship between the to-be-identified merchant and the at least one representative labeled merchant. Similarity.
- the vector representation of the merchants of a certain labeled merchant cluster is established based on the user correlation coefficient and the same number of users of the merchant and the labeled merchant cluster.
- the similarity between the business to be identified and the representative marked business can be characterized by any one of Euclidean distance, Manhattan distance, angle cosine distance, and the like.
- the following uses the Euclidean distance as an example to illustrate an example of calculating the similarity determination process between the merchant to be identified and the merchant on behalf of the mark.
- the correlation coefficient between the merchant to be identified or the user of the marked merchant cluster and the marked merchant cluster is the proportion of the users of the merchant who are the same as the marked merchant cluster.
- the vector representation of the merchant to be identified can be (M(A, G),
- G p represents the number of users who mark the merchant. Since the users representing the marked merchant are all from the marked merchant cluster, the value of the user correlation coefficient M(G p , G) between the marked merchant and the marked merchant cluster is 1, which represents the same user of the marked merchant and the marked merchant cluster The quantity
- is G p .
- the similarity between the to-be-identified merchant and the representative marked merchant can be determined by the following mathematical formula 2.
- G p represents the number of users representing the marked merchant
- D(A, G p ) is the Euclidean distance between the merchant to be identified and the representative marked merchant.
- the similarity between the to-be-identified merchant and the marked merchant cluster is determined.
- the value of the user correlation coefficient in the vector representation is small (for example, the value range is [0, 1]).
- the user correlation coefficient dimension may be covered by the same user number dimension. This may lead to insufficient accuracy of the determined similarity.
- the same user quantity dimension and user correlation coefficient dimension in the vector representation can be processed as having the same value range. Then, the operations of block 202 to block 204 are performed based on the processed vector representation.
- the dimension of the same number of users in the vector representation can be normalized.
- the following mathematical formula 3 can be used to normalize the same user quantity dimension.
- the vector of the merchant to be identified after normalization is expressed as The vector representing the marked merchant is represented as (1,1).
- the similarity between the to-be-identified merchant and the representative marked merchant can be modified to the following mathematical formula 4.
- the value range of D (A, G p ) in Math 4 is The similarity calculation formula can be further modified based on Mathematical formula 4. For example, the final similarity can be calculated based on the following mathematical formula 5.
- S(A, G p ) is the final similarity between the merchant to be identified and the merchant with the representative mark, and its value range is [0,1].
- the similarity between the merchant to be identified and the cluster of marked merchants can be determined. For example, if the marked merchant cluster has a representative marked merchant, the similarity between the merchant to be identified and the representative marked merchant may be determined as the similarity between the merchant to be identified and the marked merchant cluster. In one example, a cluster of marked merchants may have more than two representative marked merchants. At this time, the similarity between the business to be identified and each representative marked business can be averaged to obtain the similarity between the business to be identified and the marked business cluster.
- the representative marked merchants are selected based on the number of users of each marked merchant in the marked merchant cluster
- multiple representative marked merchants can be assigned different weights based on the number of users of each representative marked merchant, so as to be recognized
- a weighted average or a weighted sum of the similarity between the business and each representative marked business is taken to obtain the similarity between the business to be identified and the marked business cluster.
- Fig. 3 is a flowchart of an example of a risky merchant identification process in a method for identifying a risky merchant according to an embodiment of the present disclosure.
- the maximum value of the similarity between the merchant to be identified and each marked merchant cluster may also be determined as the risk coefficient of the merchant to be identified. Then, the previously predetermined number of businesses to be identified with the largest value in the crime risk coefficient of the businesses to be identified are determined as risk businesses. Alternatively, the previously predetermined number of businesses to be identified with the largest value greater than the predetermined threshold among the crime risk coefficients of businesses to be identified is determined as risk businesses.
- the determined risky merchants can be further subjected to verification processing to exclude the merchants with lower risk.
- the data of the identified risky merchants can be sent to the anti-risk monitoring team, and the professionals in the team can further analyze these merchants to finally determine the risky merchants.
- FIG. 4 is a structural block diagram of a risky merchant identification device 400 according to an embodiment of the present disclosure. As shown in FIG. 4, the risky merchant identification device 400 includes a similarity determination unit 410 and a risky merchant identification unit 420.
- the similarity determination unit 410 is configured to determine the similarity between the to-be-identified merchant and the labeled merchant cluster based on the user correlation metric value of the to-be-identified merchant and the labeled merchant cluster for each labeled merchant cluster in the at least one labeled merchant cluster.
- Each marked merchant in the marked merchant cluster is marked as a risk merchant of a corresponding risk category.
- the risk merchant identification unit 420 is configured to identify at least one risk merchant among the merchants to be identified based on the similarity between the merchant to be identified and each of the marked merchant clusters.
- the risky merchant identification unit 420 may be configured to, for each marked merchant cluster, calculate the maximum similarity value of the at least two merchants to be identified and the marked merchant cluster that is greater than a predetermined threshold.
- the merchants to be identified are identified as risk merchants in the category corresponding to the marked merchant cluster.
- the user correlation metric value between the merchant to be identified and the marked merchant cluster may include the user correlation coefficient between the merchant to be identified and the marked merchant cluster and the number of identical users of the merchant to be identified and the marked merchant cluster.
- the user correlation coefficient may be the proportion of users of the merchant to be identified that are the same as the users of the marked merchant cluster.
- the user correlation coefficient may also be the feature similarity between the user of the merchant to be identified and the user who marks the merchant cluster.
- FIG. 5 is a structural block diagram of an example of the similarity determination unit 410 in the risky merchant identification device 400 shown in FIG. 4.
- each cluster of marked merchants may have at least one representative marked merchant.
- the similarity determination unit 410 may include a normalization processing module 411, a first similarity determination module 412 and a second similarity determination module 413.
- the normalization processing module 411 may not be included.
- the first similarity determination module 412 and the second similarity determination module 413 will be described first.
- the first similarity determination module 412 is configured to determine the relationship between the to-be-identified merchant and the at least one representative of the labeled merchant based on the vector representation of the to-be-identified merchant for the labeled merchant cluster and the vector representation of at least one representative of the labeled merchant in the labeled merchant cluster. Similarity.
- the second similarity determination module 413 is configured to determine the similarity between the merchant to be identified and the cluster of marked merchants based on the similarity between the merchant to be identified and the at least one representative marked merchant.
- the vector representation for the merchants of the marked merchant cluster is a vector representation based on the user correlation coefficient of the merchant and the marked merchant cluster and the same number of users of the merchant and the marked merchant cluster.
- the normalization processing module 411 can be used to determine the difference between the business to be identified in at least one business to be identified and the marked business cluster.
- User correlation metric value before determining the similarity between the business to be identified and the marked business cluster, the same user number dimension and user correlation coefficient dimension in the vector representation are processed as having the same value range. The example of normalization processing has been described in the above content, and will not be repeated here.
- the device for identifying risky merchants of the present disclosure can be implemented by hardware, or by software or a combination of hardware and software. Taking software implementation as an example, as a logical device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory through the processor of the device where it is located.
- the means for identifying application program controls displayed on the terminal device can be implemented using a computing device, for example.
- FIG. 6 is a structural block diagram of a computing device 600 of a method for identifying a risky merchant according to another embodiment of the present disclosure.
- the computing device 600 may include at least one processor 610, a memory 620, a memory 630, a communication interface 640, and an internal bus 650.
- the at least one processor 610 executes on a computer-readable storage medium (ie, the memory 620).
- At least one computer-readable instruction ie, the above-mentioned element implemented in the form of software stored or encoded in the computer.
- computer-executable instructions are stored in the memory 620, which, when executed, cause at least one processor 610 to: for each of the at least one tagged merchant cluster, based on the at least one to-be-identified merchant cluster
- the merchant associates a metric value with the user of the marked merchant cluster to determine the similarity between the to-be-identified merchant and the marked merchant cluster, and each marked merchant in each marked merchant cluster is marked as a risk merchant of a corresponding risk category; and
- the similarity between the merchant to be identified and the clusters of each marked merchant identifies the risky merchant among the at least one merchant to be identified.
- a program product such as a non-transitory machine-readable medium.
- the non-transitory machine-readable medium may have instructions (ie, the above-mentioned elements implemented in the form of software), which when executed by a machine, cause the machine to execute the various embodiments described above in conjunction with FIGS. 1-5 in the various embodiments of the present disclosure. Operation and function.
- a system or device equipped with a readable storage medium may be provided, and the software program code for realizing the function of any one of the above embodiments is stored on the readable storage medium, and the computer or device of the system or device The processor reads out and executes the instructions stored in the readable storage medium.
- the program code itself read from the readable medium can realize the function of any one of the above embodiments, so the machine readable code and the readable storage medium storing the machine readable code constitute the present invention a part of.
- Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tape, Volatile memory card and ROM.
- the program code can be downloaded from a server computer or cloud via a communication network.
Landscapes
- Business, Economics & Management (AREA)
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Accounting & Taxation (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Strategic Management (AREA)
- General Business, Economics & Management (AREA)
- Finance (AREA)
- Data Mining & Analysis (AREA)
- Economics (AREA)
- Human Resources & Organizations (AREA)
- Marketing (AREA)
- Entrepreneurship & Innovation (AREA)
- Development Economics (AREA)
- Computer Security & Cryptography (AREA)
- Life Sciences & Earth Sciences (AREA)
- Evolutionary Biology (AREA)
- Artificial Intelligence (AREA)
- Bioinformatics & Computational Biology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Engineering & Computer Science (AREA)
- Evolutionary Computation (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Game Theory and Decision Science (AREA)
- Educational Administration (AREA)
- Operations Research (AREA)
- Quality & Reliability (AREA)
- Tourism & Hospitality (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
一种识别风险商家的方法及装置。该方法包括:针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度(120),所述各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家;以及基于所述待识别商家与所述各个标记商家集群的相似度,识别待识别商家是否为风险商家(140)。
Description
本公开涉及计算机技术领域,具体地,涉及识别风险商家的方法及装置。
随着电商技术的发展,越来越来的交易通过互联网进行。在互联网交易过程中,反风险团队一般会对可疑交易行为进行监控,以及时识别出有风险的行为。
与此同时,可疑分子也会竭尽手段规避反风险团队的监控。因此,监控风险行为的任务量变得越来越庞大。此外,互联网交易快速便捷的特点也使得需要监控的交易行为的数量越来越大。因此,现有技术亟需能够协助反风险团队快速有效识别风险交易的技术。
发明内容
鉴于上述,本公开提供了一种识别风险商家的方法及装置,利用该方法和装置,通过基于被标记为风险商家的标记商家集群与待识别商家的相似度,来识别待识别商家中的风险商家,不仅能够准确识别风险商家,而且能够降低识别过程的计算量,提高识别效率。
根据本公开的一个方面,提供了一种识别风险商家的方法,包括:针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度,所述各个标记商家集群中的各个标记商家被标记为风险商家;以及基于所述待识别商家与所述各个标记商家集群的相似度,识别所述待识别商家是否为风险商家。
可选的,在一个示例中,所述待识别商家与所述标记商家集群的用户关联度量值可以包括所述待识别商家与所述标记商家集群的用户关联系数和所述待识别商家与所述标记商家集群的相同用户数量。其中,所述用户关联系数可以包括:在所述待识别商家的用户中,与所述标记商家集群的相同用户所占的比例;和/或所述待识别商家与所述标记商家集群的用户特征相似度。
可选的,在一个示例中,所述各个标记商家集群具有至少一个代表标记商家,基 于至少一个待识别商家中的待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度包括:基于针对该标记商家集群的所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度;以及基于所述待识别商家与该至少一个代表标记商家的相似度,确定所述待识别商家与该标记商家集群的相似度。其中,针对该标记商家集群的商家的向量表示是基于所述商家与该标记商家集群的用户关联系数和所述商家与该标记商家集群的相同用户数量的向量表示。
可选的,在一个示例中,在基于所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度之前,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度还包括:将所述向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。
可选的,在一个示例中,所述各个标记商家集群的至少一个代表标记商家是基于该标记商家集群中的各个标记商家的用户数量确定的。
可选的,在一个示例中,所述待识别商家与所述各个标记商家集群的至少一个代表标记商家的相似度用以下中的任一者来表征:欧氏距离、曼哈顿距离和夹角余弦距离。
可选的,在一个示例中,基于所述待识别商家与所述各个标记商家集群的相似度,识别所述至少一个待识别商家中的风险商家包括:针对所述各个标记商家集群,将至少两个待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。
根据本公开的另一方面,还提供一种识别风险商家的装置,包括:相似度确定单元,被配置为针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度,所述各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家;以及风险商家识别单元,被配置为基于所述待识别商家与所述各个标记商家集群的相似度,识别所述待识别商家是否为风险商家。
可选的,在一个示例中,所述待识别商家与所述标记商家集群的用户关联度量值包括所述待识别商家与所述标记商家集群的用户关联系数和所述待识别商家与所述标记商家集群的相同用户数量。其中,所述用户关联系数可以包括:在所述待识别商家的 用户中,与所述标记商家集群的相同用户所占的比例;和/或所述待识别商家与所述标记商家集群的用户特征相似度。
可选的,在一个示例中,所述各个标记商家集群具有至少一个代表标记商家,所述相似度确定单元包括:第一相似度确定模块,被配置为基于针对该标记商家集群的所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度;以及第二相似度确定模块,被配置为基于所述待识别商家与该至少一个代表标记商家的相似度,确定所述待识别商家与该标记商家集群的相似度。其中,针对该标记商家集群的商家的向量表示是基于所述商家与该标记商家集群的用户关联系数和所述商家与该标记商家集群的相同用户数量的向量表示。
可选的,在一个示例中,所述相似度确定单元还包括:归一化处理模块,被配置为在基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度之前,将所述向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。
可选的,在一个示例中,所述风险商家识别单元被配置为:针对所述各个标记商家集群,将至少两个待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。
根据本公开的另一方面,还提供一种计算设备,包括:至少一个处理器;以及存储器,所述存储器存储指令,当所述指令被所述至少一个处理器执行时,使得所述至少一个处理器执行如上所述的识别风险商家的方法。
根据本公开的另一方面,还提供一种非暂时性机器可读存储介质,其存储有可执行指令,所述指令当被执行时使得所述机器执行如上所述的识别风险商家的方法。
利用本公开的识别风险商家的方法和装置,通过基于待识别商家与风险的标记商家集群的用户关联度量值来计算待识别商家与标记商家集群的相似度,进而基于该相似度识别待识别商家中的风险商家,从而能够综合地基于标记商家集群中的多个标记商家的特征来进行识别,由此不仅能准确识别出风险商家,而且在识别过程中的运算量较小,因而还能够提高识别效率。
利用本公开的识别风险商家的方法和装置,基于在待识别商家与标记商家集群的用户关联系数和待识别商家与标记商家集群的相同用户数量,来计算待识别商家与标记 商家集群的相似度,从而能够基于待识别商家与标记商家集群的相对关联属性和绝对关联属性来确定待识别商家与标记商家集群的相似度,以提高风险识别的准确性。
利用本公开的识别风险商家的方法和装置,以代表标记商家来代表各个标记商家集群,从而基于待识别商家和各个标记商家集群的代表标记商家的向量表示来确定待识别商家与各个标记商家集群的相似度,能够降低风险商家识别过程的复杂度,进一步提高识别效率。
利用本公开的识别风险商家的方法和装置,在确定代表标记商家与待识别商家间的相似度之前,将向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围,从而避免因为各维度的取值范围差异导致忽略某一维度的特征,由此能够进一步提高识别风险商家的准确性。
利用本公开的识别风险商家的方法和装置,基于各个标记商家集群中的各个标记商家的用户数量来确定各个标记商家集群的代表标记商家,从而能够根据实际情况选取能够代表该标记商家集群的代表标记商家,以提高识别过程的灵活性。
通过参照下面的附图,可以实现对于本公开内容的本质和优点的进一步理解。在附图中,类似组件或特征可以具有相同的附图标记。附图是用来提供对本发明实施例的进一步理解,并且构成说明书的一部分,与下面的具体实施方式一起用于解释本公开的实施例,但并不构成对本公开的实施例的限制。在附图中:
图1是根据本公开的一个实施例的识别风险商家的方法的流程图;
图2是根据本公开的一个实施例的识别风险商家的方法中的相似度确定过程的一个示例的流程图;
图3是根据本公开的一个实施例的识别风险商家的方法中的风险商家识别过程的一个示例的流程图;
图4是根据本公开的一个实施例的识别风险商家的装置的结构框图;
图5是根据本公开的一个实施例的识别风险商家的装置中的相似度确定单元的一个示例的结构框图;
图6是根据本公开的另一实施例的用于实现用于识别用户意图的方法的计算设备 的结构框图。
以下将参考示例实施方式讨论本文描述的主题。应该理解,讨论这些实施方式只是为了使得本领域技术人员能够更好地理解从而实现本文描述的主题,并非是对权利要求书中所阐述的保护范围、适用性或者示例的限制。可以在不脱离本公开内容的保护范围的情况下,对所讨论的元素的功能和排列进行改变。各个示例可以根据需要,省略、替代或者添加各种过程或组件。另外,相对一些示例所描述的特征在其它例子中也可以进行组合。
如本文中使用的,术语“包括”及其变型表示开放的术语,含义是“包括但不限于”。术语“基于”表示“至少部分地基于”。术语“一个实施例”和“一实施例”表示“至少一个实施例”。术语“另一个实施例”表示“至少一个其他实施例”。术语“第一”、“第二”等可以指代不同的或相同的对象。下面可以包括其他的定义,无论是明确的还是隐含的。除非上下文中明确地指明,否则一个术语的定义在整个说明书中是一致的。
现在结合附图来描述本公开的识别风险商家的方法及装置。
图1是根据本公开的一个实施例的识别风险商家的方法的流程图。
如图1所示,在块120,针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定待识别商家与该标记商家集群的相似度,各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家。风险商家是指有进行异常交易风险(例如洗钱犯罪交易行为等违法交易行为)的商家。
可以由人工对所收集的风险商家进行标注,并将被标注为相同类别的风险商家聚合为同一个标记商家集群。还可以利用经过训练的分类模型将已知的风险商家集群为各个标记商家集群。分类模型可以利用所收集的商家的用户数据来训练。各个标记商家集群对应的风险类别例如可以是毒品风险、走私风险、赌博风险、期诈风险等类别。在获得各个标记商家集群之后,可以合并该标记商家集群中的各个标记商家的用户群体以作为该标记商家集群的用户群体。由此,能够将各个标记商家的用户特征综合为对应的标记商家集群的用户特征。各个商家的用户可以基于用户与商家之间的资金往来关系来确定。
合并同类别的风险商家以获得标记商家集群时,可以通过确定待识别商家与各个标记商家集群的相似度来识别待识别商家中的风险商家,而不需要一一计算待识别商家与每个已知的标记商家的相似度,从而大幅降低了运算量,提高了运算效率。此外,基于标记商家集群中的各个标记商家的用户群体合并而生成的用户群体来计算相似度,能够丰富各个标记商家集群的特征,从而能够提高识别风险商家的准确性。
在一个示例中,待识别商家与标记商家集群的用户关联度量值可以包括待识别商家与标记商家集群的用户关联系数和待识别商家与标记商家集群的相同用户数量。其中,用户关联系数可以是在待识别商家的用户中,与标记商家集群的相同用户所占的比例。即,用户关联系数可以利用如一数学式一来计算。
数学式一:
在数学式一中,A表示待识别商家的用户数量,G表示标记商家集群的用户数量,M(A,G)表示待识别商家与标记商家集群的用户关联系数,A∩G表示待识别商家与标记商家集群的相同用户数量。对于数学式一,可以定为当A的值为0时,M(A,G)的值为0。
用户关联系数还可以包括待识别商家与标记商家集群的用户特征相似度。在一个示例中,可以基于待识别商家的用户数据,利用经过训练的分类模型来确定待识别商家与各个标记商家集群的用户特征相似度,进而基于所确定的用户特征相似度来确定待识别商家与各个标记商家集群的相似度。在另一示例中,还可以利用特征提取模型(例如逻辑回归模型等)来从待识别商家和标记商家集群的用户数据(例如用户基本信息、用户行为数据等)中提取用户特征,进而基于所提取的待识别商家的用户特征和各个标记商家集群的用户特征,来确定待识别商家与各个标记商家集群的用户特征相似度。
在该示例中,用户关联系数可以看作待识别商家与标记商家集群之间的相对用户关联度量值,而相同用户数量可以看作待识别商家与标记商家集群的绝对用户关联度量值。因此,该示例可以基于待识别商家与标记商家集群的相对关联属性和绝对关联属性来确定相似度,从而使所计算的相似度更加准确。
在确定待识别商家与各个标记商家集群的相似度之后,在块140,基于待识别商家与各个标记商家集群的相似度,识别待识别商家是否为风险商家。
在一个示例中,可以将待识别商家与各个标记商家集群的相似度中的最大值确定 为该待识别商家的风险系数。然后,可以基于待识别商家的风险系数和预定风险阈值,来确定该待识别商家是否为风险商家。例如,可以设置为在待识别商家的风险系数高于某一风险阈值时,确定其为风险商家。
在一个示例中,各个标记商家集群可以有至少一个代表标记商家。代表标记商家可以基于标记商家集群中的各个标记商家的用户数量来确定。例如,可以将各个标记商家集群中的标记商家按照用户数量进行排序。代表标记商家例如可以是排序结果中序数为中位数所对应的标记商家,还可以是提排序结果中居于中间位置的两个以上标记商家。此外,还可以是排序中间隔预定序数的两个以上标记商家。例如,如果某一标记商家集群中的标记商家按用户数量排序后的序数为1至300,可以每隔50个序数从第1个至第300个标记商家中选取代表标记商家。当各个标记商家集群具有代表标记商家时,可以通过确定待识别商家与代表标记商家的相似度,来确定待识别商家与对应标记商家集群的相似度。
图2是根据本公开的一个实施例的识别风险商家的方法中的基于代表标记商家来确定相似度的示例的流程图。
如图2所示,在块202,基于针对该标记商家集群的待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定待识别商家与该至少一个代表标记商家的相似度。针对某一标记商家集群的商家的向量表示是基于该商家与该标记商家集群的用户关联系数和相同用户数量而建立的。
待识别商家与代表标记商家的相似度可以用欧氏距离、曼哈顿距离、夹角余弦距离等中的任一者来表征。以下以利用欧氏距离的情形为例,来说明计算待识别商家与代表标记商家的相似度确定过程的示例。在下述示例中,待识别商家或代表标记商家与标记商家集群的用户关联系数为在该商家的用户中与该标记商家集群的相同用户所占的比例。
针对某一标记商家集群,待识别商家的向量表示可以是(M(A,G),|A∩G|),该标记商家集群的代表标记商家的向量表示可以是(M(G
p,G),|G
p∩G|)。其中,G
p表示代表标记商家的用户数量。由于代表标记商家的用户全部来自于该标记商家集群,因而代表标记商家与该标记商家集群的用户关联系数M(G
p,G)的值为1,代表标记商家与该标记商家集群的相同用户数量|G
p∩G|为G
p。待识别商家与代表标记商家之间的相似度可以利用如下数学式二来确定。
数学式二:
在数学式二中,G
p表示代表标记商家的用户数量,D(A,G
p)为待识别商家与代表标记商家之间的欧氏距离。
在块204,基于待识别商家与至少一个代表标记商家的相似度,确定待识别商家与该标记商家集群的相似度。
通常向量表示中的用户关联系数的取值较小(例如,取值范围为[0,1])。当待识别商家与某一标记商家集群的相同用户数量较大时,由于在向量表示中相同用户数量维度的取值远大于用户关联系数,因而用户关联系数维度可能会被相同用户数量维度覆盖。这会导致所确定的相似度不够准确。
因而,在执行块202的操作之前,可以将向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。然后基于处理后的向量表示执行块202至块204的操作。
例如,可以对向量表示中的相同用户数量维度进行归一化处理。可以利用如下数学式三对相同用户数量维度进行归一化处理。
数学式三:
数学式四:
数学式五:
在数学式五中,S(A,G
p)为最终得出的待识别商家与代表标记商家的相似度,其取值范围为[0,1]。
以上虽然示出了利用数学式三进行归一化处理的情形,但应当理解的是,还可以采用其它方式来对相同用户数量维度进行归一化处理。例如,可以将各个向量表示中的相同用户数量的值除以相同用户数量维度的最大值,以进行归一化处理。
在确定出待识别商家与代表标记商家之间的相似度之后,可以确定待识别商家与标记商家集群之间的相似度。例如,如果标记商家集群具有一个代表标记商家,则可以将待识别商家与该代表标记商家的相似度确定为待识是商家与该标记商家集群的相似度。在一个示例中,标记商家集群可以具有两个以上代表标记商家。此时,可以对待识别商家与各个代表标记商家的相似度取平均值,从而得出待识别商家与该标记商家集群的相似度。此外,当代表标记商家是基于标记商家集群中的各个标记商家的用户数量排序而选取的多个代表标记商家时,还可以基于各个代表标记商家的用户数量为其赋予不同的权重,从而对待识别商家与各个代表标记商家的相似度取加权平均值或加权求和,以得到待识别商家与该标记商家集群的相似度。
在确定出待识别商家与各个标记商家集群的相似度时,可以针对各个标记商家集群,将待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。图3是根据本公开的一个实施例的识别风险商家的方法中的风险商家识别过程的一个示例的流程图。
如图3所示,在块302,针对各个标记商家集群,对至少两个待识别商家与该标记商家集群的相似度进行排序。
然后,在块304,将排序结果中大于预定阈值的相似度值最大的前预定个数的待识别商家识别为该标记商家集群所对应的类别的风险商家。
此外,还可以将待识别商家与各个标记商家集群的相似度中的最大值确定为该待识别商家的风险系数。然后将待识别商家的犯罪风险系数中的值最大的前预定个数的待识别商家确定为风险商家。或者,将待识别商家的犯罪风险系数中的大于预定阈值的值最大的前预定个数的待识别商家确定为风险商家。
在确定出风险商家之后,还可以进一步对所确定出的风险商家进行验证征处理,以排除其中风险较低的商家。例如,可以将所确定出的风险商家的数据发送至反风险监控团队,由团队中的专业人员进一步对这些商家进行分析,以最终确定风险商家。
图4是根据本公开的一个实施例的风险商家识别装置400的结构框图。如图4所示,风险商家识别装置400包括相似度确定单元410和风险商家识别单元420。
相似度确定单元410被配置为针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定待识别商家与该标记商家集群的相似度,各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家。风险商家识别单元420被配置为基于待识别商家与各个标记商家集群的相似度,识别至少一个待识别商家中的风险商家。
在一个示例中,风险商家识别单元420可以被配置为针对各个标记商家集群,将至少两个待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。
在一个示例中,待识别商家与标记商家集群的用户关联度量值可以包括待识别商家与标记商家集群的用户关联系数和待识别商家与标记商家集群的相同用户数量。其中,用户关联系数可以为在待识别商家的用户中,与标记商家集群的相同用户所占的比例。用户关联系数还可以是待识别商家的用户与标记商家集群的用户的特征相似度。
图5是图4所示的风险商家识别装置400中的相似度确定单元410的一个示例的结构框图。在该示例中,各个标记商家集群可以具有至少一个代表标记商家。如图5所示,相似度确定单元410可以包括归一化处理模块411、第一相似度确定模块412和第二相似度确定模块413。其中,在另一示例中,可以不包括归一化处理模块411。以下为了说明上的方便性,首先对第一相似度确定模块412和第二相似度确定模块413进行说明。
第一相似度确定模块412被配置为基于针对该标记商家集群的待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定待识别商家与该至少一个代表标记商家的相似度。第二相似度确定模块413被配置为基于待识别商家与该至少一个代表标记商家的相似度,确定待识别商家与该标记商家集群的相似度。其中,针对该标记商家集群的商家的向量表示是基于商家与该标记商家集群的用户关联系数和所述商家与该标记商家集群的相同用户数量的向量表示。
为了避免向量表示中不同维度的取值范围差异过大而导致所确定的相似度不准确,可以利归一化处理模块411在基于至少一个待识别商家中的待识别商家与该标记商家集群的用户关联度量值,确定待识别商家与该标记商家集群的相似度之前,将向量表示中 的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。关于归一化处理的示例已在上述内容中进行了说明,此处不再赘述。
以上参照图1-5对识别风险商家的方法和装置进行了说明。需要说明的是,在以上对识别风险商家的方法的说明中提及的细节同样适用于识别风险商家的装置。本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见。
本公开的识别风险商家的装置可以采用硬件实现,也可以采用软件或者硬件和软件的组合来实现。以软件实现为例,作为一个逻辑意义上的装置,是通过其所在设备的处理器将非易失性存储器中对应的计算机程序指令读取到内存中运行形成的。在本公开中,识别终端设备上显示的应用程序控件的装置例如可以利用计算设备实现。
图6是根据本公开的另一实施例的识别风险商家的方法的计算设备600的结构框图。如图6所示,计算设备600可以包括至少一个处理器610、存储器620、内存630、通信接口640以及内部总线650,该至少一个处理器610执行在计算机可读存储介质(即,存储器620)中存储或编码的至少一个计算机可读指令(即,上述以软件形式实现的元素)。
在一个实施例中,在存储器620中存储计算机可执行指令,其当执行时使得至少一个处理器610:针对至少一个标记商家集群中的各个标记商家集群,基于至少一个待识别商家中的待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度,所述各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家;以及基于所述待识别商家与所述各个标记商家集群的相似度,识别所述至少一个待识别商家中的风险商家。
应该理解,在存储器620中存储的计算机可执行指令当执行时使得至少一个处理器610进行本公开的各个实施例中以上结合图1-5描述的各种操作和功能。
根据一个实施例,提供了一种例如非暂时性机器可读介质的程序产品。非暂时性机器可读介质可以具有指令(即,上述以软件形式实现的元素),该指令当被机器执行时,使得机器执行本公开的各个实施例中以上结合图1-5描述的各种操作和功能。
具体地,可以提供配有可读存储介质的系统或者装置,在该可读存储介质上存储着实现上述实施例中任一实施例的功能的软件程序代码,且使该系统或者装置的计算机或处理器读出并执行存储在该可读存储介质中的指令。
在这种情况下,从可读介质读取的程序代码本身可实现上述实施例中任何一项实 施例的功能,因此机器可读代码和存储机器可读代码的可读存储介质构成了本发明的一部分。
可读存储介质的实施例包括软盘、硬盘、磁光盘、光盘(如CD-ROM、CD-R、CD-RW、DVD-ROM、DVD-RAM、DVD-RW、DVD-RW)、磁带、非易失性存储卡和ROM。可选择地,可以由通信网络从服务器计算机上或云上下载程序代码。
上述对本说明书特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
在整个本说明书中使用的术语“示例性”意味着“用作示例、实例或例示”,并不意味着比其它实施例“优选”或“具有优势”。出于提供对所描述技术的理解的目的,具体实施方式包括具体细节。然而,可以在没有这些具体细节的情况下实施这些技术。在一些实例中,为了避免对所描述的实施例的概念造成难以理解,公知的结构和装置以框图形式示出。
以上结合附图详细描述了本公开的实施例的可选实施方式,但是,本公开的实施例并不限于上述实施方式中的具体细节,在本公开的实施例的技术构思范围内,可以对本公开的实施例的技术方案进行多种简单变型,这些简单变型均属于本公开的实施例的保护范围。
本公开内容的上述描述被提供来使得本领域任何普通技术人员能够实现或者使用本公开内容。对于本领域普通技术人员来说,对本公开内容进行的各种修改是显而易见的,并且,也可以在不脱离本公开内容的保护范围的情况下,将本文所定义的一般性原理应用于其它变型。因此,本公开内容并不限于本文所描述的示例和设计,而是与符合本文公开的原理和新颖性特征的最广范围相一致。
Claims (14)
- 一种识别风险商家的方法,包括:针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度,所述各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家;以及基于所述待识别商家与所述各个标记商家集群的相似度,识别所述待识别商家是否为风险商家。
- 如权利要求1所述的方法,其中,所述待识别商家与所述标记商家集群的用户关联度量值包括所述待识别商家与所述标记商家集群的用户关联系数和所述待识别商家与所述标记商家集群的相同用户数量,其中,所述用户关联系数包括:在所述待识别商家的用户中,与所述标记商家集群的相同用户所占的比例;和/或所述待识别商家与所述标记商家集群的用户特征相似度。
- 如权利要求2所述的方法,其中,所述各个标记商家集群具有至少一个代表标记商家,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度包括:基于针对该标记商家集群的所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度;以及基于所述待识别商家与该至少一个代表标记商家的相似度,确定所述待识别商家与该标记商家集群的相似度,其中,针对该标记商家集群的商家的向量表示是基于所述商家与该标记商家集群的用户关联系数和所述商家与该标记商家集群的相同用户数量的向量表示。
- 如权利要求3所述的方法,其中,在基于所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度之前,基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度还包括:将所述向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。
- 如权利要求3所述的方法,其中,所述各个标记商家集群的至少一个代表标记商家是基于该标记商家集群中的各个标记商家的用户数量确定的。
- 如权利要求3所述的方法,其中,所述待识别商家与所述各个标记商家集群的至少一个代表标记商家的相似度用以下中的任一者来表征:欧氏距离、曼哈顿距离和夹角余弦距离。
- 如权利要求1所述的方法,其中,基于所述待识别商家与所述各个标记商家集群的相似度,识别所述待识别商家中的风险商家包括:针对所述各个标记商家集群,将所述至少两个待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。
- 一种识别风险商家的装置,包括:相似度确定单元,被配置为针对至少一个标记商家集群中的各个标记商家集群,基于待识别商家中的待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度,所述各个标记商家集群中的各个标记商家被标记为相应风险类别的风险商家;以及风险商家识别单元,被配置为基于所述待识别商家与所述各个标记商家集群的相似度,识别所述待识别商家是否为风险商家。
- 如权利要求8所述的装置,其中,所述待识别商家与所述标记商家集群的用户关联度量值包括所述待识别商家与所述标记商家集群的用户关联系数和所述待识别商家与所述标记商家集群的相同用户数量,其中,所述用户关联系数包括:在所述待识别商家的用户中,与所述标记商家集群的相同用户所占的比例;和/或所述待识别商家与所述标记商家集群的用户特征相似度。
- 如权利要求8或9所述的装置,其中,所述各个标记商家集群具有至少一个代表标记商家,所述相似度确定单元包括:第一相似度确定模块,被配置为基于针对该标记商家集群的所述待识别商家的向量表示和该标记商家集群的至少一个代表标记商家的向量表示,确定所述待识别商家与该至少一个代表标记商家的相似度;以及第二相似度确定模块,被配置为基于所述待识别商家与该至少一个代表标记商家的相似度,确定所述待识别商家与该标记商家集群的相似度,其中,针对该标记商家集群的商家的向量表示是基于所述商家与该标记商家集群的用户关联系数和所述商家与该标记商家集群的相同用户数量的向量表示。
- 如权利要求10所述的装置,其中,所述相似度确定单元还包括:归一化处理模块,被配置为在基于待识别商家与该标记商家集群的用户关联度量值,确定所述待识别商家与该标记商家集群的相似度之前,将所述向量表示中的相同用户数量维度和用户关联系数维度处理为具有相同的取值范围。
- 如权利要求8所述的装置,其中,所述风险商家识别单元被配置为:针对所述各个标记商家集群,将至少两个待识别商家与该标记商家集群的相似度中的大于预定阈值的相似度值最大的前预定个数的待识别商家,识别为与该标记商家集群对应类别的风险商家。
- 一种计算设备,包括:至少一个处理器;以及存储器,所述存储器存储指令,当所述指令被所述至少一个处理器执行时,使得所述至少一个处理器执行如权利要求1到7中任一所述的方法。
- 一种非暂时性机器可读存储介质,其存储有可执行指令,所述指令当被执行时使得所述机器执行如权利要求1到7中任一所述的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/348,953 US11379845B2 (en) | 2019-03-14 | 2021-06-16 | Method and device for identifying a risk merchant |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910192832.0 | 2019-03-14 | ||
| CN201910192832.0A CN110033170B (zh) | 2019-03-14 | 2019-03-14 | 识别风险商家的方法及装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/348,953 Continuation US11379845B2 (en) | 2019-03-14 | 2021-06-16 | Method and device for identifying a risk merchant |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020181910A1 true WO2020181910A1 (zh) | 2020-09-17 |
Family
ID=67236012
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/070672 Ceased WO2020181910A1 (zh) | 2019-03-14 | 2020-01-07 | 识别风险商家的方法及装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11379845B2 (zh) |
| CN (1) | CN110033170B (zh) |
| TW (1) | TWI724552B (zh) |
| WO (1) | WO2020181910A1 (zh) |
Families Citing this family (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110033170B (zh) * | 2019-03-14 | 2022-06-03 | 创新先进技术有限公司 | 识别风险商家的方法及装置 |
| CN110610365A (zh) * | 2019-09-17 | 2019-12-24 | 中国建设银行股份有限公司 | 一种识别交易请求的方法和装置 |
| CN110688375B (zh) * | 2019-09-26 | 2022-09-27 | 招商局金融科技有限公司 | 客户渗透分析的方法、装置及计算机可读存储介质 |
| CN110874786B (zh) * | 2019-10-11 | 2022-10-18 | 支付宝(杭州)信息技术有限公司 | 虚假交易团伙识别方法、设备及计算机可读介质 |
| CN110910198A (zh) * | 2019-10-16 | 2020-03-24 | 支付宝(杭州)信息技术有限公司 | 非正常对象预警方法、装置、电子设备及存储介质 |
| CN112788351B (zh) * | 2019-11-01 | 2022-08-05 | 武汉斗鱼鱼乐网络科技有限公司 | 一种目标直播间的识别方法、装置、设备和存储介质 |
| US20230196367A1 (en) * | 2020-05-13 | 2023-06-22 | Paypal, Inc. | Using Machine Learning to Mitigate Electronic Attacks |
| CN115049446A (zh) * | 2021-03-09 | 2022-09-13 | 腾讯科技(深圳)有限公司 | 商户识别方法、装置、电子设备及计算机可读介质 |
| CN113222736A (zh) * | 2021-05-24 | 2021-08-06 | 北京城市网邻信息技术有限公司 | 一种异常用户的检测方法、装置、电子设备及存储介质 |
| CN114693303B (zh) * | 2022-03-03 | 2024-08-30 | 支付宝(杭州)信息技术有限公司 | 对风险交易进行引导举报的方法和装置 |
| CN115080855A (zh) * | 2022-06-30 | 2022-09-20 | 上海掌门科技有限公司 | 风险用户识别方法、设备及计算机可读介质 |
| CN115187387B (zh) * | 2022-07-25 | 2024-02-09 | 山东浪潮爱购云链信息科技有限公司 | 一种风险商家的识别方法及设备 |
| CN120875932B (zh) * | 2025-07-17 | 2026-04-17 | 中鸿云智(浙江)科技有限公司 | 一种数字化基层治理的方法与系统 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105159879A (zh) * | 2015-08-26 | 2015-12-16 | 北京理工大学 | 一种网络个体或群体价值观自动判别方法 |
| US20180025428A1 (en) * | 2016-07-22 | 2018-01-25 | Xerox Corporation | Methods and systems for analyzing financial risk factors for companies within an industry |
| CN109360089A (zh) * | 2018-11-20 | 2019-02-19 | 四川大学 | 贷款风险预测方法及装置 |
| CN109461068A (zh) * | 2018-09-13 | 2019-03-12 | 深圳壹账通智能科技有限公司 | 欺诈行为的判断方法、装置、设备及计算机可读存储介质 |
| CN110033170A (zh) * | 2019-03-14 | 2019-07-19 | 阿里巴巴集团控股有限公司 | 识别风险商家的方法及装置 |
Family Cites Families (31)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7376618B1 (en) * | 2000-06-30 | 2008-05-20 | Fair Isaac Corporation | Detecting and measuring risk with predictive models using content mining |
| US8386377B1 (en) | 2003-05-12 | 2013-02-26 | Id Analytics, Inc. | System and method for credit scoring using an identity network connectivity |
| US8290836B2 (en) * | 2005-06-22 | 2012-10-16 | Early Warning Services, Llc | Identification and risk evaluation |
| US10115153B2 (en) | 2008-12-31 | 2018-10-30 | Fair Isaac Corporation | Detection of compromise of merchants, ATMS, and networks |
| US20110093324A1 (en) | 2009-10-19 | 2011-04-21 | Visa U.S.A. Inc. | Systems and Methods to Provide Intelligent Analytics to Cardholders and Merchants |
| US8738418B2 (en) | 2010-03-19 | 2014-05-27 | Visa U.S.A. Inc. | Systems and methods to enhance search data with transaction based data |
| US9471926B2 (en) | 2010-04-23 | 2016-10-18 | Visa U.S.A. Inc. | Systems and methods to provide offers to travelers |
| US8554653B2 (en) | 2010-07-22 | 2013-10-08 | Visa International Service Association | Systems and methods to identify payment accounts having business spending activities |
| US9747644B2 (en) | 2013-03-15 | 2017-08-29 | Mastercard International Incorporated | Transaction-history driven counterfeit fraud risk management solution |
| US9811830B2 (en) * | 2013-07-03 | 2017-11-07 | Google Inc. | Method, medium, and system for online fraud prevention based on user physical location data |
| US20160027051A1 (en) | 2013-09-27 | 2016-01-28 | Real Data Guru, Inc. | Clustered Property Marketing Tool & Method |
| US9563894B2 (en) * | 2014-03-21 | 2017-02-07 | Ca, Inc. | Controlling eCommerce authentication based on comparing merchant information of eCommerce authentication requests |
| US20160086185A1 (en) | 2014-10-15 | 2016-03-24 | Brighterion, Inc. | Method of alerting all financial channels about risk in real-time |
| CN106529953B (zh) * | 2015-09-15 | 2020-07-31 | 阿里巴巴集团控股有限公司 | 一种对业务属性进行风险识别的方法及装置 |
| US10713660B2 (en) * | 2015-09-15 | 2020-07-14 | Visa International Service Association | Authorization of credential on file transactions |
| US9852427B2 (en) | 2015-11-11 | 2017-12-26 | Idm Global, Inc. | Systems and methods for sanction screening |
| US9818116B2 (en) | 2015-11-11 | 2017-11-14 | Idm Global, Inc. | Systems and methods for detecting relations between unknown merchants and merchants with a known connection to fraud |
| CN105608179B (zh) * | 2015-12-22 | 2019-03-12 | 百度在线网络技术(北京)有限公司 | 确定用户标识的关联性的方法和装置 |
| CN107274042A (zh) * | 2016-04-06 | 2017-10-20 | 阿里巴巴集团控股有限公司 | 一种业务参与对象的风险识别方法及装置 |
| US9888007B2 (en) | 2016-05-13 | 2018-02-06 | Idm Global, Inc. | Systems and methods to authenticate users and/or control access made by users on a computer network using identity services |
| US10366378B1 (en) | 2016-06-30 | 2019-07-30 | Square, Inc. | Processing transactions in offline mode |
| TWI801334B (zh) * | 2017-01-24 | 2023-05-11 | 香港商阿里巴巴集團服務有限公司 | 風險識別方法及裝置 |
| CN106845999A (zh) | 2017-02-20 | 2017-06-13 | 百度在线网络技术(北京)有限公司 | 风险用户识别方法、装置和服务器 |
| CN107679985B (zh) | 2017-09-12 | 2021-01-05 | 创新先进技术有限公司 | 风险特征筛选、描述报文生成方法、装置以及电子设备 |
| CN108053214B (zh) * | 2017-12-12 | 2021-11-23 | 创新先进技术有限公司 | 一种虚假交易的识别方法和装置 |
| CN107958341A (zh) * | 2017-12-12 | 2018-04-24 | 阿里巴巴集团控股有限公司 | 风险识别方法及装置和电子设备 |
| CN108564386B (zh) * | 2018-04-28 | 2020-06-02 | 腾讯科技(深圳)有限公司 | 商户识别方法及装置、计算机设备及存储介质 |
| CN108564467A (zh) * | 2018-05-09 | 2018-09-21 | 平安普惠企业管理有限公司 | 一种用户风险等级的确定方法及设备 |
| CN109118051A (zh) * | 2018-07-17 | 2019-01-01 | 阿里巴巴集团控股有限公司 | 基于网络舆情的风险商户识别及处置方法、装置及服务器 |
| CN109191129A (zh) * | 2018-07-18 | 2019-01-11 | 阿里巴巴集团控股有限公司 | 一种风控方法、系统及计算机设备 |
| CN109461073A (zh) * | 2018-12-14 | 2019-03-12 | 深圳壹账通智能科技有限公司 | 智能识别的风险管理方法、装置、计算机设备及存储介质 |
-
2019
- 2019-03-14 CN CN201910192832.0A patent/CN110033170B/zh active Active
- 2019-09-20 TW TW108134040A patent/TWI724552B/zh not_active IP Right Cessation
-
2020
- 2020-01-07 WO PCT/CN2020/070672 patent/WO2020181910A1/zh not_active Ceased
-
2021
- 2021-06-16 US US17/348,953 patent/US11379845B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105159879A (zh) * | 2015-08-26 | 2015-12-16 | 北京理工大学 | 一种网络个体或群体价值观自动判别方法 |
| US20180025428A1 (en) * | 2016-07-22 | 2018-01-25 | Xerox Corporation | Methods and systems for analyzing financial risk factors for companies within an industry |
| CN109461068A (zh) * | 2018-09-13 | 2019-03-12 | 深圳壹账通智能科技有限公司 | 欺诈行为的判断方法、装置、设备及计算机可读存储介质 |
| CN109360089A (zh) * | 2018-11-20 | 2019-02-19 | 四川大学 | 贷款风险预测方法及装置 |
| CN110033170A (zh) * | 2019-03-14 | 2019-07-19 | 阿里巴巴集团控股有限公司 | 识别风险商家的方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20210312460A1 (en) | 2021-10-07 |
| TWI724552B (zh) | 2021-04-11 |
| CN110033170B (zh) | 2022-06-03 |
| US11379845B2 (en) | 2022-07-05 |
| TW202034256A (zh) | 2020-09-16 |
| CN110033170A (zh) | 2019-07-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| TWI724552B (zh) | 識別風險商家的方法及裝置 | |
| US10748217B1 (en) | Systems and methods for automated body mass index calculation | |
| CN108229330A (zh) | 人脸融合识别方法及装置、电子设备和存储介质 | |
| US11126827B2 (en) | Method and system for image identification | |
| WO2021072876A1 (zh) | 证件图像分类方法、装置、计算机设备及可读存储介质 | |
| JP5785667B1 (ja) | 人物特定システム | |
| US20160125404A1 (en) | Face recognition business model and method for identifying perpetrators of atm fraud | |
| CN110046201A (zh) | 用于处理业务交易的总账科目数据的方法、装置及系统 | |
| WO2017157165A1 (zh) | 信用分数模型训练方法、信用分数计算方法、装置及服务器 | |
| CN108229493A (zh) | 对象验证方法、装置和电子设备 | |
| CN110096859A (zh) | 用户验证方法、装置、计算机设备及计算机可读存储介质 | |
| CN114820476A (zh) | 基于合规性检测的身份证识别方法 | |
| CN115223022B (zh) | 一种图像处理方法、装置、存储介质及设备 | |
| CN110119980A (zh) | 一种用于信贷的反欺诈方法、装置、系统和记录介质 | |
| CN114066564A (zh) | 服务推荐时间确定方法、装置、计算机设备、存储介质 | |
| WO2023029758A1 (zh) | 一种企业经济犯罪侦查方法、系统及设备 | |
| CN111768286B (zh) | 风险预测方法、装置、设备及存储介质 | |
| CN110458024B (zh) | 活体检测方法及装置和电子设备 | |
| CN111447082B (zh) | 关联账号的确定方法、装置和关联数据对象的确定方法 | |
| CN110009056A (zh) | 一种社交账号的分类方法及分类装置 | |
| CN112700235B (zh) | 线下支付用户的识别方法、装置和电子设备 | |
| CN115033880B (zh) | 一种基于互联网的计算机软件管理系统 | |
| CN117010893B (zh) | 基于生物识别的交易风险控制方法、装置和计算机设备 | |
| CN115391400B (zh) | 基于区块链地址的画像识别方法、存储介质和电子设备 | |
| CN114549179B (zh) | 风险名单生成的方法、装置、存储介质及处理器 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20769470 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20769470 Country of ref document: EP Kind code of ref document: A1 |

